Mel-TansWeD-CNN: Respiratory Diseases Detection Framework Through Acoustic Features and Fusion of CNN with Transformer
Keywords:
Respiratory diseases, CNN, Transformer, Encoder, MFCC, Lung soundsAbstract
Respiratory diseases represent a significant global health challenge, which requires a reliable framework for early diagnosis detection framework to assist clinicians in making their clinical decisions. Elderly people and people living in industrial areas are commonly most affected due to pollution. In this study, we designed Mel-TransWed-CNN, a hybrid deep learning (DL) framework that fuses Mel-frequency cepstral coefficients (MFCC) with Transformer and Convolutional Neural Network (CNN) to capture local spectral and global temporal patterns, respectively. We used the standard ICBHI 2017 respiratory sound dataset for experiments. The proposed approach consists of three main stages. First, the raw audio recordings are resampled and normalized to ensure consistency and improve signal quality. Next, MFCC with 20 coefficients are extracted to capture discriminative acoustic characteristics of respiratory sounds. Afterward, the extracted features are fed into the proposed fusion architecture, which combines a CNN with a transformer, where an initial CNN layer learns local spectral patterns, followed by a transformer encoder that models long-range temporal dependencies. A subsequent CNN stage further refines representations before pooling and Softmax-based multiclass classification. Furthermore, the proposed Mel-TransWeD-CNN achieved a mean accuracy of 95% with a standard deviation of 0.64% and a 95% confidence interval of 92.91%–95.43%, while outperforming the baseline classifiers, including Naive Bayes (88.0%), Gaussian Process Classifier (87.5%), Nearest Centroid (77.9%), and AdaBoost (71.7%). The proposed method contributes a practical and reliable automated framework for respiratory disease detection, with potential for early diagnosis and remote patient monitoring.
References
“Automatic adventitious respiratory sound analysis: A systematic review - PubMed.” Accessed: Jul. 17, 2026. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov
/28552969/
W. Sun, F. Zhang, P. Sun, Q. Hu, J. Wang, and M. Zhang, “Respiratory Sound Classification Based on Swin Transformer,” 2023 8th Int. Conf. Signal Image Process. ICSIP 2023, pp. 511–515, 2023, doi: 10.1109/ICSIP57908.2023.10270823.
“Classification of Chronic Obstructive Pulmonary Disease (COPD) Through Respiratory Pattern Analysis - PubMed.” Accessed: Jul. 17, 2026. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/39941243/
S. Sharma, S. Pandey, and D. Shah, “Enhancing Medical Diagnosis with AI: A Focus on Respiratory Disease Detection,” Indian J. Community Med., vol. 48, no. 5, pp. 709–714, 2023, doi: 10.4103/IJCM.IJCM_976_22.
“Respiratory sound classification for crackles, wheezes, and rhonchi in the clinical field using deep learning | Scientific Reports.” Accessed: Jul. 17, 2026. [Online]. Available: https://www.nature.com/articles/s41598-021-96724-7
“Deep Learning in Multi-Class Lung Diseases’ Classification on Chest X-ray Images - PubMed.” Accessed: Jul. 17, 2026. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/35453963/
A. Rahman et al., “Federated learning-based AI approaches in smart healthcare: concepts, taxonomies, challenges and open issues,” Clust. Comput. 2022 264, vol. 26, no. 4, pp. 2271–2311, Aug. 2022, doi: 10.1007/S10586-022-03658-4.
Georgios Petmezas, Grigorios Aris Cheimariotis, “Automated Lung Sound Classification Using a Hybrid CNN-LSTM Network and Focal Loss Function,” Sensors, vol. 22, no. 3, p. 1232, 2022, doi: https://doi.org/10.3390/s22031232.
G. Serbes, S. Ulukaya, and Y. P. Kahya, “An automated lung sound preprocessing and classification system based onspectral analysis methods,” IFMBE Proc., vol. 66, pp. 45–49, 2018, doi: 10.1007/978-981-10-7419-6_8.
K. Kochetov, E. Putin, M. Balashov, A. Filchenkov, and A. Shalyto, “Noise masking recurrent neural network for respiratory sound classification,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 11141 LNCS, pp. 208–217, 2018, doi: 10.1007/978-3-030-01424-7_21/SAVE-RESEARCH.
“Lung Sounds (Breath Sounds): Types, Causes & Treatment.” Accessed: Jul. 17, 2026. [Online]. Available: https://my.clevelandclinic.org/health/symptoms/25193-lung-sounds
J. Acharya and A. Basu, “Deep Neural Network for Respiratory Sound Classification in Wearable Devices Enabled by Patient Specific Model Tuning,” IEEE Trans. Biomed. Circuits Syst., vol. 14, no. 3, pp. 535–544, Jun. 2020, doi: 10.1109/TBCAS.2020.2981172.
P. Faustino, J. Oliveira, and M. Coimbra, “Crackle and wheeze detection in lung sound signals using convolutional neural networks,” Annu. Int. Conf. IEEE Eng. Med. Biol. Soc. IEEE Eng. Med. Biol. Soc. Annu. Int. Conf., vol. 2021, pp. 345–348, 2021, doi: 10.1109/EMBC46164.2021.9630391.
J. Wang, G. Dong, Y. Shen, X. Xu, M. Zhang, and P. Sun, “REDT: a specialized transformer model for the respiratory phase and adventitious sound detection,” Physiol. Meas., vol. 13, no. 2, Feb. 2025, doi: 10.1088/1361-6579/ADAF08.
Samiul Based Shuvo, Taufiq Hasan, “A Multi-Stage Hybrid CNN-Transformer Network for Automated Pediatric Lung Sound Classification,” arXiv:2507.20408, Jul. 2025, Accessed: Jul. 22, 2026. [Online]. Available: https://arxiv.org/pdf/2507.20408
N. Fraihi, O. Karrakchou, S. Member, and M. Ghogho, “Improving Deep Learning-based Respiratory Sound Analysis with Frequency Selection and Attention Mechanism,” arXiv:2507.20052, Jul. 2025, Accessed: Jul. 22, 2026. [Online]. Available: https://arxiv.org/pdf/2507.20052
W. He, Y. Yan, J. Ren, R. Bai, and X. Jiang, “Multi-View Spectrogram Transformer for Respiratory Sound Classification,” ICASSP, IEEE Int. Conf. Acoust. Speech Signal Process. - Proc., pp. 8626–8630, 2024, doi: 10.1109/ICASSP48485.2024.10445825.
Y. Chu, Q. Wang, E. Zhou, L. Fu, Q. Liu, and G. Zheng, “CycleGuardian: a framework for automatic respiratory sound classification based on improved deep clustering and contrastive learning,” Complex Intell. Syst. 2025 114, vol. 11, no. 4, pp. 200-, Mar. 2025, doi: 10.1007/S40747-025-01800-4.
J. Saldanha, S. Chakraborty, S. Patil, K. Kotecha, S. Kumar, and A. Nayyar, “Data augmentation using Variational Autoencoders for improvement of respiratory disease classification,” PLoS One, vol. 17, no. 8, p. e0266467, Aug. 2022, doi: 10.1371/JOURNAL.PONE.0266467.
R. Zulfiqar, F. Majeed, R. Irfan, H. T. Rauf, E. Benkhelifa, and A. N. Belkacem, “Abnormal Respiratory Sounds Classification Using Deep CNN Through Artificial Noise Addition,” Front. Med., vol. 8, p. 714811, Nov. 2021, doi: 10.3389/FMED.2021.714811.
J.-W. Kim, M. Toikkanen, S. Bae, M. Kim, and H.-Y. Jung, “RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification,” May 2024, Accessed: Jul. 22, 2026. [Online]. Available: https://arxiv.org/pdf/2405.02996
A. Roy and U. Satija, “RDLINet: A Novel Lightweight Inception Network for Respiratory Disease Classification Using Lung Sounds,” IEEE Trans. Instrum. Meas., vol. 72, 2023, doi: 10.1109/TIM.2023.3292953.
S. Bae et al., “Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification,” Proc. Annu. Conf. Int. Speech Commun. Assoc. INTERSPEECH, vol. 2023-August, pp. 5436–5440, Dec. 2024, doi: 10.21437/Interspeech.2023-1426.
Y. Zhang, Q. Huang, W. Sun, F. Chen, D. Lin, and F. Chen, “Research on lung sound classification model based on dual-channel CNN-LSTM algorithm,” Biomed. Signal Process. Control, vol. 94, p. 106257, Aug. 2024, doi: 10.1016/J.BSPC.2024.106257.
J.-T. Tzeng et al., “Improving the Robustness and Clinical Applicability of Automatic Respiratory Sound Classification Using Deep Learning-Based Audio Enhancement: Algorithm Development and Validation.,” JMIR AI, vol. 4, no. 1, p. e67239, Mar. 2025, doi: 10.2196/67239.
X. Lu, J. Fang, W. Xiao, and J. Wu, “Research progress and future prospects in intelligent lung sound diagnosis: models, lightweight design, and hardware platform implementation,” Biomed. Tech. (Berl)., vol. 70, no. 6, pp. 483–501, Dec. 2025, doi: 10.1515/BMT-2025-0197.
X. Mu and C. H. Min, “MFCC as Features for Speaker Classification using Machine Learning,” 2023 IEEE World AI IoT Congr. AIIoT 2023, pp. 566–570, 2023, doi: 10.1109/AIIOT58121.2023.10174566.
“Speech Recognition: a review of the different deep learning approaches | AI Summer.” Accessed: Jul. 23, 2026. [Online]. Available: https://theaisummer.com/speech-recognition/
B. Zhu and L. Cao, “Implementation of the MFCC-Streaming Bi-LSTM Based Speech Emotion Recognition Method,” Proc. - 2025 Int. Conf. Artif. Intell. Electr. Electron. Eng. AI3E 2025, pp. 109–114, 2025, doi: 10.1109/AI3E69313.2025.00030.
C. Constantinescu and R. BRAD, “Low-Cost COPD Screening from Respiratory Sounds Using MFCCs with k-NN, SVM and Decision Trees,” Int. J. Integr. Eng., vol. 18, no. 4, pp. 130–138, Jun. 2026, doi: 10.30880/ijie.2026.18.04.009.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 50sea

This work is licensed under a Creative Commons Attribution 4.0 International License.


















