CFP last date
20 August 2026
Reseach Article

Psychoacoustic Framework for Automated Dysarthria Severity Assessment Via Hybrid Acoustic Ensembles

by Reddipalli Shashank, Gopi Kumar Jha, Nirbhay Singh, Ramesh K. Bhukya
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Number 130
Year of Publication: 2026
Authors: Reddipalli Shashank, Gopi Kumar Jha, Nirbhay Singh, Ramesh K. Bhukya
10.5120/ijca9045ccba6604

Reddipalli Shashank, Gopi Kumar Jha, Nirbhay Singh, Ramesh K. Bhukya . Psychoacoustic Framework for Automated Dysarthria Severity Assessment Via Hybrid Acoustic Ensembles. International Journal of Computer Applications. 187, 130 ( Jul 2026), 39-51. DOI=10.5120/ijca9045ccba6604

@article{ 10.5120/ijca9045ccba6604,
author = { Reddipalli Shashank, Gopi Kumar Jha, Nirbhay Singh, Ramesh K. Bhukya },
title = { Psychoacoustic Framework for Automated Dysarthria Severity Assessment Via Hybrid Acoustic Ensembles },
journal = { International Journal of Computer Applications },
issue_date = { Jul 2026 },
volume = { 187 },
number = { 130 },
month = { Jul },
year = { 2026 },
issn = { 0975-8887 },
pages = { 39-51 },
numpages = {9},
url = { https://ijcaonline.org/archives/volume187/number130/psychoacoustic-framework-for-automated-dysarthria-severity-assessment-via-hybrid-acoustic-ensembles/ },
doi = { 10.5120/ijca9045ccba6604 },
publisher = {Foundation of Computer Science (FCS), NY, USA},
address = {New York, USA}
}
%0 Journal Article
%1 2026-08-01T02:11:43.552850+05:30
%A Reddipalli Shashank
%A Gopi Kumar Jha
%A Nirbhay Singh
%A Ramesh K. Bhukya
%T Psychoacoustic Framework for Automated Dysarthria Severity Assessment Via Hybrid Acoustic Ensembles
%J International Journal of Computer Applications
%@ 0975-8887
%V 187
%N 130
%P 39-51
%D 2026
%I Foundation of Computer Science (FCS), NY, USA
Abstract

Dysarthria is a neurological motor speech disorder that affects speech production and intelligibility through impairments in articulation, phonation, and prosody. Reliable severity assessment is important for clinical diagnosis and long-term monitoring, yet existing automated systems rely heavily on Mel-Frequency Cepstral Coefficients (MFCCs), which may not adequately capture subtle pathological speech variations. This work investigates Bark-Frequency Cepstral Coefficients (BFCCs) as an alternative acoustic representation for dysarthria severity classification. A hybrid 193-dimensional feature vector is constructed by combining BFCCs with Mel-spectrogram, Chromagram, Spectral Contrast, and Tonnetz features to characterize complementary articulatory, spectral, and prosodic properties of dysarthric speech. Experiments are conducted on the TORGO corpus, with Voice Activity Detection (VAD) applied for silence removal and Adaptive Synthetic Sampling (ADASYN) used to address class imbalance. The extracted features are evaluated using a diverse set of machine learning (ML) models, including k-NN, Decision Tree (DT), Support Vector Machines (SVM), PCA, Random Forest (RF), AdaBoost, LogitBoost, CatBoost, LightGBM, XGBoost, and SGD. Results show that BFCC-based hybrid features provide strong discriminative capability across severity levels. Among the evaluated models, RF achieved the highest classification accuracy of 98.33%, followed by LightGBM (98.20%) and XGBoost (97.68%). The results indicate that BFCCs capture pathological speech characteristics more effectively than conventional cepstral representations and, when combined with ensemble learning, enable accurate and robust dysarthria severity classification.

References
  1. Bassam Ali Al-Qatab and Mumtaz Begum Mustafa. Classification of dysarthric speech according to the severity of impairment: an analysis of acoustic features. IEEE Access, 9:18183–18194, 2021.
  2. Muhammad Ridha Anshari, Triando Hamonangan Saragih, Muliadi Muliadi, Dwi Kartini, Fatma Indriani, Hasri Akbar Awal Rozaq, and Oktay Yıldız. Performance comparison of adaboost, lightgbm, and catboost for parkinson’s disease classification using adasyn balancing. Jurnal Teknik Informatika (Jutif), 6(5):3495–3508, 2025.
  3. Tomas Arias-Vergara, Juan Camilo V´asquez-Correa, and Juan Rafael Orozco-Arroyave. Parkinson’s disease and aging: analysis of their effect in phonation and articulation of speech. Cognitive Computation, 9(6):731–748, 2017.
  4. Chitralekha Bhat, Bhavik Vachhani, and Sunil Kumar Kopparapu. Automatic assessment of dysarthria severity level using audio descriptors. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5070–5074. IEEE, 2017.
  5. Ramesh K Bhukya, SR Mahadeva Prasanna, and Biswajit Dev Sarma. Robust methods for text-dependent speaker verification. Circuits, Systems, and Signal Processing, 38(11):5253– 5288, 2019.
  6. Ramesh K Bhukya, Aditya Raj, Anshul Kumar, et al. Asvspoof 2021: Detecting spoofed utterances through hybrid features. APSIPA Transactions on Signal and Information Processing, 14(1), 2025.
  7. Ramesh K Bhukya, Biswajit Dev Sarma, and SR Mahadeva Prasanna. End point detection using speech-specific knowledge for text-dependent speaker verification. Circuits, Systems, and Signal Processing, 37(12):5507–5539, 2018.
  8. Tianqi Chen and Carlos Guestrin. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, pages 785–794, 2016.
  9. Dae-Lim Choi, Bong-Wan Kim, Yong-Ju Lee, Yongnam Um, and Minhwa Chung. Design and creation of dysarthric speech database for development of qolt software technology. In 2011 International Conference on Speech Database and Assessments (Oriental COCOSDA), pages 47–50. IEEE, 2011.
  10. Steven Davis and Paul Mermelstein. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE transactions on acoustics, speech, and signal processing, 28(4):357–366, 1980.
  11. Camille Dingam, Xueying Zhang, Shufei Duan, Haifeng Li, and Xiaoyu Chen. Imbalanced data classification of pathological speech using pca, smote, and expectation maximization. In International Conference in Communications, Signal Processing, and Systems, pages 309–317. Springer, 2021.
  12. Joseph R Duffy et al. Motor speech disorders: Substrates, differential diagnosis, and management. Elsevier Health Sciences, 2012.
  13. Ashit Kumar Dutta and Abdul Rahaman Wahab Sait. A speech disorder detection model using ensemble learning approach. Journal of Disability Research, 3(3):20240026, 2024.
  14. Tiago H Falk, Wai-Yip Chan, and Fraser Shein. Characterization of atypical vocal source excitation, temporal dynamics and prosody for objective measurement of dysarthric word intelligibility. Speech Communication, 54(5):622–631, 2012.
  15. Sowmiyalakshmi Ganesh, Thillai Chithambaram, Nadesh Ramu Krishnan, Durai Raj Vincent, Jayakumar Kaliappan, and Kathiravan Srinivasan. Exploring huntington’s disease diagnosis via artificial intelligence models: a comprehensive review. Diagnostics, 13(23):3592, 2023.
  16. Christopher Harte, Mark Sandler, and Martin Gasser. Detecting harmonic change in musical audio. In Proceedings of the 1st ACM workshop on Audio and music computing multimedia, pages 21–26, 2006.
  17. Haibo He, Yang Bai, Edwardo A Garcia, and Shutao Li. Adasyn: Adaptive synthetic sampling approach for imbalanced learning. In 2008 IEEE international joint conference on neural networks (IEEE world congress on computational intelligence), pages 1322–1328. Ieee, 2008.
  18. Joseph FY Hoh. Laryngeal muscles as highly specialized organs in airway protection, respiration and phonation. In Handbook of behavioral neuroscience, volume 19, pages 13– 21. Elsevier, 2010.
  19. Amlu Anna Joshy and Rajeev Rajan. Automated dysarthria severity classification: A study on acoustic features and deep learning techniques. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 30:1147–1157, 2022.
  20. Banriskhem K Khonglah, Ramesh K Bhukya, and SR Mahadeva Prasanna. Processing degraded speech for text dependent speaker verification. International Journal of Speech Technology, 20(4):839–850, 2017.
  21. Heejin Kim, Mark Hasegawa-Johnson, Adrienne Perlman, Jon R Gunderson, Thomas S Huang, Kenneth L Watkin, Simone Frame, et al. Dysarthric speech database for universal access research. In Interspeech, volume 2008, pages 1741– 1744, 2008.
  22. Martin Lauritzen, Jens Peter Dreier, Martin Fabricius, Jed A Hartings, Rudolf Graf, and Anthony John Strong. Clinical relevance of cortical spreading depression in neurological disorders: migraine, malignant stroke, subarachnoid and intracranial hemorrhage, and traumatic brain injury. Journal of Cerebral Blood Flow & Metabolism, 31(1):17–35, 2011.
  23. Ayoub Malek. Spafe: Simplified python audio features extraction. Journal of Open Source Software, 8(81):4739, 2023.
  24. David Mart´ınez, Eduardo Lleida, Phil Green, Heidi Christensen, Alfonso Ortega, and Antonio Miguel. Intelligibility assessment and speech recognizer word accuracy rate prediction for dysarthric speakers in a factor analysis subspace. ACM Transactions on Accessible Computing (TACCESS), 6(3):1–21, 2015.
  25. Xavier Menendez-Pidal, James B Polikoff, Shirley M Peters, Jennie E Leonzio, and H Timothy Bunnell. The nemours database of dysarthric speech. In Proceeding of Fourth International Conference on Spoken Language Processing. ICSLP’ 96, volume 3, pages 1962–1965. IEEE, 1996.
  26. Belinda Ndlovu, Kudakwashe Maguraushe, and Otis Mabikwa. A comparative analysis of machine learning techniques and explainable ai on voice biomarkers for effective parkinson’s disease prediction. Journal of Information Systems and Informatics, 7(3):2196–2228, 2025.
  27. Juan Rafael Orozco-Arroyave, Juli´an David Arias-Londo˜no, Jes´us Francisco Vargas-Bonilla, Mar´ıa Claudia Gonzalez- R´ativa, and Elmar N¨oth. New spanish speech corpus database for the analysis of people suffering from parkinson’s disease. In Lrec, volume 14, pages 342–347, 2014.
  28. Milton Sarria Paja and Tiago H Falk. Automated dysarthria severity classification for improved objective intelligibility assessment of spastic dysarthric speech. In Thirteenth annual conference of the international speech communication association, pages 62–65, 2012.
  29. B Shunmuga Priya, G Chitra, and R Ramalakshmi. Enhancing parkinson’s disease diagnosis using smote and lightgbm on imbalanced speech dataset. In 2025 International Conference on Computational Robotics, Testing and Engineering Evaluation (ICCRTEE), pages 1–6. IEEE, 2025.
  30. Frank Rudzicz, Aravind Kumar Namasivayam, and Talya Wolff. The torgo database of acoustic and articulatory speech from speakers with dysarthria. Language resources and evaluation, 46(4):523–541, 2012.
  31. Arman Hossain Shanto, Upoma Deb Suchi, Raida Jafar Chowdhury, Afnan Mohammed, and Fairuz Suhala Reza. Machine learning-based approach to improving Parkinson’s disease diagnosis through voice signal analysis. PhD thesis, Brac University, 2025.
  32. Bhavik Vachhani, Chitralekha Bhat, and Sunil Kumar Kopparapu. Data augmentation using healthy speech for dysarthric speech recognition. In Interspeech, pages 471–475, 2018.
  33. Juan Camilo V´asquez-Correa, JR Orozco-Arroyave, Tobias Bocklet, and Elmar Noeth. Towards an automatic evaluation of the dysarthria level of patients with parkinson’s disease. Journal of communication disorders, 76:21–36, 2018.
  34. Eun Jung Yeo, Kwanghee Choi, Sunhee Kim, and Minhwa Chung. Cross-lingual dysarthria severity classification for english, korean, and tamil. In 2022 Asia-Pacific signal and information processing association annual summit and conference (APSIPA ASC), pages 566–574. IEEE, 2022.
  35. Eberhard Zwicker and Hugo Fastl. Psychoacoustics: Facts and models, volume 22. Springer Science & Business Media, 2013.
Index Terms

Computer Science
Information Sciences

Keywords

Dysarthria Severity Detection Dysarthric Speech Speech Disorders Machine Learning Acoustic Features Speech Analysis Classification