Uncovering the important acoustic features for detecting vocal fold paralysis with explainable machine learning
Low, D. M.; Randolph, G. W.; Rao, V.; Ghosh, S. S.; Song, P. C.
Show abstract
IntroductionDetecting voice disorders from voice recordings could allow for frequent, remote, and low-cost screening before costly clinical visits and a more invasive laryngoscopy examination. Our goals were to detect unilateral vocal fold paralysis (UVFP) from voice recordings using machine learning, to identify which acoustic variables were important for prediction to increase trust, and to determine model performance relative to clinician performance. MethodsPatients with confirmed UVFP through endoscopic examination (N=77) and controls with normal voices matched for age and sex (N=77) were included. Voice samples were elicited by reading the Rainbow Passage and sustaining phonation of the vowel "a". Four machine learning models of differing complexity were used. SHapley Additive exPlanations (SHAP) was used to identify important features. ResultsThe highest median bootstrapped ROC AUC score was 0.87 and beat clinicians performance (range: 0.74 - 0.81) based on the recordings. Recording durations were different between UVFP recordings and controls due to how that data was originally processed when storing, which we can show can classify both groups. And counterintuitively, many UVFP recordings had higher intensity than controls, when UVFP patients tend to have weaker voices, revealing a dataset-specific bias which we mitigate in an additional analysis. ConclusionWe demonstrate that recording biases in audio duration and intensity created dataset-specific differences between patients and controls, which models used to improve classification. Furthermore, clinicians ratings provide further evidence that patients were over-projecting their voices and being recorded at a higher amplitude signal than controls. Interestingly, after matching audio duration and removing variables associated with intensity in order to mitigate the biases, the models were able to achieve a similar high performance. We provide a set of recommendations to avoid bias when building and evaluating machine learning models for screening in laryngology.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- A Machine-Learning Based Objective Measure for ALS Disease Severity 93%
- Automatic Identification of Tinnitus Malingering Based on Overt and Covert Behavioral Responses During Psychoacoustic Testing 93%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 92%
Similar papers in this journal
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 90%
- Effectiveness of bimodal neuromodulation for tinnitus treatment in a real-world clinical setting in United States: A retrospective chart review 90%
- Prospective validation of smartphone-based heart rate and respiratory rate measurement algorithms 90%
Similar papers in this journal
- Integrating a host transcriptomic biomarker with a large language model for diagnosis of lower respiratory tract infection 90%
- Timbral effects on consonance illuminate psychoacoustics of music evolution 89%
- Deep neural network models reveal interplay of peripheral coding and stimulus statistics in pitch perception 89%
Similar papers in this journal
- Acoustic parameter combinations underlying mapping of pseudoword sounds to multiple domains of meaning: representational similarity analyses and machine-learning models 94%
- Effects of face masks on acoustic analysis and speech perception: Implications for peri-pandemic protocols 93%
- Focality of sound source placement by higher (9th) order ambisonics and perceptual effects of spectral reproduction errors 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.