On Estimating Age and Gender from Parkinson's Disease Diagnostic-Oriented Recordings Using Wav2Vec 2.0
Klempir, O.; Tichopad, A.; Krupicka, R.
Show abstract
Can self-supervised speech foundation models (SFMs) be used for automatic patient metadata extraction, even when no prior demographic information is available and speech is affected by pathology? SFMs show strong cross-task generalization, yet it remains unclear to what extent demographic attributes such as age and gender are intrinsically encoded, particularly in pathological speech. This study evaluated the capability of a pretrained SFM Wav2Vec 2.0 to estimate age and gender across healthy controls (HC), Parkinsons disease (PD) subjects, and related parkinsonian syndromes (multiple system atrophy, progressive supranuclear palsy), without exposing the model to any data from the evaluated datasets. A frozen, publicly available Wav2Vec 2.0 model was used to extract speech representations from three independent multilingual datasets. No machine learning model was trained or fine-tuned on the target data. The analysis solely assessed information already present in the pretrained embeddings. Multiple speech tasks (read text, diadochokinesis, sustained vowels) and diagnostic groups were evaluated using gender accuracy, correlations with true age, chi-square tests, and group-level analyses. Gender estimation achieved consistently high accuracy (min. 94%, up to 100%) across datasets and tasks. Age estimation showed significant correlations with true age for read speech, including PD speakers. Analyses of vowel data demonstrated preserved gender distributions but systematic age bias across all diagnostic groups. Without task-specific training, pretrained Wav2Vec 2.0 embeddings robustly encode gender and preserve age-related structure in connected, including pathological, speech, whereas age estimation from isolated vowel phonation remains unreliable. These findings highlight the demographic robustness of SFMs and the need to monitor age-related bias in clinical applications.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Online speech synthesis using a chronically implanted brain-computer interface in an individual with ALS 94%
- Feasibility of decoding covert speech in ECoG with aTransformer trained on overt speech 94%
- Deep Learning Restores Speech Intelligibility in Multi-Talker Interference for Cochlear Implant Users 93%
Similar papers in this journal
- Uncertainty in Deep Learning for EEG under Dataset Shifts 92%
- Subtle anomaly detection in MRI brain scans: Application to biomarkers extraction in patients with de novo Parkinson’s disease 91%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 90%
Similar papers in this journal
- Speech-driven Facial Animations Improve Speech-in-Noise Comprehension of Humans 94%
- The Effect on Speech-in-Noise Perception of Real Faces and Synthetic Faces Generated with either Deep Neural Networks or the Facial Action Coding System 92%
- Amplified cortical tracking of word-level features of continuous competing speech in older adults 91%
Similar papers in this journal
- Time-adaptive Unsupervised Auditory Attention Decoding Using EEG-based Stimulus Reconstruction 92%
- Off-body Sleep Analysis for Predicting Adverse Behavior in Individuals with Autism Spectrum Disorder 89%
- Hyperbolic graph embedding of MEG brain networks to study brain alterations in individuals with subjective cognitive decline 88%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.