Ranking Pretrained Speech Embeddings in Parkinson's Disease Detection: Does Wav2Vec 2.0 Outperform its 1.0 Version Across Speech Modes and Languages?
Klempir, O.; Skryjova, A.; Tichopad, A.; Krupicka, R.
Show abstract
Speech and language technologies are effective tools for identifying the distinct speech changes associated with Parkinsons disease (PD), enabling earlier and more accurate diagnosis. Recent advancements in self-supervised speech pretraining, particularly with Wav2Vec models, have demonstrated superior performance over traditional feature extraction methods. While Wav2Vec 2.0 has been successfully utilized for PD detection, a rigorous quantitative comparison with Wav2Vec 1.0 is needed to comprehensively evaluate its advantages, limitations, and applicability across different speech modes in PD. This study presents a systematic comparison of Wav2Vec 1.0 and Wav2Vec 2.0 embeddings across three multilingual datasets using various classification approaches in classifying normal (healthy controls; HC) and PD speech. Additionally, both Wav2Vec versions were benchmarked against traditional baseline features across diverse linguistic contexts, including spontaneous speech, non-spontaneous speech, and isolated vowels. A multicriteria TOPSIS approach was employed to rank feature extraction methods, revealing that the Wav2Vec 2.0 consistently excelled across all speech modes, with its first transformer layer demonstrating the best performance for contextual tasks (read text and monologue) and its feature extractor performing best in vowel-based classification. In contrast, the Wav2Vec 1.0, while generally outperformed by the Wav2Vec 2.0, still provided a faster alternative with competitive performance in contextual tasks, highlighting its potential for specific applications, such as federated learning. This comparative analysis furthermore underscores the strengths of each Wav2Vec architecture and informs their optimal use in PD detection.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Time-adaptive Unsupervised Auditory Attention Decoding Using EEG-based Stimulus Reconstruction 94%
- Off-body Sleep Analysis for Predicting Adverse Behavior in Individuals with Autism Spectrum Disorder 92%
- Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach 92%
Similar papers in this journal
- Direct Speech Reconstruction from Sensorimotor Brain Activity with Optimized Deep Learning Models 95%
- 'Are you even listening?' - EEG-based decoding of absolute auditory attention to natural speech 95%
- Linear versus deep learning methods for noisy speech separation for EEG-informed attention decoding 95%
Similar papers in this journal
- Uncertainty in Deep Learning for EEG under Dataset Shifts 93%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 92%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.