Back

SONIVA: Speech recOgNItion Validation in Aphasia

Sanguedolce, G.; Price, C. J.; Brook, S.; Gruia, D. C.; Parkinson, N. V.; Naylor, P. A.; Geranmayeh, F.

2025-06-04 health informatics
10.1101/2025.06.03.25328889 medRxiv
Show abstract

Post-stroke aphasia is a major contributor to language impairment and neuro-disability worldwide, making automated assessment a critical research priority. However, the development of clinically validated automatic speech recognition (ASR) systems remains limited by the lack of large, annotated datasets that capture aphasias heterogeneous and unpredictable manifestations. We introduce SONIVA (Speech recOgNItion Validation in Aphasia), the largest and most richly annotated database of pathological speech to date, comprising recordings from {approx}1,000 stroke survivors (including over 200 longitudinally) and {approx}7,000 age-matched controls. Current annotations include 576 patients (mean age: 61.23 {+/-} 13.23 years; 69.81% male) and 104 controls (mean age: 61.05 {+/-} 12.05 years; 34% male), with rich linguistic coding, orthographic and international phonetic alphabet transcriptions. Foundation models finetuned on SONIVA extract linguistic features that correlate with expert transcriptions (Spearmans r = 0.86 - 0.79; p < 0.0001), while acoustic classifiers achieve 93% stroke classification accuracy. These results position SONIVA as a critical resource that can transform rehabilitation through objective, scalable speech assessment.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.