Performance Analysis of Speech Recognition Models in Automated Scoring of the QuickSIN Test
Hassanpour, A.; Jiang, Y.; Folkeard, P.; Macpherson, E.; Scollie, S. D.; Parsa, V.
Show abstract
PurposeBest practices in audiology recommend assessing speech understanding in noisy environments, especially for those with communication difficulties. Speech-in-noise (SiN) assessments such as the QuickSIN are used for validating signal processing in hearing aids (HAs) and are linked to HA satisfaction. This project seeks to enhance QuickSIN test efficiency by applying recent advancements in automatic speech recognition (ASR) technologies. MethodTwenty-three adults with sensorineural hearing loss were fitted bilaterally with Unitron Moxi HAs and were administered the QuickSIN test in low and high reverberation environments. Testing was performed with two different HA programs: an omnidirectional program and a fixed directional microphone program. QuickSIN sentences were presented from 0{degrees} azimuth and competing babble from either 0{degrees}, laterally from 90{degrees} or 270{degrees}, or simultaneously from 90{degrees}, 180{degrees}, and 270{degrees} azimuths. Participants verbal responses to QuickSIN stimuli were scored by an audiologist and were recorded in parallel for offline transcription and scoring by ASR models from Amazon, Microsoft, NVIDIA, and Picovoice. The ASR-derived QuickSIN scores were compared to the corresponding audiologist-derived scores. ResultsRepeated Measures ANOVA results revealed that all ASR models overestimated the QuickSIN scores across most test conditions. Bland-Altman analyses showed that the Amazon ASR model had the least bias and the narrowest range for the limits of agreement, in comparison to the manual scoring by an experienced audiologist. ConclusionsSome ASR models, such as Amazon, demonstrated performance comparable to that of an audiologist in automatically scoring QuickSIN tests. However, further refinements are necessary to increase the robustness of the ASR models in scoring low SNR loss test conditions.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Air-Conduction And Bone-Conduction Reference Threshold Levels – A Multicenter Study 97%
- Magnified interaural level differences enhance spatial release from masking in bilateral cochlear implant users 97%
- Effects of face masks on acoustic analysis and speech perception: Implications for peri-pandemic protocols 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Auditory tests for characterizing hearing deficits in listeners with various hearing abilities: The BEAR test battery 97%
- Speech-driven Facial Animations Improve Speech-in-Noise Comprehension of Humans 95%
- Neural correlates of masked and unmasked tones: psychoacoustics and late auditory evoked potentials (LAEPs) 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.