Accents Still Confuse AI: Systematic Errors in Speech Transcription and LLM-Based Remedies
Fatapour, Y.; Samaan, J. S.; Kuchi, A.; Srinivasan, A. P.; Fatapour, S.; Liu, H.; Berkowitz, J. S.; Tsang, K.; Zietz, M.; Friedrich, N.; Srinivasan, N.; Thangaratnam, S.; King, R.; Czarny, R.; Nguyen, T.; Yeo, Y. H. S.; Kim, H.; Lee, Y.-T.; Wongjarupong, N.; Abiri, A.; Tatonetti, N. P.
Show abstract
Accurate and timely documentation in the electronic health record (EHR) is essential for delivering safe and effective patient care. AI-enabled medical tools powered by automatic speech recognition (ASR) offer to streamline this process by transcribing clinical conversations directly into structured notes. However, a critical challenge in deploying these technologies at scale is their variable performance across speakers with diverse accents, which leads to transcription inaccuracies, misinterpretation, and downstream clinical risks. We measured transcription accuracy of Whisper and WhisperX on clinical texts across native and non-native English speakers and found that both models have significantly higher errors for non-native speakers. Fortunately, we found that post-processing the transcripts using GPT-4o recovers the lost accuracy. Our findings indicate that using a chained model approach, WhisperX-GPT, will enhance transcription quality significantly and reduce errors associated with accented speech. We make all code, models, and pipelines freely available.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 94%
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 93%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 93%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 93%
- Development and Validation of a Machine Learning Model Integrated with the Clinical Workflow for Inpatient Discharge Date Prediction 89%
Similar papers in this journal
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 93%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 91%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.