Accuracy and error patterns of ChatGPT-4o for real-time English-Nepali voice translation: A cross-sectional field evaluation in rural Nepal
Mandich, A.; Koirala, S.; Westen, S.; Adhikari, S.; Acharya, A.; Shrestha, A.
Show abstract
Language discordance can impede community-based research and health communication where trained interpreters are limited. Although multimodal artificial intelligence systems can provide real-time spoken translation, performance with under-resourced languages during spontaneous field interactions remains poorly characterized. We evaluated ChatGPT-4o during bidirectional English-Nepali voice translation in a community setting near Dhulikhel Hospital, Nepal. In this cross-sectional field study, 30 primarily Nepali-speaking adults were recruited by convenience sampling. ChatGPT-4o mediated conversations using standardized English questions and spontaneous Nepali responses. A bilingual Nepali-English reviewer assessed 485 translated utterances using a 3-point accuracy scale and an inductively developed framework for translation and conversational deviations. Of 485 translations, 282 (58.1%) received the highest accuracy rating, 134 (27.6%) a moderate rating, and 69 (14.2%) the lowest. Mean accuracy was higher for English-to-Nepali than Nepali-to-English translation (2.63 {+/-} 0.53 vs 2.23 {+/-} 0.86); 63 of 69 low-accuracy translations (91.3%) occurred in the Nepali-to-English direction. Among 329 deviation tags, the most frequent were distortion of intended meaning (17.1%), overly formal or unnatural phrasing (14.7%), omission (14.2%), and addition of content (11.5%). Some fluent outputs substantially altered meaning or introduced information not expressed by the speaker. ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses. Accuracy was lower and more variable for Nepali-to-English translation; however, translation direction was confounded with input type because Nepali inputs were spontaneous and English inputs standardized, limiting conclusions about directional performance. These findings support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or participant understanding. As multimodal AI evolves, performance should be reevaluated across languages, real-world conditions, and model versions, with bilingual oversight and community partnership remaining central to responsible use.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- Evaluation of Large Language Models in Medical Examinations:A Scoping Review Protocol 91%
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 91%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 91%
- Human-supervised, large language model-based clinical decision support aligned to national newborn protocols in Kenya: a pragmatic, early-stage evaluation 91%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 91%
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 92%
Similar papers in this journal
- AI-Generated Clinical Summaries: Errors and Susceptibility to Speech and Speaker Variability 93%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.