Bridging Language Gaps in Neurology Patient Education Through Large Language Models: a Comparative Analysis of ChatGPT, Gemini, and Claude
Haq, M.; Mushhood Ur Rehman, M.; Derhab, M.; Saeed, R.; Kalia, J.
Show abstract
This study evaluates the capability to translate neurology patient education material using three Large Language Models (LLMs) - ChatGPT-4 Omni, Gemini 1.5 Pro, and Claude 3.5 Sonnet. Five neurological conditions (Bells palsy, multiple sclerosis, stroke, migraine, and epilepsy) were translated from English into Spanish, Urdu, and Arabic. The translations were assessed by physicians using four metrics: accuracy, clarity, comprehensiveness, and readability at a 6th grade level. Results showed that Claude outperformed both ChatGPT and Gemini overall, particularly excelling in Spanish and Urdu translations, while Gemini led in Arabic. All LLMs demonstrated superior performance in Spanish compared to Urdu and Arabic. This study highlights the potential of LLMs in enhancing patient education across languages, while also identifying areas for improvement in translation accuracy and readability.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 93%
Similar papers in this journal
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 92%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 91%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
Similar papers in this journal
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 94%
- Development of a knowledge translation platform for ataxia: Impact on readers and volunteer contributors 93%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 93%
Similar papers in this journal
- Large language models (GPT-5, Grok-4, Claude Opus 4.1, Gemini 2.5 Pro) achieved textbook-level accuracy on the Japanese medical licensing examination by 2025: A comparative study 92%
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 91%
- Choice of Intraoperative Ultrasound adjuncts for Brain Tumor Surgery 90%
Similar papers in this journal
- YouTube as an information source during the Coronavirus disease (COVID-19) pandemic: Evaluation of the Turkish and English content 92%
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 89%
- “This is a quiz” Premise Input: A Key to Unlocking Higher Diagnostic Accuracy in Large Language Models 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.