Back

Bridging Language Gaps in Neurology Patient Education Through Large Language Models: a Comparative Analysis of ChatGPT, Gemini, and Claude

Haq, M.; Mushhood Ur Rehman, M.; Derhab, M.; Saeed, R.; Kalia, J.

2024-09-24 neurology
10.1101/2024.09.23.24314229 medRxiv
Show abstract

This study evaluates the capability to translate neurology patient education material using three Large Language Models (LLMs) - ChatGPT-4 Omni, Gemini 1.5 Pro, and Claude 3.5 Sonnet. Five neurological conditions (Bells palsy, multiple sclerosis, stroke, migraine, and epilepsy) were translated from English into Spanish, Urdu, and Arabic. The translations were assessed by physicians using four metrics: accuracy, clarity, comprehensiveness, and readability at a 6th grade level. Results showed that Claude outperformed both ChatGPT and Gemini overall, particularly excelling in Spanish and Urdu translations, while Gemini led in Arabic. All LLMs demonstrated superior performance in Spanish compared to Urdu and Arabic. This study highlights the potential of LLMs in enhancing patient education across languages, while also identifying areas for improvement in translation accuracy and readability.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.