A Blinded Comparative Evaluation of Clinical and AI-Generated Responses to Otologic Patient Queries
Akinniyi, S.; Jain-Poster, K.; Evangelista, E.; Yoshikawa, N.; Rivero, A.
Show abstract
ObjectiveThe objective of this study is to assess the quality, empathy, and readability of large language model (LLM) responses regarding otologic questions from patients as they compare to verified physician responses in other patient-driven forums. This study aims to predict the potential utility of LLMs in patient-centered communication. Study DesignComparative study SettingsInternet MethodsA sample of 49 otology-related questions posted on Reddit r/AskDocs1 between January 2020 and June 2025 were selected using search terms including "hearing loss," "ear infection," "tinnitus," "ear pain," and "vertigo." Posts were retrieved using Reddits "Top" filter. Each question was answered by a verified doctor on Reddit and three AI LLMs (ChatGPT-4o, ClaudeAI, Google Gemini). Responses were scored by five evaluators. ResultsCommon otologic concerns posed in patient questions were otalgia (38.7%), vertigo (28.6%), tinnitus (24.5%), hearing loss (22.4%), and aural fullness (20.4%). LLM responses were longer than physician responses (mean 145 vs 67 words; p < .05) and rated higher in quality (10.95 vs 9.58), empathy (7.26 vs 5.18), and readability (4.00 vs 3.73); (all p < .05). Evaluators correctly identified AI versus physician responses in 89.4% of cases with higher sensitivity for detecting physician responses (93.5%). By Flesch-Kincaid grade level, ChatGPT produced the most readable content (mean 7.25), while ClaudeAI responses were more complex (11.86; p < .05). ConclusionLLM responses received higher ratings in quality, empathy, and readability than those of physicians in response to a variety of otologic concerns. When appropriately implemented, such systems may enhance access to understandable otologic information and complement clinician-delivered care.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 91%
- Automatic Identification of Tinnitus Malingering Based on Overt and Covert Behavioral Responses During Psychoacoustic Testing 91%
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 90%
Similar papers in this journal
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 91%
- Face Mask Fit Hacks: Improving the Fit of KN95 Masks and Surgical Masks with Fit Alteration Techniques 90%
- What do blind people "see" with retinal prostheses? Observations and qualitative reports of epiretinal implant users 90%
Similar papers in this journal
- Effectiveness of bimodal neuromodulation for tinnitus treatment in a real-world clinical setting in United States: A retrospective chart review 90%
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 90%
- A Comparative Survey of Functional Evidence Use in Hearing and Vision Loss Genetics 87%
Similar papers in this journal
- Performance of ChatGPT in pediatric audiology as rated by students and experts 93%
- Applicability of a short form of the Speech, Spatial und Qualities of Hearing Scale in 97 individuals with Meniere's disease in a multicentric registry 91%
- Artificial Intelligence in laryngeal endoscopy: Systematic Review and Meta-Analysis 90%
Similar papers in this journal
- Great expectations: Aligning visual prosthetic development with implantee needs 91%
- Identification of Risk Factors for Glaucoma Progression in Free-Text Clinical Notes using a Local Small Language Model 90%
- Advancing Question-Answering in Ophthalmology with Retrieval Augmented Generations (RAG): Benchmarking Open-source and Proprietary Large Language Models 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.