Can Parents and Patients Understand Myopia Using Large Language Model-Based Chatbots?
Panigrahi, S.; Shah, S.; Thakur, S.; Biswas, S.; Verkicharla, P. K.
Show abstract
PurposeThis study aimed to compare the reliability of myopia-related information from AI chatbots using a set of commonly asked questions by parents and patients on myopia, which is an emerging disease of the 21st-century. DesignProspective comparative reliability study MethodsThe study used ChatGPT(OpenAI(2025)GPT-5), Gemini(Gemini 2.0,Google,2025) and DeepSeek (DeepSeek-R1). Twenty myopia-related questions were framed from the perspective of parents and patients, covering general questions, prevention and control, and complications of myopia. Based on their experience in the field of myopia, two senior clinicians, one junior clinician and one researcher(all[≥]3 years of experience in myopia) rated the responses generated by AI chatbots on a 5-point Likert scale(1:very poor, 2:poor, 3: acceptable, 4:good and 5:very good). ResultsOverall, combined rating for tested chatbots had median score of 4("good"). Gemini received significantly lower ratings than other two chatbots (p[≤]0.001), with a median rating of 3("acceptable"). ChatGPT and DeepSeek had median score of 4("good") and there was no significant difference in ratings (p=0.48). Both ChatGPT(66.0%) and DeepSeek(67.5%) had high proportions of "good" and "very good" ratings, compared to Gemini(40.0%). Combined "poor" and "very poor" ratings were highest for Gemini(7.5%), followed by ChatGPT(5.0%) and DeepSeek(4.0%). For general questions on myopia, ChatGPT and DeepSeek were rated "good"; for complications of myopia, ChatGPT was rated as "good", while others were rated "acceptable". ConclusionsChatGPT and DeepSeek demonstrated consistently high-quality responses, while ratings for Gemini were slightly lower but remained adequate. These findings suggest AI chatbots can support patients or parents in understanding myopia.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identifying the Influencing Factors for Cataract Surgery Uptake in Malaysia 95%
- Knowledge and Awareness-based Survey of COVID-19 within the Eye Care Profession in Nepal: Misinformation is Hiding the Truth. 94%
- A psychometric evaluation of the Chinese Impact of Vision Impairment (C-IVI) questionnaire in an adult cohort with high myopia using Rasch Analysis 94%
Similar papers in this journal
- Use of assistive technology to assess distal motor function in subjects with neuromuscular disease 95%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
- Detecting papilloedema as a marker of raised intracranial pressure using artificial intelligence: a systematic review 93%
Similar papers in this journal
- A cross-sectional audit and survey of Open Science and Data Sharing practices at The Montreal Neurological Institute-Hospital 92%
- Development and Validation of a Bedside Scale for Assessing Upper Limb Function Following Stroke: A Methodological Study 91%
- Recommendations from long-term care reports, commissions, and inquiries in Canada 90%
Similar papers in this journal
- Validation of the patient reported outcome measures tool “Catquest” in Odia language 95%
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 92%
- YouTube as an information source during the Coronavirus disease (COVID-19) pandemic: Evaluation of the Turkish and English content 92%
Similar papers in this journal
- Unveiling the Clinical Incapabilities: A Benchmarking Study of GPT-4V(ision) for Ophthalmic Multimodal Image Analysis 95%
- AI-Powered Effective Lens Position Prediction Improves the Accuracy of Existing Lens Formulas 93%
- Performance of DeepSeek-R1 in Ophthalmology: An Evaluation of Clinical Decision-Making and Cost-Effectiveness 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.