Evaluating Large Language Models for ADHD Education: A Comparative Study of ChatGPT 5, DeepSeek V3, and Grok 4
HAN, X.; Xing, R.; Zhou, M.
Show abstract
BackgroundChildren with attention-deficit/hyperactivity disorder (ADHD) often face barriers to participating in organized sports, particularly when physical education (PE) is delivered by outsourced coaches with limited training in disability inclusion. Meanwhile, large language models (LLMs) such as ChatGPT, DeepSeek, and Grok are increasingly used to generate educational content, yet their readability, stability, and accuracy for non-specialist educators remain unclear. MethodsThis study systematically compared three advanced LLMs, ChatGPT-5, DeepSeek V3, and Grok 4, using identical prompts related to ADHD definitions, symptoms, and medication-exercise interactions. Thirty responses per model were collected and analyzed for content accuracy, readability (Flesch-Kincaid Reading Ease, Grade Level, and SMOG), and lexical complexity. ResultsAll models aligned with DSM-5 in describing ADHD but differed in emphasis and stability. DeepSeek V3 produced the broadest and most variable outputs, Grok 4 showed the greatest consistency and clinical structure, and ChatGPT-5 generated concise and strengths-based explanations. However, all models exhibited high reading levels (FKGL > 12), exceeding recommended public-health standards. ConclusionWhile LLMs demonstrate strong potential for generating ADHD-related educational materials, their current readability and stability limitations restrict accessibility for non-specialist educators. Future work should focus on optimizing prompt design and language calibration to enhance usability in inclusive education contexts.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Use of assistive technology to assess distal motor function in subjects with neuromuscular disease 93%
- Feasibility and preliminary efficacy of an online mindful walking intervention among COVID-19 long haulers: A mixed method study including daily diary surveys 92%
- How can digital citizen science approaches improve ethical smartphone use surveillance among youth: traditional surveys versus ecological momentary assessments 91%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 94%
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 93%
- Remote working in mental health services: a rapid umbrella review of pre-COVID-19 literature 92%
Similar papers in this journal
- Design and Formative Evaluation of a Voice-based Virtual Coach for Problem-Solving Treatment 92%
- Evaluating the Clinical Feasibility of an Artificial Intelligence-Powered Clinical Decision Support System: A Longitudinal Feasibility Study 92%
- Exploring Patient and Staff Experiences of Video Consultations During COVID-19 in an English Outpatient Care Setting: Secondary Data Analysis of Routinely Collected Feedback Data 92%
Similar papers in this journal
- How should job crafting interventions be implemented to make their effects last? Study protocol of group concept mapping 92%
- Impostor Phenomenon in the Nutrition and Dietetics Profession: An Online Cross-Sectional Survey 91%
- A Cohort Study: Evaluating Self-Efficacy in Adolescents Attending a Tailored Youth-Informed Breastfeeding Program 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.