Evaluating Large Language Models for Transparent Quality-of-Care Measurement in Children with ADHD
Bannett, Y.; Pillai, M.; Huang, T.; Luo, I.; Gunturkun, F.; Hernandez-Boussard, T.
Show abstract
ImportanceGuideline-concordant care for young children with attention-deficit/hyperactivity disorder (ADHD) includes recommending parent training in behavior management (PTBM) as first-line treatment. However, assessing guideline adherence through manual chart review is time-consuming and costly, limiting scalable and timely quality-of-care measurement. ObjectiveTo evaluate the accuracy and explainability of large language models (LLMs) in identifying PTBM recommendations in pediatric electronic health record (EHR) notes as a scalable alternative to manual chart review. Design, Setting, and ParticipantsThis retrospective cohort study was conducted in a community-based pediatric healthcare network in California consisting of 27 primary care clinics. The study cohort included children aged 4-6 years with [≥] 2 primary care visits between 2020-2024 and ICD-10 diagnoses of ADHD or ADHD symptoms (n=542 patients). Clinical notes from the first ADHD-related visit were included. A stratified subset of 122 notes, including all cases with model disagreement, was manually annotated to assess model performance in identifying PTBM recommendations and rank model explanations. ExposuresAssessment and plan sections of clinical notes were analyzed using three generative large language models (Claude-3.5, GPT-4o, and LLaMA-3.3-70B) to identify the presence of PTBM recommendations and generate explanatory rationales and documentation evidence. Main Outcomes and MeasuresModel performance in identifying PTBM recommendations (measured by sensitivity, positive predictive value (PPV), and F1-score) and qualitative explainability ratings of model-generated rationales (based on the QUEST framework). ResultsAll three models demonstrated high performance compared to expert chart review. Claude-3.5 showed balanced performance (sensitivity=0.89, PPV=0.95, and F1-score=0.92) and ranked highest in explainability. LLaMA3.3-70B achieved sensitivity=0.91, PPV=0.89, and F1-score=0.90, ranking second for explainability. GPT-4o had the highest PPV [0.97] but lowest sensitivity [0.82], with an F1-score of 0.89 and the lowest explainability ranking. Based on classifications from the best-performing model, Claude-3.5, 26.4% (143/542) of patients had documented PTBM recommendations at their first ADHD-related visit. Conclusions and RelevanceLLMs can accurately extract guideline-concordant clinician recommendations for non-pharmacological ADHD treatment from unstructured clinical notes while providing clear explanations and supporting evidence. Evaluating model explainability as part of LLM implementation for medical chart review tasks can promote transparent and scalable solutions for quality-of-care measurement.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 93%
- Electronic Health Record Documentation of Psychiatric Assessments in Massachusetts Emergency Department and Outpatient Settings During the COVID-19 Pandemic 89%
- If you build it, will they use it? Use of a Digital Assistant for Self-Reporting of COVID-19 Rapid Antigen Test Results during Large Nationwide Community Testing Initiative 89%
Similar papers in this journal
- Measuring Quality-of-Care in Treatment of Children with Attention-Deficit/Hyperactivity Disorder: A Novel Application of Natural Language Processing 98%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 92%
Similar papers in this journal
- Evaluation of a Large Language Model to Identify Confidential Content in Adolescent Encounter Notes 95%
- Clinical features and burden of post-acute sequelae of SARS-CoV-2 infection in children and adolescents: an exploratory EHR-based cohort study from the RECOVER program 90%
- Impacts of school closures on physical and mental health of children and young people: a systematic review 89%
Similar papers in this journal
- Heterogeneity of Diagnosis and Documentation of Post-COVID Conditions in Primary Care: A Machine Learning Analysis 93%
- An AI-Based Chatbot to Support Health-Related Social Needs among Pediatric Primary Care Population: Protocol for a Pilot Randomized Controlled Trial 92%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 92%
Similar papers in this journal
- Diagnostic test accuracy in longitudinal study settings: Theoretical approaches with use cases from clinical practice 90%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 90%
- Large-scale validation of the Prediction model Risk Of Bias ASsessment Tool (PROBAST) using a short form: high risk of bias models show poorer discrimination 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.