A Multi-Agent Large Language Model Reasoning Engine for Early Detection of Pediatric Growth Disorders
Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.
Show abstract
Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluation of a Large Language Model to Identify Confidential Content in Adolescent Encounter Notes 90%
- Clinical features and burden of post-acute sequelae of SARS-CoV-2 infection in children and adolescents: an exploratory EHR-based cohort study from the RECOVER program 88%
- Acute upper airway disease in children with the omicron (B.1.1.529) variant of SARS-CoV-2: a report from the National COVID Cohort Collaborative (N3C) 86%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
- Measuring Quality-of-Care in Treatment of Children with Attention-Deficit/Hyperactivity Disorder: A Novel Application of Natural Language Processing 92%
- Developing and Evaluating Pediatric Phecodes (Peds-Phecodes) for High-Throughput Phenotyping Using Electronic Health Records 91%
Similar papers in this journal
- Extraction of Crohn's Disease Clinical Phenotypes from Clinical Text Using Natural Language Processing 91%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 90%
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 88%
Similar papers in this journal
- Evidence Aggregator: AI reasoning applied to rare disease diagnostics 92%
- ClinPhen extracts and prioritizes patientphenotypes directly from medical records to accelerate genetic disease diagnosis 91%
- Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models 90%
Similar papers in this journal
- Systematic identification of rare disease patients in electronic health records enables evaluation of clinical outcomes 92%
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 91%
- A Syndromic Surveillance Tool to Detect Anomalous Clusters of COVID-19 Symptoms in the United States 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.