Grounding Health AI: Architecture and Evaluation of a Domain-Expert Metabolic Health Agent
Diament, A.; Sapir, G.; Gorodetski, M.; Wolf, A.; Rice, A.; Azouri, D.; Etzion-Fuchs, A.; Gelbard Solodkin, D.; Talmor-Barkan, Y.; Lutsker, G.; Segal, E.; Rossman, H.
Show abstract
General-purpose language models generate fluent health reports that can fabricate derived clinical metrics. In an illustrative comparison on identical two-week CGM and meal data, leading foundation models produced reports with invented MAGE values, inflated meal counts, and unreferenced complication-risk projections: failures invisible to non-expert readers and plausible enough to mislead clinicians. We describe the HPP Personal Health Agent (PHA), a metabolic health agent that grounds generation in four layers: the Human Phenotype Project (HPP), a deep-phenotyped cohort of 13,000+ participants supplying population references and trained predictive models; 21 domain-expert tools and trained-model wrappers that compute clinical metrics and risk predictions; declarative behavioural skills that constrain what the model may claim; and 21 automated evals across 8 categories developed via a test-driven cycle in which each eval encodes a failure mode discovered during iterative development. In a 210-report matrix (14 participants x 3 prompts x 5 system conditions), the gains are largest on the system's primary use case (meal-grounded metabolic reports, the report it was designed for), where the full system raises a deterministic form/provenance score from 0.37 (the same foundation model with no tools or skills) to 0.91; this score measures structural completeness, numerical accuracy, tool grounding, and clinical-language compliance: a necessary condition for trustworthy health reporting, with clinical quality as a complementary axis examined qualitatively. A skills-vs-tools decomposition shows the two layers act on different axes: tools drive numerical accuracy (from about 14% to 90% of reported metrics correct), while the declarative skills add most of the remaining gain in citations, completeness, and structure (tools alone recover only part of the gap, 0.49 from the same 0.37 baseline). The lift generalises beyond the primary use case: to a second metabolic prompt (0.72) and a cardiovascular extension (0.70), each from a 0.37-0.39 baseline. The architecture extends across clinical domains: adding a SCORE2 cardiovascular risk tool and a corresponding skill (with no changes to orchestration, eval harness, or existing tools) produced a cardiovascular risk report from the same system. Trustworthy domain-specialised health AI is a systems design problem: deep-phenotyped cohort data, domain-expert tools and models, and eval-driven development together form a replicable pattern.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 94%
- Zero Shot Health Trajectory Prediction Using Transformer 93%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 93%
Similar papers in this journal
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 92%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 90%
- Clinical trial emulation can identify new opportunities to enhance the regulation of drug safety in pregnancy 89%
Similar papers in this journal
Similar papers in this journal
- Application of Generative Artificial Intelligence to Utilise Unstructured Clinical Data for Acceleration of Inflammatory Bowel Disease Research 90%
- The Medical Action Ontology: A Tool for Annotating and Analyzing Treatments and Clinical Management of Human Disease 88%
- Multimodal surveillance of SARS-CoV-2 at a university enables development of a robust outbreak response framework 86%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.