A framework for human-artificial intelligence co-learning for disease activity labeling using electronic health records
Yang, Z.; Zhang, Y.; Love, Z.; Animashaun, A.; Zhong, K.; McDermott, G.; Cai, T.; Liao, K. P.
Show abstract
Objective To develop and evaluate a framework for human-AI interaction. This approach, SHARE (Synergistic Human-Agent REasoning system) was designed to support scalable phenotyping of complex outcomes accurately, robustly and reproducibly from real-world electronic health record (EHR) data to support real-world evidence (RWE) generation. Methods and Analysis Using rheumatoid arthritis (RA) disease activity as the use-case, we studied a multi-institutional EHR-based RA cohort of 3,167 patients. Expert reviewers and a disease activity agent labeled notes using the same review guideline. The agent combined embedding-based informative-note filtering, structured evidence extraction, and evidence-based integrated reasoning to assign disease activity categories with supporting evidence, rationale, confidence, and ambiguity flags. To support scalable deployment, we evaluated a budget-tiered configuration using GPT-5 Nano for high-volume evidence extraction, o4-mini for final reasoning, benchmarking against a GPT-5.4 high reasoning effort configuration applied at every step. Note-level discrepancies were adjudicated by reviewers into final co-produced labels that were used to refine labels and inform agent development. The main outcome measure was the mean absolute error (MAE) of the initial and final agent vs the final co-produced labels. The agreement between agent- and reviewer-flagged ambiguous notes, per-note cost and compute time across configurations were also tested. Results Expert reviewers labeled 626 notes from 273 patients; human-AI adjudication revised 127 (20%) of these initial labels and added 60 newly labeled notes, yielding a 686-note co-produced reference. Against this reference, the final agent's accuracy improved from a mean absolute error of 0.406 to 0.291 with co-learning, and its ambiguity flag agreed with expert ambiguity designations with 92.1% accuracy. Applied across the cohort, the agent labeled 101,691 notes; the budget tiered configuration matched the accuracy of GPT-5.4 at high reasoning effort while reducing estimated cost by 69% and compute time by 70%. Conclusion Adopting a framework for human-AI co-learning, SHARE, improved the overall quality of gold-standard labels, identified ambiguous cases for further review, and supported accurate and standardized chart reviews of disease activity at a scale infeasible for manual review. SHARE's resource efficiency provides a transferable approach to incorporate complex phenotypes in RWE studies.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 94%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 93%
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 93%
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 93%
- Artificial Intelligence's Contribution to Biomedical Literature Search: Revolutionizing or Complicating? 92%
- Explainable deep learning for disease activity prediction in chronic inflammatory joint diseases 91%
Similar papers in this journal
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 93%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 93%
- Understanding Data Differences across the ENACT Federated Research Network 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.