Comparative Performance of agentic AI and Physicians in Taking Clinical History across Leading Large Language Models (LLMs)
Steinbuch, S. M.; de Vos-Hillebrand, L.; O'Neill-Dee, C.; Wu, I.; Steinbuch, H.; Kulcsar, Z.; Ranjan, A.; Verma, A.; Landsberg, J.; Dietrich, D.; Hardin, C. C.; Jain, R. K.; Subudhi, S.
Show abstract
Comprehensive clinical history taking is essential for high-quality care. We hypothesized that large language models (LLMs), guided by a structured agentic framework, can efficiently obtain clinically meaningful patient histories. We developed an iterative prompting system that evaluates relevance and completeness across standard history domains and generates targeted follow-up questions until sufficient detail is obtained. We built a patient-facing application and evaluated it using 52 published case reports and 20 constructed clinical scenarios with simulated patient interactions. The framework was implemented using GPT-4o, Gemini-2.5-Flash-Lite, or Grok-3. After each interaction, the system generated an EHR-ready clinical summary, differential diagnosis, and recommended investigations. Across models, relevant history elements were captured with >85% accuracy and F1 scores, as independently assessed by three blinded physicians, and recommended investigations aligned with those used to establish final diagnoses. These findings support the potential of agentic LLM systems for structured clinical history collection and justify prospective clinical evaluation.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 94%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 94%
- Few shot learning for phenotype-driven diagnosis of patients with rare genetic diseases 93%
Similar papers in this journal
Similar papers in this journal
- Application of Generative Artificial Intelligence to Utilise Unstructured Clinical Data for Acceleration of Inflammatory Bowel Disease Research 91%
- Multimodal surveillance of SARS-CoV-2 at a university enables development of a robust outbreak response framework 88%
- Monitoring of the SARS-CoV-2 Omicron BA.1/BA.2 variant transition in the Swedish population reveals higher viral quantity in BA.2 cases 88%
Similar papers in this journal
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 90%
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 90%
- Genomic Insights for Personalized Care: Motivating At-Risk Individuals Toward Evidence-Based Health Practices 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.