MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
Fajardo Vargas, D. E.; Proniakin, O.; Gruber, V. E.; Marinescu, R.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWWe present MO_SCPLOWEDC_SCPLOWPI, a high-dimensional benchmark for evaluating large language models (LLMs) in patient-clinician conversations. Unlike single-turn question-answer (QA) benchmarks, MO_SCPLOWEDC_SCPLOWPI evaluates the medical dialogue across 105 dimensions comprising the medical process, treatment safety, treatment outcomes and doctor-patient communication across a granular, accreditation-aligned rubric. MO_SCPLOWEDC_SCPLOWPI comprises five layers: (1) PO_SCPLOWATIENTC_SCPLOW PO_SCPLOWACKETSC_SCPLOW (synthetic EHR-like ground truth); (2) an AI PO_SCPLOWATIENTSC_SCPLOW instantiated through an LLM with memory and affect; (3) a TO_SCPLOWASKC_SCPLOW MO_SCPLOWATRIXC_SCPLOW spanning encounter reasons (e.g. anxiety, pregnancy, wellness checkup) x encounter objectives (e.g. diagnosis, lifestyle advice, medication advice); (4) an EO_SCPLOWVALUATIONC_SCPLOW FO_SCPLOWRAMEWORKC_SCPLOW with 105 dimensions on a 1-4 scale mapped to the Accreditation Council for Graduate Medical Education (ACGME) competencies; and (5) AI JO_SCPLOWUDGESC_SCPLOW that are calibrated, committee-based LLMs providing scores, flags, and evidence-linked rationales. We evaluate 9 flagship models - Claude Opus 4.1, Claude Sonnet 4, MedGemma, Gemini 2.5 Pro, Llama 3.3 70b Instruct, GPT-5, GPT OSS 120b, o3, Grok-4 - across 366 AI patients and 7,097 conversations using a standardized "vanilla clinician" prompt. For all LLMs, we observe low performance across a variety of dimensions, in particular on differential diagnosis. Our work can help guide future use of LLMs for diagnosis and treatment recommendations.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Modeling physician variability to prioritize relevant medical record information 94%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 94%
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 94%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.