Back

MedPI: Evaluating AI Systems in Medical Patient-facing Interactions

Fajardo Vargas, D. E.; Proniakin, O.; Gruber, V. E.; Marinescu, R.

2026-01-01 health informatics
10.64898/2025.12.24.25342982 medRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWWe present MO_SCPLOWEDC_SCPLOWPI, a high-dimensional benchmark for evaluating large language models (LLMs) in patient-clinician conversations. Unlike single-turn question-answer (QA) benchmarks, MO_SCPLOWEDC_SCPLOWPI evaluates the medical dialogue across 105 dimensions comprising the medical process, treatment safety, treatment outcomes and doctor-patient communication across a granular, accreditation-aligned rubric. MO_SCPLOWEDC_SCPLOWPI comprises five layers: (1) PO_SCPLOWATIENTC_SCPLOW PO_SCPLOWACKETSC_SCPLOW (synthetic EHR-like ground truth); (2) an AI PO_SCPLOWATIENTSC_SCPLOW instantiated through an LLM with memory and affect; (3) a TO_SCPLOWASKC_SCPLOW MO_SCPLOWATRIXC_SCPLOW spanning encounter reasons (e.g. anxiety, pregnancy, wellness checkup) x encounter objectives (e.g. diagnosis, lifestyle advice, medication advice); (4) an EO_SCPLOWVALUATIONC_SCPLOW FO_SCPLOWRAMEWORKC_SCPLOW with 105 dimensions on a 1-4 scale mapped to the Accreditation Council for Graduate Medical Education (ACGME) competencies; and (5) AI JO_SCPLOWUDGESC_SCPLOW that are calibrated, committee-based LLMs providing scores, flags, and evidence-linked rationales. We evaluate 9 flagship models - Claude Opus 4.1, Claude Sonnet 4, MedGemma, Gemini 2.5 Pro, Llama 3.3 70b Instruct, GPT-5, GPT OSS 120b, o3, Grok-4 - across 366 AI patients and 7,097 conversations using a standardized "vanilla clinician" prompt. For all LLMs, we observe low performance across a variety of dimensions, in particular on differential diagnosis. Our work can help guide future use of LLMs for diagnosis and treatment recommendations.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.