Towards Metacognitive Clinical Reasoning: Benchmarking MD-PIE Against State-of-the-Art LLMs in Medical Decision-Making
Esteitieh, Y.; Mandal, S.; Laliotis, G.
Show abstract
The ability of large language models (LLMs) to perform clinical reasoning and cognitive tasks within medicine remains a critical measure of their overall capabilities in decision-making, with significant implications for patient outcomes and healthcare efficiency. Current AI models often face limitations in real-world clinical environments, including variability in performance, a lack of domain-specific knowledge, and black-box reasoning processes. In this study, we introduce a novel PIE framework, named MD-PIE, which emulates cognitive and reasoning abilities in medical reasoning and decision-making. We benchmark our framework and baseline methods both quantitatively and qualitatively using state-of-the-art LLMs, comparing them against OpenAIs o1, Gemini 2.0 Flash Thinking, and DeepSeek V3 across diverse benchmarks. Our results demonstrate that MD-PIE surpasses existing models in differential diagnosis and reasoning accuracy across diverse medical benchmarks. This study underscores its potential to improve clinical decision-making through adaptive and collaborative design. Future research should focus on larger benchmarks and real-world validation to confirm its reliability and effectiveness in varied clinical scenarios.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 96%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 96%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 95%
Similar papers in this journal
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Modeling physician variability to prioritize relevant medical record information 95%
- Evaluation of Patient-Level Retrieval from Electronic Health Record Data for a Cohort Discovery Task 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.