Evaluating a Large Reasoning Models Performance on Open-Ended Medical Scenarios
Hoyt, R. E.; Knight, D.; Haider, M.; Bajwa, M.
Show abstract
Large language models (LLMs) have emerged as a dominant form of generative artificial intelligence (GenAI) in multiple domains. In early 2025, DeepSeek R1 was released, which is a new large reasoning model (LRM) that includes CoT (CoT) reasoning, Mixture of Experts (MoE), and reinforcement learning. As these technologies continue to improve, evaluating the accuracy and reliability of LLMs and LRMs in medicine remains a crucial challenge. This paper reports on a follow-up study using DeepSeek R1 to evaluate medical scenarios contained in the MMLU-Pro benchmark, an enhanced benchmark designed to evaluate language understanding models across broader and more challenging tasks. In the previously reported study, the accuracy rate was 96% when multiple-choice MMLU-Pro answers were provided. In the current study, we evaluated DeepSeek R1 on 162 medical scenarios, but without multiple-choice answers provided. The overall accuracy was 92%. This approach mirrors a more realistic clinical scenario where the clinician must decide on the most likely diagnosis and differential diagnoses without any clues. Further research is necessary to determine how to deploy LRMs in clinical medicine, given their high accuracy rate, both with and without answers provided.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 93%
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 93%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.