Evaluating Large Reasoning Model Performance on Complex Medical Scenarios In The MMLU-Pro Benchmark
Hoyt, R. E.; Knight, D.; Haider, M.; Bajwa, M.
Show abstract
Large language models (LLMs) have emerged as a major force in artificial intelligence, demonstrating remarkable capabilities in natural language processing, comprehension, and text and image generation. Recent advancements have led to the development of LLMs specifically designed for medical applications, showcasing their potential to revolutionize healthcare. These models can analyze complex medical scenarios, assist in diagnoses, and provide treatment recommendations. However, evaluating the accuracy and reliability of LLMs in medicine remains a crucial challenge. The output may not be current and could suffer from inaccurate information, known as hallucinations. In early 2025, DeepSeek R1 was released, which is a large reasoning model (LRM) that includes the "chain of thought" reasoning that made it more transparent than any LLM that preceded it. This study utilized the new MMLU-Pro benchmark, which is a more complex Q&A dataset compared to the Massive Multitask Language Understanding (MMLU). DeepSeek R1 was used to analyze the Q&A dataset primarily for accuracy, but medical scenario Q&As are only one facet of a comprehensive assessment. The study found that DeepSeek R1 had an accuracy rate of 96.3% on 162 medical scenarios after reconciliation with subject matter experts on 23 questions. Our findings contribute to the growing body of knowledge on LLM applications in healthcare and provide insights into the strengths and limitations of DeepSeek R1 in this domain. DeepSeek R1 demonstrates excellent accuracy along with unique transparency. Our analysis also highlights the need for multifaceted evaluation methods that go beyond simple accuracy metrics to ensure the safe and effective deployment of LLMs in medical settings.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 96%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
Similar papers in this journal
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 92%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.