How does DeepSeek-R1 perform on USMLE?
Faray de Paiva, L.; Luijten, G.; Puladi, B.; Egger, J.
Show abstract
DeepSeek, a Chinese artificial intelligence company, released its first free chatbot app based on its DeepSeek-R1 model. DeepSeek provides its models, algorithms, and training details to ensure transparency and reproducibility. Their new model is trained with reinforcement learning, allowing it to learn through interactions and feedback rather than relying solely on supervised learning. Reports showcase that DeepSeeks model shows competitive performances against established large language models (LLMs) such as Anthropics Claude and OpenAIs GPT-4o on established benchmarks in language understanding, mathematics (AIME 2024) and programming (Codeforces) while trained at a fraction of the costs. Additionally, running inference shows significantly lower costs, leading to DeepSeek surpassing ChatGPT as the most downloaded free app on the American iOS App Store. This development contributed to a nearly 17% drop in Nvidias share price, resulting in the most significant one-day loss in U.S. history, amounting to nearly $600 billion. The open-source models also bring a significant shift in the healthcare system, allowing cost-efficient medical LLMs to be deployed within hospital networks. To understand its performance in the healthcare sector, we analyse the new DeepSeek-R1 model on the United States Medical Licensing Examination (USMLE) and compare it to ChatGPT.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 96%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 95%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 95%
Similar papers in this journal
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 93%
- Large language models (GPT-5, Grok-4, Claude Opus 4.1, Gemini 2.5 Pro) achieved textbook-level accuracy on the Japanese medical licensing examination by 2025: A comparative study 92%
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets 93%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 90%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 90%
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.