MedEvalArena: A Self-Generated, Peer-Judged Benchmark for Medical Reasoning
Prem, P.; Shidara, K.; Kuppa, V.; Wheeler, E.; Liu, F.; Alaa, A.; Bernardo, D.
Show abstract
Large Language Models (LLMs) demonstrate strong performance at medical specialty board multiple-choice question (MCQ) answering, however, underperform in more complex medical reasoning scenarios. This gap indicates a need for improving both LLM medical reasoning and evaluation paradigms. We introduce MedEvalArena, a framework in which LLMs engage in a symmetric round-robin format. Each model generates challenging board-style medical MCQs, then serves in an ensemble LLM-as-judge bench to adjudicate validity of generated questions, and finally completes the validated exam as an examinee. We compared performance of leading LLMs across the OpenAI, Grok, Gemini, Claude, Kimi, and DeepSeek families on both question generation validity and exam-taking performance. Across frontier models, we observe no statistically significant differences in exam-taking performance, suggesting convergence in current medical reasoning ability across frontier LLMs for question-answering. We observed higher question validity rates in questions generated by OpenAI, Gemini, and Claude frontier models compared to Kimi, Grok, and DeepSeek models. When jointly considering accuracy and inference cost, multiple frontier models lie on the Pareto frontier with no single model dominating across both dimensions. MedEvalArena provides a dynamic and scalable framework for benchmarking LLM medical reasoning.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 92%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 91%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 91%
Similar papers in this journal
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 92%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 92%
Similar papers in this journal
- Actionability of Synthetic Data in a Heterogeneous and Rare Healthcare Demographic; Adolescents and Young Adults (AYAs) with Cancer 91%
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 89%
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 89%
Similar papers in this journal
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 92%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 91%
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.