Back

The accuracy and repeatability of OpenEvidence on complex medical subspecialty scenarios: a pilot study

Jagarapu, J.; Babata, K.; Chamarthi, S.; Hoyt, R. E.

2025-12-04 health informatics
10.64898/2025.11.29.25341091 medRxiv
Show abstract

OpenEvidence is a popular artificial intelligence (AI) based medical search engine that generates evidence-based answers. It includes a quick search engine method (OE) that takes only seconds to respond, along with a limited number of references. In mid-2025, the platform introduced "Deep Consult" (DC), which takes several minutes to respond and provides more comprehensive answers with additional references. OpenEvidence scored 100% on USMLE-type multiple-choice questions, but it has not been tested on more complex medical scenarios. We tested the OE and DC models using questions primarily derived from medical specialty board exams, specifically, the MedXpertQA dataset. In a prior published study, this dataset was evaluated with eleven large language models (LLMs), and the results indicated poor accuracy (14-46%) for all LLMs. We evaluated the performance of OpenEvidence on a sample of the MedXpertQA dataset, comprising 100 medical subspecialty scenarios and using two independent evaluators. The highest accuracy for DC was 41%, and for OE, 34%. Repeatability testing revealed an evaluator concordance rate of 77% for OE and 72% for DC.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.