Autonomous generation of decision-grade clinical evidence
Yang, S.; Wu, J.; Xie, H.; Xin, Z.; Wang, W.
Show abstract
Medical practice is bottlenecked by the slow production of high-quality clinical evidence. Despite progress in automating selected stages, autonomous conduct of the entire research life cycle remains beyond reach. Here we present OpenEBM, the first autonomous system to generate decision-grade clinical evidence by conducting evidence-synthesis research end to end. To enable and evaluate this, we develop OpenEBM-Corpus, a foundation resource of expert-annotated research trajectories that enables training of a specialist model, and OpenEBM-Bench, a multidisciplinary benchmark that evaluates the entire research life cycle. Our compact specialist model generates valid clinical evidence in 90.7% of end-to-end evaluations and matches expert performance across the research trajectory, whereas GPT-5 falls to 3.8% as failures propagate through dependent stages. In blinded evaluations across clinical domains, independent evaluators prefer OpenEBM at multiple stages and cannot distinguish its reasoning traces from expert-conducted work above chance. Applied to a question left unresolved by current guidelines, OpenEBM produces de novo evidence addressing the efficacy and safety of neoadjuvant chemotherapy for locally advanced rectal cancer. OpenEBM brings within reach the founding aspiration of evidence-based medicine and establishes a paradigm for scalable evidence generation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 94%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 92%
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 92%
Similar papers in this journal
Similar papers in this journal
- Zero-shot drug repurposing with geometric deep learning and clinician centered design 93%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 91%
- Doctors and Nurses Social Media Ads Reduced Holiday Travel and COVID-19 infections: A cluster randomized controlled trial in 13 States 90%
Similar papers in this journal
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 91%
- Results dissemination from clinical trials conducted at German university medical centres was delayed and incomplete 91%
- Strength of Statistical Evidence for the Efficacy of Cancer Drugs: A Bayesian Re-Analysis of Trials Supporting FDA Approval 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.