Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting
Hayat, H.; Kudrautsau, M.; Makarov, E.; Melnichenko, V.; Tsykunou, T.; Varaksin, P.; Pavelle, M.; Oskowitz, A. Z.
Show abstract
BackgroundGlobally we face a projected shortage of 11 million healthcare practitioners by 2030, and administrative burden consumes 50% of clinical time. Artificial intelligence (AI) has the potential to help alleviate these problems. However, no end-to-end autonomous large language model (LLM)-based AI system has been rigorously evaluated in real-world clinical practice. In this study, we evaluated whether a multi-agent LLM-based AI framework can function autonomously as an AI doctor in a virtual urgent care setting. MethodsWe retrospectively compared the performance of the multi-agent AI system Doctronic and board-certified clinicians across 500 consecutive urgent-care telehealth encounters. The primary end points: diagnostic concordance, treatment plan consistency, and safety metrics, were assessed by blinded LLM-based adjudication and expert human review. ResultsThe top diagnosis of Doctronic and clinician matched in 81% of cases, and the treatment plan aligned in 99.2% of cases. No clinical hallucinations occurred (e.g., diagnosis or treatment not supported by clinical findings). In an expert review of discordant cases, AI performance was superior in 36.1%, and human performance was superior in 9.3%; the diagnoses were equivalent in the remaining cases. ConclusionsIn this first large-scale validation of an autonomous AI doctor, we demonstrated strong diagnostic and treatment plan concordance with human clinicians. These findings indicate that multi-agent AI systems can achieve comparable clinical decision-making to human providers and offer a potential solution to healthcare workforce shortages. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=117 SRC="FIGDIR/small/25331406v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@c00e55org.highwire.dtl.DTLVardef@ecfe38org.highwire.dtl.DTLVardef@1263910org.highwire.dtl.DTLVardef@6c6ce7_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 97%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 95%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 95%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 93%
Similar papers in this journal
- Evaluation of Patient-Level Retrieval from Electronic Health Record Data for a Cohort Discovery Task 93%
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 93%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 93%
Similar papers in this journal
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 94%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 94%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.