AI vs Human Performance in Conversational Hospital-Based Neurological Diagnosis
Sorka, M.; Gorenshtein, A.; Abramovitch, H.; Soontrapa, P.; Shelly, S.; Aran, D.
Show abstract
BackgroundMost evaluations of artificial intelligence (AI) in medicine rely on static, multiple-choice benchmarks that fail to capture the dynamic, sequential nature of clinical diagnosis. While conversational AI has shown promise in telemedicine, these systems rarely test the iterative decision-making process in which clinicians gather information, order tests, and refine diagnoses. MethodsWe developed DiagnosticXchange, a web-based platform simulating realistic clinical interactions between providers and specialist consultants. A nurse agent responds to requests from human physicians or AI systems acting as diagnosticians. Sixteen neurological diagnostic challenges of varying complexity were drawn from diverse educational and peer-reviewed sources. We evaluated 14 neurologists at different training stages and multiple state-of-the-art large language models (LLMs) using efficiency metrics, including: diagnostic accuracy, procedural cost efficiency (based on CPT codes and hospital pricing), and time to diagnosis (using actual procedure durations). We also developed Gregory, a specialized multi-agent system that systematically generates differential diagnoses, challenges initial hypotheses, and strategically selects high-yield diagnostic tests. ResultsHuman neurologists achieved 81% diagnostic accuracy (79% residents, 88% specialists) across 97 sessions; base LLMs ranged from 81-94%. Gregory achieved perfect diagnostic accuracy with markedly lower diagnostic costs (average $1,423; 95% CI: $450-$2,860) compared with human neurologists (average $3,041; 95% CI: $2,464-$3,677; p=0.008) and base LLMs (average $2,759; 95% CI: $2,137-$3,476; p=0.002). Time to diagnosis was also shorter with Gregory (23 days; 95% CI: 6-48) versus human neurologists (43 days; 95% CI: 31-58; p=0.002) and base models (41 days; 95% CI: 31-51; p=0.07). The platform revealed distinct diagnostic patterns: human users and some base LLMs frequently ordered broad and expensive testing, while Gregory employed targeted strategies that avoided unnecessary procedures without sacrificing thoroughness. ConclusionsA well-designed multi-agent AI system outperformed both human physicians and base LLMs in diagnostic accuracy, while reducing costs and time. DiagnosticXchange enables systematic evaluation of diagnostic efficiency and reasoning in realistic, interactive scenarios, offering a clinically relevant alternative to static benchmarks and a pathway toward more effective AI-assisted diagnosis.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 92%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 94%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 92%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 91%
Similar papers in this journal
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 91%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 90%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 89%
Similar papers in this journal
Similar papers in this journal
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 93%
- Automating Evaluation of LLM-generated Responses to Patient Questions about Rare Diseases 90%
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.