Artificial intelligence-driven virtual tumorboard enhances precision care in myelodysplasticsyndromes
Swoboda, D. M.; DeZern, A. E.; England, J. T.; Venugopal, S.; Kehoe, T.; Aubrey, B. J.; Raddi, M. G.; Consagra, A.; Wang, J.; Andreadakis, J.; Rivero, G.; Stahl, M.; Zeidan, A. M.; Haferlach, T.; Brunner, A. M.; Buckstein, R.; Santini, V.; Della Porta, M. G.; Sekeres, M. A.; Nazha, A.
Show abstract
Background: Large language models (LLMs) perform well on standardized medical exam questions, but their reliability for complex hematology decision making is uncertain. We compared four general-purpose LLMs (GPT-4o, GPT-o3, Claude Sonnet 4, and DeepSeek-V3) with a Virtual MDS Panel (VMP), a coordinated multi-agent AI system in which domain-specialized, rule-bound software agents (WHO/ICC guidelines; IPSS-R/IPSS-M; NCCN) collaborate to generate tumor-board-level recommendations. Methods: Each model generated diagnostic, prognostic, and treatment recommendations for 30 myelodysplastic syndrome cases. Nine international MDS experts from five institutions, blinded to model identity, completed 3,000 structured ratings using 5-point Likert scales for diagnosis, prognosis, and therapy and classified errors by severity. Results: General-purpose LLMs achieved modest expert ratings (overall mean scores: 3.7 for GPT-o3, 3.2 for GPT-4o, 3.1 for DeepSeek, and 3.0 for Claude) and contained major factual errors in at least 24% of responses. The VMP increased the proportion of outputs rated 4 or higher to 87% (vs. 34-66% for general-purpose models), improved mean scores to 4.3 overall (4.3 for diagnosis, 4.4 for prognosis, and 4.1 for therapy), and reduced major errors to 8%. Conclusions: In this blinded evaluation of 30 complex MDS cases, general-purpose LLMs produced clinically important errors at rates that raise safety concerns for autonomous hematology decision making. The VMP, a rule-bound, multi-agent architecture, approached expert-level accuracy supporting its potential role as an effective decision-support tool for MDS in the future.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genetic program activity delineates risk, relapse, and therapy responsiveness in Multiple Myeloma 91%
- MatchMiner: An open source platform for cancer precision medicine 91%
- Image-based Explainable Artificial Intelligence Accurately Identifies Myelodysplastic Neoplasms Beyond Conventional Signs of Dysplasia 90%
Similar papers in this journal
- International Electronic Health Record-Derived COVID-19 Clinical Course Profiles: The 4CE Consortium 90%
- CT-based Rapid Triage of COVID-19 Patients: Risk Prediction and Progression Estimation of ICU Admission, Mechanical Ventilation, and Death of Hospitalized Patients 89%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 89%
Similar papers in this journal
- Exploring the role of Large Language Models (LLMs) in hematology: a systematic review of applications, benefits, and limitations 92%
- Validation of the IMPEDE VTE Score for Prediction of Venous Thromboembolism in Multiple Myeloma: A Retrospective Cohort Study 90%
- Resistance to vincristine in cancerous B-cells by disruption of p53-dependent mitotic surveillance 88%
Similar papers in this journal
- Real World Predictors of Response and 24-month survival in high-grade TP53 -mutated Myeloid Neoplasms 92%
- High WEE1 expression is independently linked to poor survival in multiple myeloma 90%
- Gene interaction network analysis in multiple myeloma detects complex immune dysregulation associated with shorter survival 89%
Similar papers in this journal
- Monosomy 7/del(7q) Cause Sensitivity to Inhibitors of Nicotinamide Phosphoribosyltransferase in Acute Myeloid Leukemia 90%
- Molecular mechanisms promoting long-term cytopenia after BCMA CAR-T therapy in Multiple Myeloma 90%
- Serum Flt3 ligand is a biomarker of progenitor cell mass and prognosis in acute myeloid leukemia 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.