Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients
Hamdan, M.; Harati, A.; Al-Bakheet, A.; Fuetterer, I.; Alshaer, I.
Show abstract
Objective: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease. Methods: We retrospectively analyzed 147 consecutive patients. Each case included clinical documentation and MRI presented as two composite PNG images. Two resident doctors and three multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the surgical level. Analyses used Cochran's Q, McNemar tests with Holm correction, and Bayesian methods. Results: LLMs achieved higher therapy-decision accuracy (66.0%-68.0%; 97-100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for residents versus 33.3%-41.1% for LLMs. Conclusion: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Artificial Intelligence for Surgical Scene Understanding: A Systematic Review and Reporting Quality Meta-Analysis 91%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 91%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 89%
Similar papers in this journal
- Investigating the relationship between internal spinal alignment and back shape in patients with scoliosis using PCdare: a comparative, reliability and validation study 94%
- Improving the learning curve in monoportal endoscopic lumbar surgery: development and validation of a porcine training model. 91%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 91%
Similar papers in this journal
- Loon Lens 1.0 Validation: Agentic AI for Title and Abstract Screening in Systematic Literature Reviews 86%
- Statistical Decision Properties of Imprecise Trials Assessing COVID-19 Drugs 85%
- The roadmap for implementing value based healthcare in European university hospitals - consensus report and recommendations 85%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 92%
- Integrating Multidimensional Data Analytics for Precision Diagnosis of Chronic Low Back Pain 92%
- Chronic Low Back Pain Patient Satisfaction with Lumbar Steroid Injection: a Data-Driven Analysis 90%
Similar papers in this journal
- Validity of intraoperative imageless navigation (Naviswiss™) for component positioning accuracy in primary total hip arthroplasty: Protocol for a prospective observational cohort study in a single-surgeon practice 91%
- Protocol of the observational study STRATUM-OS: First step in the development and validation of the STRATUM tool based on multimodal data processing to assist surgery in patients affected by intra-axial brain tumours 90%
- Comparative Computed Tomography with Stress Manoeuvres for Diagnosing Distal Isolated Tibiofibular Syndesmotic Injury in Acute Ankle Sprain: a Protocol for an Accuracy-Test Prospective Study. 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.