Retrospective Evaluation of an AI System to Classify Negative Musculoskeletal Trauma Radiographs
Czolbe, S.; Christophersen, J.; Bachmann, R.; Mudannayake, J.; Nexmann, A.; Perria, F.; Knauer, N.; Lisouski, P.; Lundemann, M.
Show abstract
Study objectivesTo evaluate an AI system for musculoskeletal (MSK) radiography in confidently identifying examinations without injury-related pathologies to support AI-guided discharge of patients without traumatic findings. MethodsWe retrospectively sampled radiographic examinations, including one or more radiographs and a radiological report, of suspected MSK trauma from >1,000 clinical sites in two countries. Medically trained professionals independently classified all exams. When disagreements between the classification and the original radiological report arose, adjudication was performed by a third professional. Annotators were also asked to record their confidence level of each classification. The AI system analyzed all exams and assigned them to the categories AI Positive, AI Negative, and AI Negative (very high confidence). Performance of the AI system was assessed using error rate, false negative rate (FNR), and qualitative review of misclassified exams. ResultsA total of 2,962 exams were included. The AI classified 27.6% (818/2,961; 95% CI: 26.0-29.3) of exams as AI Negative (very high confidence). Of all exams, 0.7% (21/2,962; 95% CI: 0.0-1.0) were falsely classified as highconfidence negatives, corresponding to a false negative rate (FNR) of 2.0% (21/1,026; 95% CI: 0.0-2.9). Qualitative review of false negatives showed that the majority had no clinical consequence if correctly diagnosed during routine follow-up the following day, and no clearly high-risk exams were missed. ConclusionThe AI system identified over one-quarter of MSK trauma radiographs as confidently negative with a very low rate of false negatives, performing on par or better than the reported standard-of-care. These results suggest potential for safe, AI-driven decision support and workflow optimization for the discharge of patients with clearly negative examinations.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Investigating the relationship between internal spinal alignment and back shape in patients with scoliosis using PCdare: a comparative, reliability and validation study 94%
- Internal calibration for opportunistic computed tomography muscle density analysis 93%
- Early user experience and lessons learned using ultra-portable digital X-ray with computer-aided detection (DXR-CAD) products: A qualitative study from the perspective of healthcare providers 93%
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 96%
- Observer agreement and clinical significance of chest CT reporting in patients suspected of COVID-19 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 95%
- MyoVision-US: an Artificial Intelligence-Powered Software for Automated Analysis of Skeletal Muscle Ultrasonography 94%
- Inconsistency of AI in Intracranial Aneurysm Detection with Varying Dose and Image Reconstruction 92%
Similar papers in this journal
Similar papers in this journal
- Validity of intraoperative imageless navigation (Naviswiss™) for component positioning accuracy in primary total hip arthroplasty: Protocol for a prospective observational cohort study in a single-surgeon practice 92%
- Large language model-based information extraction from free-text radiology reports: a scoping review protocol 91%
- Point-of-care lung ultrasonography for early identification of mild COVID-19: a prospective cohort of outpatients in a Swiss screening center 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.