Data Extraction from Oncology Imaging Reports by Large Language Models: A Comparative Accuracy Study
Passweg, L. P.; Schwenke, J. M.; Schoenenberger, C. M.; Locher, F.; Picker, J.; Dieterle, M.; Thiele, B.; Hasler, D.; Danelli, A.; Schmitt, A. M.; Heye, T.; Stojanov, T.; Briel, M.; Kasenda, B.
Show abstract
ImportanceManual data extraction from clinical text is resource intensive. Locally hosted large language models (LLMs) may offer a privacy-preserving solution, but their performance on non-English data remains unclear. ObjectiveTo investigate whether the classification accuracy of locally hosted LLMs is non-inferior to human accuracy when determining metastasis status and treatment response from German radiology reports. DesignIn this retrospective comparative accuracy study, five locally hosted LLMs (llama3.3:70b, mistral-small:24b, qwq:32b, qwen3:32b, and gpt-oss:120b) were compared against humans. To calculate accuracy, a ground truth was established via duplicate human extraction and adjudication of discrepancies by a senior oncologist. Both initial human extraction and LLM outputs were compared against this ground truth. SettingThe study was conducted at a tertiary referral hospital in Switzerland; data processing and analyses took place inside the hospital network. Participants400 randomly sampled radiology reports from adult cancer patients (CT, MRI, PET) generated between January 2023 and May 2025. ExposuresAutomated classification of metastasis status and treatment response by LLMs using a standardized prompt pipeline compared to manual human review. Main Outcomes and MeasuresPrimary outcomes were non-inferiority (5 percentage points [pp] margin) of LLM classification accuracy compared with human accuracy for metastasis status (presence/absence by anatomical site) and treatment response categories. Secondary outcomes included accuracy for primary tumor diagnosis, radiological absence of tumor, and extraction time per report. ResultsThe analysis included 400 reports from 317 patients (mean age 63 years, 32% women). On the test set (n=300), human accuracy for metastasis status was 98.4% (95% CI 98.0%-98.8%). All LLMs were non-inferior; gpt-oss:120b performed best (97.6% accuracy; difference:xs -0.8pp [90% CI, -1.3 to -0.3 pp]). For response to treatment, human accuracy was 86.0% (95% CI 83.2%-88.8%). All LLMs were inferior; the most accurate model, gpt-oss:120b, achieved 78.3% (difference -7.7 pp [90% CI, -11.6 to -3.8 pp]). Mean human time per report was 120 seconds vs 11-63 seconds for LLMs. Conclusion and RelevanceIn this study, LLMs were non-inferior to human accuracy for classification of metastasis status but were inferior for response to treatment assessment. gpt-oss:120b was the most accurate among tested LLMs. Study RegistrationOSF: 45PVQ Key PointsO_ST_ABSQuestionC_ST_ABSCan locally hosted large language models (LLMs) match human performance when extracting sites of metastases and response to treatment from radiology reports of cancer patients? FindingsIn this preregistered, single center study of 300 German radiology reports, all evaluated LLMs were non-inferior to humans in extracting the presence or absence of metastasis by organ site, but LLMs were inferior to humans in classification of response to treatment. MeaningLLMs can be suitable for classification of metastasis status, whereas more caution is warranted for more complex tasks where additional clinical reasoning may be required.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models to help appeal denied radiotherapy services 95%
- Histology-based Prediction of Therapy Response to Neoadjuvant Chemotherapy for Esophageal and Esophagogastric Junction Adenocarcinomas Using Deep Learning 94%
- Simple Linear Cancer Risk Prediction Models with Novel Features Outperform Complex Approaches 92%
Similar papers in this journal
- Development of an artificial intelligence-generated, explainable treatment recommendation system for urothelial carcinoma and renal cell carcinoma to support multidisciplinary cancer conferences 93%
- Repurposing cardiovascular disease prediction models for cancer 89%
- Osteosarcoma: novel prognostic biomarkers using circulating and cell-free tumour DNA 89%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 92%
- Multicenter Validation of a Machine Learning Algorithm for Diagnosing Pediatric Patients with Multisystem Inflammatory Syndrome and Kawasaki Disease 90%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 90%
Similar papers in this journal
- Large-scale validation of the Prediction model Risk Of Bias ASsessment Tool (PROBAST) using a short form: high risk of bias models show poorer discrimination 91%
- Strength of Statistical Evidence for the Efficacy of Cancer Drugs: A Bayesian Re-Analysis of Trials Supporting FDA Approval 90%
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 89%
Similar papers in this journal
- Diagnostic Accuracy of Artificial Intelligence in Classifying HER2 Status in Breast Cancer Immunohistochemistry Slides and Implications for HER2-Low Cases: A Systematic Review and Meta-Analysis 92%
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 92%
- High-Sensitivity Pan-Cancer AI Assessment of Lymph Node Metastasis via Uncertainty Quantification 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.