Comparative analysis of discriminative and generative natural language processing pipelines for automated prostate magnetic resonance imaging reports
Lee, D. J.; McCoy, N.; Haroldsen, C.; Gilkey, M.; Verma, S.; Pyarajan, S.; Maxwell, K.; Nickols, N.; Rettig, M.; Silvestri, G.; Garraway, I.
Show abstract
Objectives: Natural language processing (NLP) can enable scalable extraction of clinically relevant information from unstructured radiology reports retrieved from electronic healthcare data warehouses, but reliance on externally hosted models may pose cost, privacy, and deployment challenges. We compared self-hosted discriminative and generative NLP pipelines for automated extraction of Prostate Imaging and Reporting Data System (PIRADS) scores from multiparametric magnetic resonance imaging (mpMRI) reports used in prostate cancer risk assessment. Materials and Methods: We identified 44,511 mpMRI reports across 68 Veterans Affairs (VA) healthcare systems. A stratified random sample of 1,973 reports was used to train, test, and evaluate multiple pipeline configurations combining Named Entity Recognition (NER) models and large language models (LLMs). Performance was assessed by accuracy of maximum PI-RADS extraction and processing speed using self-hosted implementations of spaCy NER, Transformers NER, and generative LLMs Llama 3, Qwen3, and Gemma3. Results: Across the top 10 pipeline configurations, accuracy for maximum PI-RADS extraction ranged from 89.3% to 95.5%, with processing times spanning 150 milliseconds to 70 seconds per report. Generative LLM pipelines achieved the highest accuracy (up to 95.5%) but were substantially slower (2 to 70 seconds), whereas NER based pipelines demonstrated lower accuracy (88.5%) with faster performance (50 to 150 milliseconds). Discussion: Discriminative NER pipelines achieved high accuracy while offering advantages in speed and potential scalability. Accuracy gains from LLMs were accompanied by significantly higher computational cost, potentially limiting feasibility in high-volume clinical environments. Conclusion: Discriminative methods were more efficient than generative models in annotating PIRADS from mpMRI report text, providing insights into configurations for optimal clinical deployment when volume is a limiting factor. However, generative AI offered improved accuracy with less upfront development.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Natural language inference for clinical registry curation 93%
- Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review 90%
- Quantification of abdominal fat from computed tomography using deep learning and its association with electronic health records in an academic biobank 90%
Similar papers in this journal
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 92%
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 91%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 91%
Similar papers in this journal
- The effect of digital-enabled multidisciplinary therapy conferences on efficiency and quality of the decision making in prostate-cancer care 91%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 88%
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 87%
Similar papers in this journal
- Large-Scale Deep Learning for Metastasis Detection in Pathology Reports 92%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 91%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 90%
Similar papers in this journal
- Implementation and prospective real-time evaluation of a generalized system for in-clinic deployment and validation of machine learning models in radiology 91%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 91%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.