Back

Comparative analysis of discriminative and generative natural language processing pipelines for automated prostate magnetic resonance imaging reports

Lee, D. J.; McCoy, N.; Haroldsen, C.; Gilkey, M.; Verma, S.; Pyarajan, S.; Maxwell, K.; Nickols, N.; Rettig, M.; Silvestri, G.; Garraway, I.

2026-07-14 health informatics
10.64898/2026.07.12.26357886 medRxiv
Show abstract

Objectives: Natural language processing (NLP) can enable scalable extraction of clinically relevant information from unstructured radiology reports retrieved from electronic healthcare data warehouses, but reliance on externally hosted models may pose cost, privacy, and deployment challenges. We compared self-hosted discriminative and generative NLP pipelines for automated extraction of Prostate Imaging and Reporting Data System (PIRADS) scores from multiparametric magnetic resonance imaging (mpMRI) reports used in prostate cancer risk assessment. Materials and Methods: We identified 44,511 mpMRI reports across 68 Veterans Affairs (VA) healthcare systems. A stratified random sample of 1,973 reports was used to train, test, and evaluate multiple pipeline configurations combining Named Entity Recognition (NER) models and large language models (LLMs). Performance was assessed by accuracy of maximum PI-RADS extraction and processing speed using self-hosted implementations of spaCy NER, Transformers NER, and generative LLMs Llama 3, Qwen3, and Gemma3. Results: Across the top 10 pipeline configurations, accuracy for maximum PI-RADS extraction ranged from 89.3% to 95.5%, with processing times spanning 150 milliseconds to 70 seconds per report. Generative LLM pipelines achieved the highest accuracy (up to 95.5%) but were substantially slower (2 to 70 seconds), whereas NER based pipelines demonstrated lower accuracy (88.5%) with faster performance (50 to 150 milliseconds). Discussion: Discriminative NER pipelines achieved high accuracy while offering advantages in speed and potential scalability. Accuracy gains from LLMs were accompanied by significantly higher computational cost, potentially limiting feasibility in high-volume clinical environments. Conclusion: Discriminative methods were more efficient than generative models in annotating PIRADS from mpMRI report text, providing insights into configurations for optimal clinical deployment when volume is a limiting factor. However, generative AI offered improved accuracy with less upfront development.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.2%
18.1%
2
npj Digital Medicine
118 papers in training set
Top 0.5%
11.7%
3
Scientific Reports
3612 papers in training set
Top 11%
6.6%
4
JAMIA Open
42 papers in training set
Top 0.2%
6.6%
5
BMJ Health & Care Informatics
15 papers in training set
Top 0.1%
5.4%
6
PLOS Digital Health
106 papers in training set
Top 1%
5.4%
50% of probability mass above
7
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.5%
4.0%
8
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.2%
3.2%
9
Communications Medicine
113 papers in training set
Top 1%
2.6%
10
International Journal of Medical Informatics
26 papers in training set
Top 0.5%
2.4%
11
JMIR Medical Informatics
18 papers in training set
Top 0.4%
2.1%
12
Biology Methods and Protocols
61 papers in training set
Top 0.6%
2.1%
13
Journal of Biomedical Informatics
47 papers in training set
Top 0.7%
1.9%
14
GigaScience
212 papers in training set
Top 2%
1.9%
15
BMJ Open
601 papers in training set
Top 10%
1.7%
16
PLOS ONE
5266 papers in training set
Top 51%
1.5%
17
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 0.6%
1.4%
18
JAMA Network Open
130 papers in training set
Top 3%
1.3%
19
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.5%
1.1%
20
Diagnostics
50 papers in training set
Top 2%
1.0%
21
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 1%
1.0%
22
BMC Medical Research Methodology
47 papers in training set
Top 1%
0.8%
23
Frontiers in Digital Health
24 papers in training set
Top 1%
0.8%
24
JMIR Public Health and Surveillance
45 papers in training set
Top 2%
0.8%
25
Patterns
78 papers in training set
Top 3%
0.8%
26
Artificial Intelligence in Medicine
17 papers in training set
Top 0.9%
0.6%
27
Bioinformatics
1204 papers in training set
Top 10%
0.6%
28
Medical Physics
14 papers in training set
Top 0.6%
0.6%
29
Annals of Internal Medicine
28 papers in training set
Top 0.8%
0.6%
30
European Journal of Nuclear Medicine and Molecular Imaging
20 papers in training set
Top 0.2%
0.6%