Back

Large language models for cancer registry abstraction: a real-world evaluation across models, variables, and cancer types

Fuchs, J.; Satusky, M. J.; Leese, P. J.; Nag, S.; Zipple, I. W.; Baggett, C. D.; Lash, S.; Reeder-Hayes, K.; Wood, W. A.; Johnson, C. T.; Critchley, C.; Krishnamurthy, A. K.; Elston Lafata, J.; Thompson, C. A.; Troester, M. A.; Pfaff, E. R.

2026-06-29 health informatics
10.64898/2026.06.25.26356626 medRxiv
Show abstract

Cancer registries enable cancer surveillance at the population level. These registries require significant human-time to read through many different parts of the electronic health record, including structured data and lengthy, free-text clinical reports, to abstract values for hundreds of required variables. Large language models (LLMs) offer the possibility to significantly improve this process by supporting and speeding up cancer registry data abstraction. However, it is unclear how well these models perform at real-world cancer registry abstraction involving multiple cancer types and large patient volumes. Here, we evaluate five foundational LLMs for their ability to reliably abstract cancer registry variables. We leverage hospital cancer registry data from a large regional health system as the ground truth and use LLMs to abstract from clinical reports eight registry variables for 5,939 patients with seven different cancer types. We use a zero-shot prompting strategy to compare LLM ability on commonly abstracted cancer variables with different data types. The results show that larger and more advanced models (Claude Sonnet 4.5, GPT-OSS-120b, GPT-OSS-20b) generally outperform smaller models (Gemma 12b, LLaMA 3.1 8b). The best performing models show F1 scores around 0.8 for cancer registry variables with low cardinality (grade, summary stage, laterality), with only slightly lower F1 scores for variables with high cardinality (primary site, regional nodes examined, regional nodes positive). On the more complex task of precise date extraction, all models showed decreased performance on both diagnosis and treatment dates (exact accuracy ~0.55 for the best performing models), which increased to ~0.85 for a tolerance within {+/-}30 days. These results quantify the performance of various models as well as the potential and limitations of LLMs in cancer registry abstraction tasks.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.1%
18.6%
2
JAMIA Open
42 papers in training set
Top 0.1%
12.7%
3
Journal of Biomedical Informatics
47 papers in training set
Top 0.1%
11.1%
4
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.1%
8.9%
50% of probability mass above
5
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.2%
7.9%
6
npj Digital Medicine
118 papers in training set
Top 0.9%
6.3%
7
Scientific Reports
3612 papers in training set
Top 23%
4.3%
8
JMIR Medical Informatics
18 papers in training set
Top 0.1%
4.1%
9
Journal of Medical Internet Research
87 papers in training set
Top 0.8%
3.3%
10
PLOS ONE
5266 papers in training set
Top 45%
2.1%
11
Bioinformatics
1204 papers in training set
Top 7%
1.7%
12
Patterns
78 papers in training set
Top 1%
1.7%
13
Artificial Intelligence in Medicine
17 papers in training set
Top 0.4%
1.3%
14
International Journal of Medical Informatics
26 papers in training set
Top 0.9%
1.3%
15
Communications Medicine
113 papers in training set
Top 3%
1.1%
16
BMJ Health & Care Informatics
15 papers in training set
Top 0.7%
1.1%
17
iScience
1154 papers in training set
Top 29%
1.0%
18
Frontiers in Digital Health
24 papers in training set
Top 1%
0.8%
19
GigaScience
212 papers in training set
Top 4%
0.8%
20
PLOS Digital Health
106 papers in training set
Top 4%
0.8%
21
Biology Methods and Protocols
61 papers in training set
Top 3%
0.6%
22
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 1%
0.6%