Back

Retrospective cohort study extracting coexisting background breast-lesion features from stage I-III invasive breast cancer

Lim, R. J. Y.; Nitar, P.; Lau, K. W.; Leong, L. C. H.; Lim, G. H.; Tan, V. K. M.; Tan, B. K. T.; Tan, E. Y.; Goh, S. S. N.; Hartman, M.; Wong, F. Y.; Li, J.; Joint Breast Cancer Registry,

2026-05-22 oncology
10.64898/2026.05.19.26353633 medRxiv
Show abstract

Background Background breast features are frequently noted in pathology reports alongside invasive breast cancer but rarely factor into prognosis or treatment decisions. Their relationship to tumor characteristics and patient outcomes remains incompletely characterised. Methods We conducted a retrospective cohort study of 7,603 patients with Stage I-III invasive breast cancer (diagnosed 1991-2022, age <80 years) from the Joint Breast Cancer Registry in Singapore. Natural language processing (NLP) was applied to 9,754 free-text pathology reports to extract co-existing background breast features, with accuracy validated by dual-reviewer assessment of 200 reports. Unsupervised hierarchical clustering grouped extracted features into three categories. Associations with tumor characteristics were assessed by multinomial logistic regression, and ten-year overall survival by Cox proportional hazards models (median follow-up 9.6 years; 620 deaths). Results Here we show that NLP-based extraction of background breast features from routine pathology reports achieves an accuracy of over 90% across features. Lobular neoplasia and benign proliferative changes are associated with less aggressive tumor characteristics, whereas early neoplastic and papillary lesions are more prevalent in HER2-enriched and luminal B tumor subtypes. Benign proliferative changes are associated with better survival in age- and year-adjusted models (hazard ratio 0.91, 95% CI 0.86-0.97), but this association is attenuated after adjustment for stage and subtype. Conclusions NLP-enabled extraction of background breast features from pathology text is feasible at scale. These features reflect tumor biology but do not independently add prognostic information beyond established clinical variables.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Breast Cancer Research
36 papers in training set
Top 0.1%
14.9%
2
JNCI Cancer Spectrum
10 papers in training set
Top 0.1%
7.8%
3
PLOS ONE
5266 papers in training set
Top 25%
6.7%
4
npj Breast Cancer
23 papers in training set
Top 0.1%
6.2%
5
British Journal of Cancer
49 papers in training set
Top 0.2%
5.1%
6
Cancer Epidemiology, Biomarkers & Prevention
20 papers in training set
Top 0.1%
4.3%
7
Scientific Reports
3612 papers in training set
Top 24%
4.3%
8
Cancers
213 papers in training set
Top 1%
4.3%
50% of probability mass above
9
BMC Cancer
67 papers in training set
Top 0.5%
4.0%
10
JAMA Network Open
130 papers in training set
Top 0.9%
3.4%
11
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.2%
3.2%
12
Diagnostics
50 papers in training set
Top 0.7%
2.8%
13
Communications Medicine
113 papers in training set
Top 1%
2.6%
14
Clinical Cancer Research
64 papers in training set
Top 0.8%
2.4%
15
European Journal of Cancer
11 papers in training set
Top 0.1%
2.4%
16
Frontiers in Oncology
103 papers in training set
Top 2%
1.7%
17
The Journal of Pathology
26 papers in training set
Top 0.4%
1.7%
18
BMJ Open
601 papers in training set
Top 10%
1.7%
19
Nature Communications
5641 papers in training set
Top 46%
1.7%
20
International Journal of Cancer
49 papers in training set
Top 0.7%
1.5%
21
PeerJ
308 papers in training set
Top 6%
1.5%
22
Molecular Oncology
55 papers in training set
Top 1%
1.1%
23
npj Digital Medicine
118 papers in training set
Top 3%
1.0%
24
Journal of Clinical Pathology
15 papers in training set
Top 0.5%
0.8%
25
BMC Medical Genomics
50 papers in training set
Top 1%
0.8%
26
Journal of Personalized Medicine
28 papers in training set
Top 1%
0.8%
27
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 42%
0.8%
28
npj Precision Oncology
53 papers in training set
Top 2%
0.8%
29
PLOS Computational Biology
1863 papers in training set
Top 22%
0.6%
30
Cancer Medicine
26 papers in training set
Top 1%
0.6%