Pretrained transformers applied to population cancer registries improve survival prediction in label-scarce and previously unseen cancers
Gao, Y.; Yu, S.; Xia, Y.; Chen, S.; Xia, S.; An, R.; Zeng, J.; Zhao, F.; Ma, Y.; Wang, Y.; Xie, X.; Zhang, J.
Show abstract
Prognostic models in oncology are developed one cancer at a time, from that cancer's own labelled outcomes, and fail where prognostic information is scarcest. Rare cancers account for roughly a fifth of diagnoses and most paediatric malignancies, yet seldom supply enough events for a reliable time-to-event model. We therefore asked whether a representation learned without outcome labels can supply what those cohorts cannot. A Transformer encoder was pretrained by masked field-value modelling on 9425135 tumour records from the SEER 17 registries, diagnosed in 2000 to 2023. Only diagnosis-time fields passing a fail-closed coding-verification gate were admitted, and each record was emitted as an era-specific and a harmonised view, keeping two decades of recoding auditable. The encoder was then frozen and read by a linear Cox head for overall survival. Nine rare cancers were removed from the pretraining corpus entirely, each requiring an independent pretraining run. On a sealed test partition, all nine exceeded an architecture-identical random frozen encoder in Harrell concordance by +0.0034 to +0.0368, every lower confidence limit above zero. At 256 labelled patients, all 67 cancers favoured the pretrained representation over budget-matched Cox regression, median difference +0.0283. The advantage was bounded: given the entire training set, Cox regression was favoured in seven of nine rare cancers. The encoder did not outperform a field-frequency baseline on its own objective, so upstream reconstruction did not predict downstream transfer. Outcome-agnostic registry pretraining carries prognostic signal into cancers it has never seen, and is most useful where labels are fewest, without establishing clinical utility.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep learning inference of cell type-specific gene expression from breast tumor histopathology 94%
- Explainable, federated deep learning model predicts disease progression risk of cutaneous squamous cell carcinoma 93%
- Predicting the Tumor Microenvironment Composition and Immunotherapy Response in Non-Small Cell Lung Cancer from Digital Histopathology Images 92%
Similar papers in this journal
- Integrative ensemble modelling of cetuximab sensitivity in colorectal cancer PDXs 93%
- Integration of clinical, pathological, radiological, and transcriptomic data improves the prediction of first-line immunotherapy outcome in metastatic non-small cell lung cancer 92%
- Artificial intelligence-based histopathology image analysis identifies a novel subset of endometrial cancers with distinct genomic features and unfavourable outcome 92%
Similar papers in this journal
- Simple Linear Cancer Risk Prediction Models with Novel Features Outperform Complex Approaches 94%
- Histology-based Prediction of Therapy Response to Neoadjuvant Chemotherapy for Esophageal and Esophagogastric Junction Adenocarcinomas Using Deep Learning 93%
- Error reduction in leukemia machine learning classification with conformal prediction 92%
Similar papers in this journal
Similar papers in this journal
- AI-Driven Predictive Biomarker Discovery with Contrastive Learning to Improve Clinical Trial Outcomes 93%
- Evolutionary states and trajectories characterized by distinct pathways stratify ovarian high-grade serous carcinoma patients 91%
- The Tumor Profiler Study: Integrated, multi-omic, functional tumor profiling for clinical decision support 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.