Large Language Models Improve Cancer Survival Prediction Using Real-World Clinical Notes
Kiermeyer, N.; Lenfers, T.; Dada, A.; Friedrich, J.; Khattab, S.; Knop, E.; Egger, J.; Pauly, M.; Jung, A.; Montavon, G.; Siveke, J. T.; Wiesweg, M.; Kasper, S.; Neumann, U. P.; Klauschen, F.; Hartmann, S.; Schuler, M.; Keyl, P.; Kleesiek, J.; Keyl, J.
Show abstract
In medical documentation, vast amounts of unstructured text are generated that are still underutilized in current prognostic models. We investigate the potential of self-hosted large language models (LLM) to extract clinically meaningful, patient-specific information from routine clinical notes for personalized risk stratification in cancer care. We collected real-world medical notes from 2,708 non-small cell lung cancer (NSCLC) patients and 814 colon cancer patients documented before treatment at a large comprehensive cancer center. LLMs extracted key prognostic indicators, including comorbidities, metastatic sites, and qualitative descriptors of patient condition, in a zero-shot manner without prior task-specific training. Integrating these LLM-derived features into machine learning models significantly improved the prediction of overall survival compared to TNM staging alone (C-Index: NSCLC, 0.72 vs 0.64; colon cancer, 0.70 vs 0.59), and surpassed models using text embeddings. Based on the LLM-informed risk scores, patients were stratified into four distinct risk groups, enabling reclassification of 61.4% of NSCLC and 68.3% of colon cancer patients. Analysis of model drivers revealed that LLM-derived factors, such as the physical condition, substantially modulated the prognostic impact of TNM stage. These findings highlight the potential of self-hosted LLM to extract clinically meaningful information from unstructured clinical documentation and support clinical decision-making.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 94%
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 94%
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 94%
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 95%
- Self-Supervised Learning Reveals Clinically Relevant Histomorphological Patterns for Therapeutic Strategies in Colon Cancer 94%
- Integration of clinical, pathological, radiological, and transcriptomic data improves the prediction of first-line immunotherapy outcome in metastatic non-small cell lung cancer 94%
Similar papers in this journal
- Explainable, federated deep learning model predicts disease progression risk of cutaneous squamous cell carcinoma 94%
- Predicting the Tumor Microenvironment Composition and Immunotherapy Response in Non-Small Cell Lung Cancer from Digital Histopathology Images 94%
- Image-Based Consensus Molecular Subtyping in Rectal Cancer Biopsies and Response to Neoadjuvant Chemoradiotherapy 94%
Similar papers in this journal
Similar papers in this journal
- Histology-based Prediction of Therapy Response to Neoadjuvant Chemotherapy for Esophageal and Esophagogastric Junction Adenocarcinomas Using Deep Learning 94%
- Simple Linear Cancer Risk Prediction Models with Novel Features Outperform Complex Approaches 94%
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.