Early detection of non-small cell lung cancer using electronic health record data
Li, X.; Yuan, E.; Kuperberg, S. J.; Bonzel, C.-L.; Jeffway, M. I.; Cai, T.; Liao, K. P.; Aguiar-Ibanez, R.; Kao, Y.-H.; Santorelli, M. L.; Christiani, D. C.; Cai, T.; Duan, R.
Show abstract
RationaleSpecific patient characteristics increase the risk of cancer, necessitating personalized healthcare approaches. For high-risk individuals, tailored clinical management ensures proactive monitoring and timely interventions. Electronic Health Records (EHR) data are crucial for supporting these personalized approaches, improving cancer prevention and early diagnosis. ObjectivesWe leverage EHR data and build a prediction model for early detection of non-small cell lung cancer (NSCLC). MethodsWe utilize data from Mass General Brighams EHR and implement a three-stage ensemble learning approach. Initially, we generate risk scores using multivariate logistic regression in a self-control and case-control design to distinguish between cases and controls. Subsequently, these risk scores are integrated and calibrated using a prospective Cox model to develop the risk prediction model. ResultsWe identified 127 EHR-derived features predictive for early detection of NSCLC. The highly predictive features include smoking, relevant lab test results, and chronic lung diseases. The predictive model reached area under the ROC curve (AUC) of 0.801 (positive predictive value (PPV) 0.0173 with specificity 0.02) for predicting one-year NSCLC risk in a population aged 18 and above, and AUC of 0.757 (PPV 0.0196 with specificity 0.02) in a population aged 40 and above. ConclusionsThis study identified EHR derived features which are predictive of early NSCLC diagnosis. The developed risk prediction model exhibits superior performance for early detection of NSCLC compared to a baseline model that only relies on demographic and smoking information, demonstrating the potential of incorporating EHR derived features for personalized cancer screening recommendations and early detection.
Matching journals
The top 13 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 92%
- Clinical risk factors and blood protein biomarkers of 10-year pneumonia risk 92%
- Predicting Mechanical Ventilation Effects on Six Human Tissue Transcriptomes 92%
Similar papers in this journal
- Genome-wide investigation of gene-cancer associations for the prediction of novel therapeutic targets in oncology 93%
- Predicting bloodstream infection outcome using machine learning 92%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 92%
Similar papers in this journal
- Accurate diagnosis of high-risk pulmonary nodules using a non-invasive epigenetic biomarker test 94%
- Dysregulation of lncRNA MALAT1 Contributes to Lung Cancer in African Americans by modulating the tumor immune microenvironment 93%
- Functional signatures in non-small-cell lung cancer: a systematic review and meta-analysis of sex-based differences in transcriptomic studies 91%
Similar papers in this journal
- Design and methodological considerations for biomarker discovery and validation in the Integrative Analysis of Lung Cancer Etiology and Risk (INTEGRAL) Program 93%
- Deprivation and Segregation in Ovarian cancer survival among African American Women: a mediated analysis 89%
- Racial/Ethnic Disparities in the Observed COVID-19 Case Fatality Rate Among the U.S. Population 88%
Similar papers in this journal
- Asian Lung Cancer Absolute Risk Models for lung cancer mortality based on China Kadoorie Biobank 96%
- Stratifying Lung Adenocarcinoma Risk with Multi-ancestry Polygenic Risk Scores in East Asian Never-Smokers 93%
- Genetic analysis of lung cancer reveals novel susceptibility loci and germline impact on somatic mutation burden 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.