Machine learning-based proteogenomic data modeling identifies circulating plasma biomarkers for early detection of lung cancer
Johnson, M. A.; Hou, L.; Huang, B. E.; Saadatpour, A.; Doostparast Torshizi, A.
Show abstract
Identifying genetic variants associated with lung cancer (LC) risk and their impact on plasma protein levels is crucial for understanding LC predisposition. The discovery of risk biomarkers can enhance early LC screening protocols and improve prognostic interventions. In this study, we performed a genome-wide association analysis using the UK Biobank and FinnGen. We identified genetic variants associated with LC and protein levels leveraging the UK Biobank Pharma Proteomics Project. The dysregulated proteins were then analyzed in pre-symptomatic LC cases compared to healthy controls followed by training machine learning models to predict future LC diagnosis. We achieved median AUCs ranging from 0.79 to 0.88 (0-4 years before diagnosis/YBD), 0.73 to 0.83 (5-9YBD), and 0.78 to 0.84 (0-9YBD) based on 5-fold cross-validation. Conducting survival analysis using the 5-9YBD cohort, we identified eight proteins, including CALCB, PLAUR/uPAR, and CD74 whose higher levels were associated with worse overall survival. We also identified potential plasma biomarkers, including previously reported candidates such as CEACAM5, CXCL17, GDF15, and WFDC2, which have shown associations with future LC diagnosis. These proteins are enriched in various pathways, including cytokine signaling, interleukin regulation, neutrophil degranulation, and lung fibrosis. In conclusion, this study generates novel insights into our understanding of the genome-proteome dynamics in LC. Furthermore, our findings present a promising panel of non-invasive plasma biomarkers that hold potential to support early LC screening initiatives and enhance future diagnostic interventions.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genetic analyses of inflammatory polyneuropathy and chronic inflammatory demyelinating polyradiculoneuropathy identified candidate genes 91%
- Pleiotropy-guided transcriptome imputation from normal and tumor tissues identifies new candidate susceptibility genes for breast and ovarian cancer 91%
- Functional genomics implicates natural killer cells as potential key drivers in the pathogenesis of ankylosing spondylitis 90%
Similar papers in this journal
- Association between circulating inflammatory markers and adult cancer risk: a Mendelian randomization analysis 92%
- COVID-19 lung disease shares driver AT2 cytopathic features with Idiopathic pulmonary fibrosis 92%
- DAGM: a novel modelling framework to assess the risk of HER2-negative breast cancer based on germline rare coding mutations 91%
Similar papers in this journal
- Comprehensive Study of Germline Mutations and Double-Hit Events in Esophageal Squamous Cell Cancer 92%
- Detection of Somatic Copy Number Deletion of CDKN2A Gene for Clinical Practices Based on Discovery of A Base-Resolution Common Deletion Region 91%
- Integrated molecular and pharmacological characterization of patient-derived xenografts from bladder and ureteral cancers identifies new potential therapies. 91%
Similar papers in this journal
- Biological insights from plasma proteomics of non-small cell lung cancer patients treated with immunotherapy 93%
- A functional genomics approach to understand host genetic regulation of COVID-19 severity 92%
- Profiling tumor immune microenvironment of non-small cell lung cancer using multiplex immunofluorescence 91%
Similar papers in this journal
- Cross-Cancer Evaluation of Polygenic Risk Scores for 17 Cancer Types in Two Large Cohorts 93%
- Identifying therapeutic targets for cancer: 2,094 circulating proteins and risk of nine cancers 93%
- Integrative spatial omics reveals distinct tumor-promoting multicellular niches and immunosuppressive mechanisms in Black American and White American patients with TNBC 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.