Integration of lung tissue proteomics and genome-wide association data to identify lung cancer susceptibility proteins and potential drug targets
Xu, S.; Shi, J.; Shu, X.-O.; Tao, R.; Dou, Y.; Guo, X.; Wen, W.; Yang, Y.; Zhang, B.; Wu, J.; Deppen, S. A.; Li, B.; Zheng, W.; Long, J.; Cai, Q.
Show abstract
Background: Proteins directly impact disease development and act as drug targets. Therefore, we integrated genomic and lung tissue proteomics data to identify lung cancer susceptibility proteins, elucidating genetic mechanisms and candidate drug targets. Method: We profiled the proteome and genome in non-neoplastic lung tissue from 200 lung cancer patients. Using this data, we constructed genetic models to predict abundance across the proteome in lung tissue. We applied these models to genome-wide association study (GWAS) data from 55,174 lung cancer cases and 1,294,174 controls to evaluate their associations with the risk of lung cancer, overall and by major histological subtypes. Bayesian colocalization and Mendelian randomization (MR) analyses were used to prioritize putative causal proteins, which were cross-referenced with three main drug-protein databases to identify potential therapeutic targets. Results: We identified 29 proteins associated with lung cancer risk at a false discovery rate < 5%, including 25 for overall lung cancer, two (AQP3 and IL18) specifically for adenocarcinoma, and another two (HMGN2 and HLA-DMB) for squamous cell carcinoma. Of them, genes encoding 17 proteins reside at least 2Mb away from any known GWAS risk loci, including 14 for overall lung cancer (HYI, GPX1, GMPPB, DSP, HDDC2, MTCH2, SUOX, JMJD7, PDIA3, IL16, IQGAP1, SULT1A2, ARHGAP27, and TYMP) and three for subtypes (AQP3, IL18, and HMGN2). Among the 12 proteins located within the known risk loci, EPHX2, CLDN18, PSMD5, and CYP2S1 proteins showed an association independent of the proximal GWAS-identified lead variant. Colocalization and/or MR analysis suggested 11 potential causal proteins. Five of these candidate causal proteins (DSP, CLDN18, IQGAP1, IL18 and TYMP) are targeted by nine drugs already approved by the FDA or in phase III trials. Conclusion: Our study identified novel lung cancer susceptibility proteins and potential drug targets, offering valuable insights into lung cancer biology and future translational utilities.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Proteomic associations with forced expiratory volume – a Mendelian randomisation study 93%
- The UIP honeycomb airway cells are the site of mucin biogenesis with deranged cilia 92%
- Cigarette smoke exposed airway epithelial cell-derived EVs promote pro-inflammatory macrophage activation in alpha-1 antitrypsin deficiency 90%
Similar papers in this journal
- Co-expression in tissue-specific gene networks links genes in cancer-susceptibility loci to known somatic driver genes 92%
- Exome-wide analysis of copy number variation shows association of the human leukocyte antigen region with asthma in UK Biobank 89%
- Lung Disease Network Reveals the Impact of Comorbidity on SARS-CoV-2 infection 89%
Similar papers in this journal
- COVID-19 lung disease shares driver AT2 cytopathic features with Idiopathic pulmonary fibrosis 92%
- Heterogeneous expression of the SARS-Coronavirus-2 receptor ACE2 in the human respiratory tract 91%
- DAGM: a novel modelling framework to assess the risk of HER2-negative breast cancer based on germline rare coding mutations 90%
Similar papers in this journal
- Overlap between COPD genetic association results and transcriptional quantitative trait loci 91%
- Pleiotropy-guided transcriptome imputation from normal and tumor tissues identifies new candidate susceptibility genes for breast and ovarian cancer 91%
- Genetic analyses of inflammatory polyneuropathy and chronic inflammatory demyelinating polyradiculoneuropathy identified candidate genes 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.