Enhancing genotype-phenotype association with optimized machine learning and biological enrichment methods
Jangale, V.; Sharma, J.; Shekhawat, R. S.; Yadav, P.
Show abstract
Genome-wide association studies (GWAS) are surging again owing to newer high-quality T2T-CHM13 and human pangenome references. Conventional GWAS methods have several limitations, including high false negatives. Non-conventional machine learning-based methods are warranted for analyzing newly sequenced, albeit complex, genomic regions. We present a robust machine learning-based framework for feature selection and association analysis, incorporating functional enrichment analysis to avoid false negatives. We benchmarked four popular single nucleotide polymorphism (SNP) feature selection methods: least absolute shrinkage and selection operator, ridge regression, elastic-net, and mutual information. Furthermore, we evaluated four association methods: linear regression, random forest, support vector regression (SVR), and XGBoost. We assessed proposed framework on diverse datasets, including subsets of publicly available PennCATH datasets as well as imputed, rare-variants, and simulated datasets. Low-density lipoprotein (LDL) cholesterol level was used as a phenotype for illustration. Our analysis revealed elastic-net combined with SVR consistently outperformed other methods across various datasets. Functional annotation of top 100 SNPs from PennCATH-real dataset revealed their expression in LDL cholesterol-related tissues. Our analysis validated three previously known genes (APOB, TRAPPC9, and EEPD1) implicated in cholesterol-regulated pathways. Also, rare-variant dataset analysis confirmed 37 known genes associated with LDL cholesterol. We identified several important genes, including APOB (familial-hypercholesterolemia), PTK2B (Alzheimers disease), and PTPN12 (myocardial ischemia/reperfusion injuries) as potential drug targets for cholesterol-related diseases. Our comprehensive analyses highlight elastic-net combined with SVR for association analysis could overcome limitations of conventional GWAS approaches. Our framework effectively detects common and rare variants associated with complex traits, enhancing the understanding of complex diseases.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Computational Strategies in Nutrigenetics: Constructing a Reference Dataset of Nutrition-Associated Genetic Polymorphisms 94%
- A Utility-Based Machine Learning-Driven Personalized Lifestyle Recommendation for Cardiovascular Disease Prevention 93%
- Integration of Mendelian randomisation and systems biology models to identify novel blood-based biomarkers for stroke 92%
Similar papers in this journal
Similar papers in this journal
- Development and validation of decision rules models to stratify coronary artery disease, diabetes, and hypertension risk in preventive care: cohort study of returning UK Biobank participants 92%
- Intersections between copper, β-arrestin-1, calcium, FBXW7, CD17, insulin resistance and atherogenicity mediate depression and anxiety due to type 2 diabetes mellitus: a nomothetic network approach 92%
- Body mass index and birth weight improve polygenic risk score for type 2 diabetes 91%
Similar papers in this journal
- Multimodal AI/ML for discovering novel biomarkers and predicting disease using multi-omics profiles of patients with cardiovascular diseases 96%
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 95%
- Can machine learning improve risk prediction of incident hypertension? An internal method comparison and external validation of the Framingham risk model using HUNT Study data 95%
Similar papers in this journal
- Machine Learning Based Refined Differential Gene Expression Analysis of Pediatric Sepsis 93%
- Characterizing sensitivity and coverage of clinical WGS as a diagnostic test for genetic disorders 93%
- Novel feature selection method via kernel tensor decomposition for improved multi-omics data analysis 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.