Improved Type 2 Diabetes Risk Stratification in the Qatar Biobank Cohort by Ensemble Learning Classifier Incorporating Multi-Trait, Population-Specific, Polygenic Risk Scores
Ahmed, I.; Ziab, M.; Mbarek, H.; Taheri, S.; Chagoury, O. L.; Hussain, S. A.; Lakshmi, J.; Bhat, A.; Fakhro, K. A.; Al-Shabeeb AKIL, A. S.
Show abstract
BackgroundType 2 Diabetes (T2D) is a pervasive chronic disease influenced by a complex interplay of environmental and genetic factors. To enhance T2D risk prediction, leveraging genetic information is essential, with polygenic risk scores (PRS) offering a promising tool for assessing individual genetic risk. Our study focuses on the comparison between multi-trait and single-trait PRS models and demonstrates how the incorporation of multi-trait PRS into risk prediction models can significantly augment T2D risk assessment accuracy and effectiveness. MethodsWe conducted genome-wide association studies (GWAS) on 12 distinct T2D-related traits within a cohort of 14,278 individuals, all sequenced under the Qatar Genome Programme (QGP). This in-depth genetic analysis yielded several novel genetic variants associated with T2D, which served as the foundation for constructing multiple weighted PRS models. To assess the cumulative risk from these predictors, we applied machine learning (ML) techniques, which allowed for a thorough risk assessment. ResultsOur research identified genetic variations tied to T2D risk and facilitated the construction of ML models integrating PRS predictors for an exhaustive risk evaluation. The top-performing ML model demonstrated a robust performance with an accuracy of 0.8549, AUC of 0.92, AUC-PR of 0.8522, and an F1 score of 0.757, reflecting its strong capacity to differentiate cases from controls. We are currently working on acquiring independent T2D cohorts to validate the efficacy of our final model. ConclusionOur research underscores the potential of PRS models in identifying individuals within the population who are at elevated risk of developing T2D and its associated complications. The use of multi-trait PRS and ML models for risk prediction could inform early interventions, potentially identifying T2D patients who stand to benefit most based on their individual genetic risk profile. This combined approach signifies a stride forward in the field of precision medicine, potentially enhancing T2D risk prediction, prevention, and management.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Hardy-Weinberg Equilibrium in the Large Scale Genomic Sequencing Era 93%
- Genetic landscape of rare autoinflammatory disease variants in Qatar and Middle Eastern populations through the integration of whole-genome and exome datasets 92%
- MetaPhat: Detecting and decomposing multivariate associations from univariate genome-wide association statistics 92%
Similar papers in this journal
Similar papers in this journal
- Assessment of ability of AlphaMissense to identify variants affecting susceptibility to common disease 94%
- A Tool for Translating Polygenic Scores onto the Absolute Scale Using Summary Statistics 94%
- Lifestyle Risk Score for aggregating multiple lifestyle factors: Handling missingness of individual lifestyle components in meta-analysis of gene-by-lifestyle interactions 92%
Similar papers in this journal
- Genome-wide Polygenic Risk Score for Type 2 Diabetes in Indian Population 94%
- Identification of metabolomics biomarkers for type 2 diabetes: triangulating evidence from longitudinal and Mendelian randomization analyses 92%
- Longitudinal pathway analysis using structural information with case studies in early type 1 diabetes 92%
Similar papers in this journal
- Evaluation of Bayesian Linear Regression Models for Gene Set Prioritization in Complex Diseases 94%
- Analyzing human knockouts to validate GPR151 as a therapeutic target for reduction of body mass index 92%
- Large scale sequence-based screen for recessive variants allows for identification and monitoring of rare deleterious variants in pigs 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.