DRB1 Subtyping Reveals Divergent Risk and Protection for Type 1 Diabetes in Middle Eastern Populations
Ahmed, I.; Hagopian, W.; Sunni, M.; Dauleh, H.; Azzam, H.; Ali, H.; Bashir, M.; Baggar, K.; Mohanadi, D.; Bhat, A.; Lakshmi, J.; Salim, S.; Chin-Smith, E.; Hussain, S.; Weedon, M. N.; Sharp, S.; Fakhro, K.; Oram, R.; Al-Shabeeb AKIL, A. S.
Show abstract
BackgroundType 1 diabetes (T1D) is strongly influenced by HLA variation, yet current genetic risk models developed largely in European cohorts perform suboptimally in Middle Eastern populations due to region-specific allele frequencies, DR4 subtype heterogeneity, and distinct haplotype structures. We aimed to characterize HLA diversity in the Qatar Biobank (QBB) cohort and develop a Middle East optimized, machine learning-based T1D risk model (MENA T1D-GRS). MethodsWe analyzed high-coverage whole-genome sequencing data from >14,000 individuals comprising 7,359 healthy controls, and 410 clinically diagnosed T1D patients plus 230 first-degree relatives (FDR). High-resolution HLA typing was performed using HLA-LA, HLA-HD, and Kourami. Haplotype phasing, LD estimation, and association testing identified population-specific risk and protective configurations. We computed GRS2 using 66 genomic variants and trained an XGBoost classifier integrating 79 weighted HLA features and GRS2 components. Synthetic data augmentation (ADASYN) was applied to correct the class imbalance between T1D cases and controls, thereby enhancing model sensitivity. Model discrimination was evaluated by AUCROC. ResultsThe QBB cohort exhibited exceptional HLA diversity, with 305 DRB1-DQA1-DQB1 haplotypes. Established risk haplotypes, DR3-DQ2.5 and DR4-DQ8.1 were significantly enriched in T1D cases, with compound heterozygosity conferring >12-fold increased odds. Importantly, DRB1*04:03 was protective (OR=0.54), contrasting sharply with DRB1*04:02 and *04:05. GRS2 achieved an AUC of 0.74 vs. population controls and 0.65 vs. FDRs; AUC improved to 0.81 in autoantibody-positive cases. The MENA T1D-GRS model achieved AUC 0.79 (baseline) and 0.82 with ADASYN. Sensitivity improved to 75-80% in autoantibody-positive subgroups. SHAP analysis revealed allele-specific effects, highlighting the opposing roles of DR4 subtypes. ConclusionThe MENA T1D-GRS provides a population-tailored genomic risk prediction tool, outperforming existing scores and capturing non-linear HLA interactions. It supports early screening, differential diagnosis, and precision medicine efforts in Middle Eastern populations.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and validation of a Trans-Ancestry polygenic risk score for Type 1 Diabetes 93%
- Young onset diabetes in Asian Indians is associated with lower measured and genetically determined beta-cell function: an INSPIRED study 92%
- Subgroups of young type 2 diabetes in India reveal insulin deficiency as a major driver 92%
Similar papers in this journal
- Genetic landscape of rare autoinflammatory disease variants in Qatar and Middle Eastern populations through the integration of whole-genome and exome datasets 94%
- Hardy-Weinberg Equilibrium in the Large Scale Genomic Sequencing Era 92%
- Admixture Mapping of Peripheral Artery Disease in a Dominican Population Reveals a Novel Risk Locus on 2q35 91%
Similar papers in this journal
- Expression levels of HLA-DRB and HLA-DQ are associated with MHC Class II haplotypes in healthy individuals and rheumatoid arthritis patients 93%
- HLA-A*11:01:01:01, HLA*C*12:02:02:01-HLA-B*52:01:02:02, age and sex are associated with severity of Japanese COVID-19 with respiratory failure 92%
- “Immunogenetics of resistance to SARS-CoV-2 infection in discordant couples” 91%
Similar papers in this journal
Similar papers in this journal
- Genome-wide Polygenic Risk Score for Type 2 Diabetes in Indian Population 91%
- Variant landscape of the RYR1 gene based on whole genome sequencing of the Singaporean population 90%
- Identification of metabolomics biomarkers for type 2 diabetes: triangulating evidence from longitudinal and Mendelian randomization analyses 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.