Early Prediction of Gestational Diabetes Using Integrated Cell-free DNA Features and Omics-derived Genetic Scores
Dao, V. N.; Tran, N. T.; Vo, S. T.; Le, T. H.; Nguyen, T. H. T.; Nguyen, Q. H. V.; Ha, M. T. T.; Le, T. M.; Hoang, D. T. T.; Huynh, K. T. N.; Nguyen, N. V.; Nguyen, C. C.; Chi, T. B.; Nguyen, X. T.; Le, S. V.; Tran, V. D.; Nguyen, M. N. B.; Nguyen, T. V.; Nguyen, T. A. T.; Hoang, B. P.; Nguyen, T. V.; Nguyen, T. A. T.; Nguyen, T. T.; Duong, T. D.; Pham, C. H.; Luong, K. O. T.; Dao, C. N.; Hoang, K. V.; Huynh, T. T. T.; Nguyen, K. M.; Tran, S. T. T.; Tran, H. T.; Nguyen, S. C.; Tran, T. D.; Nguyen, L. P. T.; Pham, V. T.; Pham, K. C.; Thai, M. D.; Truong, M. H. T.; Pham, H. H.; Do, T. T. T.; Tan
Show abstract
BackgroundGestational diabetes mellitus (GDM) affects 15.6% of pregnancies globally, with Vietnam exhibiting one of the highest prevalences at 21%. Current diagnostic approaches at 24-28 weeks limit early intervention opportunities. We developed a multi-modal machine learning framework integrating cell-free DNA (cfDNA) structural features and genetic information for early GDM prediction at 10-12 weeks of gestation in Vietnamese women. MethodsWe analyzed blood samples from 1,086 pregnant women (435 GDM cases, 651 controls) collected at 9-12 weeks. Two parallel analytical pathways were employed: cfDNA profiling extracting cfDNA-specific features (fragment length, end motifs, GC content, nucleosome patterns), and whole-genome imputation generating predictions for [~]19,000 omics traits. Component scores were developed using TabPFN classifier and integrated via logistic regression into a unified master score. ResultsGenome-wide analysis identified five omics traits with significant GDM associations: HSD11B1, NEK7, COMMD10, KLRC4, and OCEL1. Component score optimization revealed distinct patterns--cfDNA scores peaked at 200 features (AUC=71.53), while genetics-based scores improved with up to 2,000 omics traits (AUC=77.21). The final master score, integrating three components (gbSC2000, gbSCBH, cfSC200), achieved AUCs of 86.82 - 87.19 across validation cohorts with 70% sensitivity and 89% specificity. Addition-deletion analysis confirmed that both cfDNA and genetic components provided essential, non-redundant contributions. ConclusionsThis multi-modal framework demonstrates superior performance compared to single-biomarker approaches, enabling risk stratification from very low (4% GDM prevalence) to very high risk (90% prevalence). At the cutoff 0.4, the model identifies 78% of future GDM cases at 10-12 weeks while maintaining an 18% false-positive rate, potentially enabling early interventions to prevent GDM development and associated complications.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Crowdsourcing assessment of maternal blood multi-omics for predicting gestational age and preterm birth 93%
- A multi-level investigation of the genetic relationship between endometriosis and ovarian cancer histotypes 92%
- Pancreatic cancer risk prediction using deep sequential modeling of longitudinal diagnostic and medication records 90%
Similar papers in this journal
Similar papers in this journal
- Validation of a Trans-Ancestry Polygenic Risk Score for Type 2 Diabetes in Diverse Populations 94%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 90%
- Finding associations in a heterogeneous setting: Statistical test for aberration enrichment 90%
Similar papers in this journal
Similar papers in this journal
- Using Genomic Context Informed Genotype Data and Within-model Ancestry Adjustment to Classify Type 2 Diabetes 91%
- Deep plasma proteomics identifies and validates an eight-protein biomarker panel that separate benign from malignant tumors in ovarian cancer 91%
- Genome-wide DNA methylation, imprinting, and gene expression in human placentas derived from Assisted Reproductive Technology 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.