Training with synthetic data provides accurate and openly-available DNA methylation classifiers for developmental disorders and congenital anomalies via MethaDory
Ferraro, F.; Drost, M.; van der Linde, H.; Bardina, L.; Smits, D.; de Graaf, B. M.; Schot, R.; van Bever, Y.; Brooks, A.; Donker Kaat, L.; Brosens, E.; Deng, R.; Barakat, T. S.; Verhoeven, V. J. M.; van Ham, T. J.; Kleefstra, T.; Rots, D.
Show abstract
Multiple developmental and congenital disorders due to genetic variants or environmental exposures are associated with unique genome-wide alterations in DNA methylation (DNAm). Consequently, these patterns referred to as DNAm signatures, can be leveraged for diagnostic purposes by developing artificial intelligence (AI) models that enable molecular subclassification of individuals. Notably, DNAm signature application has been particularly successful for diagnosing individuals affected by developmental disorders and congenital anomalies, especially those with defects in genes encoding the Mendelian epigenetic machinery. So far, over 100 DNAm distinctive signatures have been reported for these disorders. However, the translation into diagnostic practice remains challenging not only because of the scarcity of samples from these (ultra)rare disorders needed to train DNAm AI models, but also due to the privacy regulations that restrict the sharing of affected individuals data, lack of methods for standardization, limited replication across different centers, and the emergence of commercial entities with competing interests. In this study, we show that synthetic cases, meaning in silico cases generated from publicly available DNAm data from unaffected individuals, and summarized data derived from anonymized study cohorts of affected individuals with certain disorders, can be used to train DNAm classifiers. We demonstrate that these DNAm classifiers trained on a large cohort of synthetic cases have an improved performance compared to previously published classifiers trained on cohorts of affected individuals only, which typically are limited in size due to the rarity of these conditions. Furthermore, they improve the classification of variants with intermediate effect and mosaic cases and do not require any private affected individual data for training. Finally, to facilitate dissemination of these models, we release 169 synthetic cases-trained DNAm classifiers for 89 disorders with MethaDory, an open-access tool for simultaneous testing of these DNAm signatures.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Loss-of-function of the Zinc Finger Homeobox 4 ( ZFHX4 ) gene underlies a neurodevelopmental disorder 95%
- Transcriptome-wide outlier approach identifies individuals with minor spliceopathies 95%
- A survey of rare epigenetic variation in 23,116 human genomes identifies disease-relevant epivariations and novel CGG expansions 94%
Similar papers in this journal
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 95%
- IGenomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes 95%
- Poison exon annotations improve the yield of clinically relevant variants in genomic diagnostic testing 94%
Similar papers in this journal
- Comprehensive reanalysis for CNVs in ES data from unsolved rare disease cases results in new diagnoses 96%
- Discordance between a deep learning model and clinical-grade variant pathogenicity classification in a rare disease cohort 94%
- Association between genes regulating neural pathways for quantitative traits of speech and language disorders 92%
Similar papers in this journal
- Diagnostic Utility of Genome-wide DNA Methylation Analysis in Genetically Unsolved Developmental and Epileptic Encephalopathies and Refinement of a CHD2 Episignature 95%
- Functional annotation of rare structural variation in the human brain 94%
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.