CpGPT: a Foundation Model for DNA Methylation
de Lima Camillo, L. P.; Sehgal, R.; Armstrong, J.; Miller, H. E.; Higgins-Chen, A. T.; Horvath, S.; Wang, B.
Show abstract
DNA methylation is a type of epigenetic modification that plays a significant role in development, aging, and disease. Despite extensive research into the molecular mechanisms of DNA methylation, they remain poorly understood today. Foundation models are a class of machine learning model that leverage vast quantities of data to make sense of complex data types, such as genome sequences or single-cell transcriptomes. Here, we present the Cytosine-phosphate-Guanine Pretrained Transformer (CpGPT), a novel foundation model pretrained on CpGCorpus, a novel database with more than 2,000 DNA methylation datasets encompassing over 150,000 samples from diverse conditions. CpGPT leverages an improved transformer architecture to learn comprehensive representations of methylation patterns, allowing it to impute and reconstruct genome-wide methylation profiles from limited input data. By capturing sequence, positional, and epigenetic contexts, CpGPT outperforms specialized models when finetuned for agingrelated tasks, such as mortality risk and morbidity assessments. The model is highly adaptable and can impute beta values across different methylation platforms, tissue types, mammalian species, and even single-cell data. As a foundation model, CpGPT can be leveraged as a new tool for biological discovery in the field of epigenetics. The open-source code and model can be found at http://github.com/lcamillo/CpGPT. HighlightsO_LICpGPT is a novel foundation model for DNA methylation analysis, pretrained on over 2,000 datasets encompassing 150,000+ samples. C_LIO_LIThe model demonstrates strong performance in zero-shot tasks including imputation, array conversion, and reference mapping. C_LIO_LICpGPT achieves state-of-the-art results in mortality prediction and chronological age estimation. C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 97%
- MethylBERT: A Transformer-based model for read-level DNA methylation pattern identification and tumour deconvolution 96%
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.