Back

Human whole epigenome modelling for clinical applications with Pleiades

Niki, P.; Nalmpantis, C.; Ganbat, J.-O.; Byrne, D.; Babikir, H.; Jhutty, A.; Rowe, W.; Liu, T.; Loyfer, N.; Toniolo, S.; Husain, M.; Manohar, S. G.; Thompson, S.; Koychev, I.; Zetterberg, H.; Sugar, R.; Malinauskas, A.; Saab, K.; Madan, H.; Wan, J. C. M.; Solanki, R.

2025-07-21 genomics
10.1101/2025.07.16.665231 bioRxiv
Show abstract

Gene regulation in humans extends beyond the four letter genetic code. Cytosine methylation, in particular, functions as a critical epigenetic switchboard, dynamically programming cellular identity, adapting gene expression in response to environmental cues, and underpinning the onset and progression of numerous diseases. Here we present Pleiades, a series of whole-genome epigenetic foundation models spanning three sizes: 90M, 600M, and 7B parameters. Pleiades is trained upon an extensive proprietary dataset of methylated and unmethylated human DNA sequences totalling 1.9T tokens. We introduce alignment embeddings and stacked hierarchical attention techniques to provide precise epigenetic modelling without the need for extended context lengths. Collectively, these advances enable Pleiades to perform a diverse range of downstream biological and clinical tasks, including nucleotide-level regulatory prediction, realistic generation of cell-free DNA fragments and fragment-level celltype-of-origin classification, within a unified and scalable computational framework. We specifically apply Pleiades to the early detection of real-world cohorts of clinical Alzheimers disease and Parkinsons disease, achieving high-accuracy. We integrate Pleiades with leading protein biomarkers, achieving state-of-the-art results, underscoring the complementary value of epigenomic and proteomic multi-modal approaches. By advancing beyond the modelling of pure DNA sequences and relying on limited genomic regions, Pleiades establishes genome-wide epigenomic modelling as a new paradigm for clinical diagnostics, synthetic biology, and precision medicine.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.