Back

Human panepigenome represents epigenomic diversity

Dong, Z.; Macias-Velasco, J. F.; Zhuo, X.; Zhang, W.; Tomlinson, C.; Jiang, J.; Belter, E. A.; Jony, S. R.; Li, R.; Fulton, R. S.; Dong, S.; Liu, T.; Li, D.; Wang, T.

2026-08-23 genomics
10.64898/2026.08.19.745838 bioRxiv
Show abstract

The human pangenome captures genetic diversity beyond a single linear reference, yet conventional DNA methylome maps largely assume that the underlying CpG substrate is fixed. Here we present a first draft of the human panepigenome on DNA methylation, generated from 440 haplotype-resolved long-read methylomes spanning 26 globally distributed populations. This map jointly captures CpG presence, absence and methylation across diverse human haplotypes, revealing 12.8 million CpGs absent from GRCh38 and an average of 1.2 million variant-associated CpGs (var-CpGs) per methylome. Var-CpGs captured population-associated epigenomic variation beyond that represented by shared reference CpGs, and structural variants could create or remove entire CpG islands, frequently through mobile element insertions. Promoter var-CpG methylation was associated with transcript expression, with most associations persisting after adjustment for the underlying CpG-altering variant, whereas a subset showed evidence of mediation or variant-by-methylation interactions. CpG-altering variants were also enriched among molecular QTLs in HPRC2 and across GTEx tissues, and among clinically and pharmacogenomically annotated loci. Together, these results establish a haplotype-resolved framework in which human epigenomic diversity reflects not only variation in methylation level but also genetic variation in presence or absence of CpG substrates, providing a foundation for panepigenomic studies across tissues, populations and disease contexts.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.