Back

Harnessing methylation signals inherent in long-read sequencing data for improved variant phasing

Pfennig, A.; Akey, J. M.

2026-03-12 bioinformatics
10.64898/2026.03.11.710820 bioRxiv
Show abstract

Accurate phasing of genetic and epigenetic variation is crucial for many downstream analyses, including association testing, clinical variant interpretation, and inference of population history. Although long-read sequencing significantly improves the continuity and completeness of genome sequencing, reconstructing chromosome-scale haplotypes remains challenging, often requiring the integration of multiple technologies, such as PacBio HiFi and Oxford Nanopore Technologies (ONT) sequencing. While these sequencing platforms detect the epigenetic modification 5-methylcytosine (5mC), current read-based phasing algorithms do not incorporate this information. We developed a read-based phasing method named LongHap that seamlessly integrates sequence and methylation data and shows that it significantly improves haplotype reconstruction. LongHap first creates phase blocks based on overlapping heterozygous sequence variants, accurately phasing complex variants by embedding them into the broader haplotype context through belief propagation. LongHap then dynamically identifies differentially methylated sites that are informative for phasing to refine and extend initial phase blocks. Through extensive analyses, we demonstrate that LongHap outperforms existing tools, including WhatsHap, HapCUT2, LongPhase, and MethPhaser, by achieving lower switch error rates and greater phase block contiguity. Crucially, we show that LongHap also improves variant phasing in challenging, medically relevant genes. In summary, by leveraging native methylation signals from long-read sequencing data, LongHap enhances long-range haplotype reconstruction, enabling more accurate haplotype-based genome analysis. LongHap is available from: https://github.com/AkeyLab/LongHap.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.