Back

Personalized reference genome-based pipeline reveals comprehensive haplotype-resolved views of cancer genomes

Sakamoto, Y.; Ochi, Y.; Kogure, Y.; Kato, S.; Sato-Otsubo, A.; Sugawa, M.; Tanaka, Y.; Tsujimura, T.; Mikami, T.; Nagae, G.; Chiba, K.; Okada, A.; Ito, Y.; Suzuki, H.; Aburatani, H.; Koga, Y.; Kato, I.; Takita, J.; Mano, H.; Ogawa, S.; Kataoka, K.; Kato, M.; Shiraishi, Y.

2026-05-30 bioinformatics
10.64898/2026.05.28.728591 bioRxiv
Show abstract

Cancer genome analysis relies on standard human reference genomes but detecting somatic alterations in highly repetitive or individual-specific regions remains challenging. We developed the Personalized Reference genome-based Cancer Genome Analysis Pipeline (PRCGAP), to our knowledge, the first comprehensive pipeline integrating haplotype-resolved analyses of somatic point mutations, structural variants, copy number, and DNA methylation on personalized diploid reference genomes. We applied PRCGAP to eight tumor-normal cell line pairs and three pediatric B-cell acute lymphoblastic leukemia (B-ALL) clinical samples. PRCGAP detected most variants identified by GRCh38- and T2T-CHM13-based pipelines while uncovering variants in centromeric and telomeric regions. Building on PRCGAP outputs, we identified L1 retrotransposition source sites absent from standard references and showed that the IGH::DUX4 fusion in a B-ALL sample originated from an internal DUX4 pseudogene within an internal D4Z4 repeat unit, rather than the canonical full-length DUX4 gene. PRCGAP extends comprehensive cancer genome analysis from cell lines to clinical settings.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.