Back

Leveraging Human Pangenome for Improved Somatic Variant Detection

Fu, Q.; Xin, Z.; Miao, B.; Zhang, W.; Kong, N.; Tang, Z.; Ruttenberg, A.; Albracht, D.; Belter, E. A.; Garza, J. E.; Tomlinson, C.; Mehinovic, E.; Shen, J.; Zhuo, X.; Dong, S.; Johnson, B. K.; Majewski, M. F.; Palmer, T.; Jang, H. J.; Cheng, Y.; Li, Z.; Lawson, H. A.; Lindsay, T.; Li, D.; Fulton, R.; Shen, H.; Jin, S. C.; SMaHT Network Assembly/Pangenome Working Group, ; Macias-Velasco, J. F.; Wang, T.

2026-01-04 genomics
10.64898/2026.01.04.697580 bioRxiv
Show abstract

Somatic variant detection is technically challenging due to low variant allele fractions, the confounding presence of germline variation, and reference bias. Linear references such as GRCh38 miss sample-specific variation, causing misalignments and incorrect variant calls. Although telomere-to-telomere donor-specific assemblies (DSAs) accurately represent individual genomes, their application is limited by cost and technical barriers. Alternatively, the graph-based human pangenome provides a scalable framework to improve read alignment and perform genome inference. Here, we benchmarked somatic variant detection using GRCh38, graph-based pangenomes, and pangenome-inferred DSAs with a HapMap mixture dataset and the COLO829 melanoma cell line. Pangenome-guided alignment improves read mapping and somatic variant calling accuracy. Furthermore, personalized pangenomes partially reconstruct donor-specific genomic content, improving accuracy, reducing germline contamination, and enabling detection of events in loci absent or poorly represented in GRCh38. These findings demonstrate that graph-based and personalized pangenomes are effective strategies for enhancing somatic variant detection compared with GRCh38.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.