Back

Comprehensive benchmarking of somatic single-nucleotide variant and indel detection at ultra-low allele fractions using short- and long-read data

Ha, Y.-J. J.; Maziec, D.; Markowski, J.; Georges, S. J.; Parmalee, N. L.; Berselli, M.; Coorens, T. H.; Dong, S.; Gardiner, S.; Kalra, D.; Li, D.; Miao, B.; Musunuri, R.; Xue, L.; Yu, Z.; Walker, K.; Anderson, L.; Au, N. Y.; Cibulskis, C.; Doddapaneni, H.; Grochowski, C. M.; Jensen, D. M.; Lindsay, T.; Loy, K.; Narayan, A.; Narzisi, G.; Ou, J.; Pham, M. M.; Runnels, A. M.; Stergachis, A. B.; Sutherlin, L. M.; Wang, T.; Jin, H.; Feng, W. C.; Zhang, Y.; Veit, A. D.; Kim, C. T.; Chun, H.-J. E.; Ardlie, K.; Fulton, R. S.; Germer, S.; Gibbs, R. A.; Marth, G. T.; Bennett, J. T.; Park, P. J.

2025-10-14 bioinformatics
10.1101/2025.10.13.681545 bioRxiv
Show abstract

Mosaic mutations in normal tissues occur at low variant allele fractions (VAFs), complicating detection. To benchmark strategies, the SMaHT Network created a cell-line mixture (1:49) and produced ultra-deep whole-genome sequencing using short and long reads (five centers, 180-500x each). We assembled a reference of 44,008 mosaic SNVs and 2,059 Indels, cross-validation between platforms to expose limits of short-read analysis. We also partitioned the genome by mappability to examine the impact of genomic context, added a negative reference set, and accounted for culture-derived mutations. When seven institutions applied eleven algorithms to mixture data, call sets were largely discordant across tools and replicates, partly reflecting stochastic presence of low-VAF mutations in biological replicants. For >2% VAF SNVs, sensitivity and precision approached [~]80% at [≥]300x, with little gain from additional sequencing. This work provides a comprehensive framework for reliable detection of low-VAF mutations in non-cancer tissues and a valuable resource for the community.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.