Back

Whole-genome variant detection in long-read sequencing data from ultra-low input patient samples

Wang, K.; Lee, H.; Aex, C. J.; Finot, L.; Zhu, K.; Chang, J. R.; Horning, A. M.; Rowell, W. J.; Li, P.; Kingan, S. B.; Snyder, M. P.; Erwin, G. S.

2025-07-27 genetic and genomic medicine
10.1101/2025.07.25.25332067 medRxiv
Show abstract

Long-read sequencing provides a more complete view of the genome than short-read sequencing, with improved detection of structural variants, tandem repeats, and small variants (single nucleotide variants and insertions and deletions) in difficult-to-map regions. One limitation of long-read sequencing has been high input DNA requirements, with several micrograms required per sample. Here, we evaluated two methods of amplification-based long-read, whole-genome sequencing: Ultra-Low Input HiFi (ULI-HiFi) sequencing and droplet multiple displacement amplification (dMDA) sequencing. When benchmarked against the Genome in a Bottle reference set (NA24385), we observed high precision and recall of single nucleotide variants (SNVs) with ULI-HiFi compared to the dMDA-amplified samples (F1 scores for SNVs of 99.82% for ULI-HiFi compared to 89.46% for dMDA). Across a catalog of >1.6 million tandem repeats (TRs), ULI-HiFi achieved 90.4% perfect concordance and 98.9% accuracy when allowing for single motif differences. ULI-HiFi also illuminated medically-important genes that were poorly mapped by short-read sequencing. We extended ULI-HiFi to analyze a normal, polyp, and adenocarcinoma sample from a patient with familial adenomatous polyposis (FAP), a hereditary form of colorectal cancer. We identified a TR that progressively expanded in length from normal to polyp to adenocarcinoma. This repeat is located in the 5' UTR of LIMD1, a reported tumor suppressor. Luciferase reporter assays revealed that increasing repeat length significantly reduced expression in colorectal cancer cell lines. We conclude that ULI-HiFi improves the characterization of genetic variants in dark regions of genomes from patient samples, enabling a better understanding of human disease.

Published in Genome Research (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.