MosaicSim: A Novel Mosaic Variant Simulator Reveals Diminishing Returns of Ultra-High Coverage for Mosaic Variant Detection
Stricker, E.; Jaryani, F.; Izydorczyk, M.; Poon, C.-L.; Sanio, P.; Alexander, A.; Deb, S.; Sedlazeck, F.; Rogers, J.; Atkinson, E. G.
Show abstract
Genetic mutations within select cells of a tissue, termed mosaic variants (MV), are being increasingly recognized for their role in human disease. This growing interest underscores the need for specialized tools to detect and analyze MVs. However, such detection methods still lack thorough evaluation, largely due to missing benchmarking datasets that are large, reliable, and reflective of the complexity of biological samples. To address this gap, we developed MosaicSim, a tool for simulating variants in realistic sequencing data. The TweakVar workflow is at the tools core and represents a unique simulation pipeline that layers simulated MVs onto empirical whole genome sequencing data, generating a large, realistic ground truth dataset that combines the strengths of both simulation and biological data. To demonstrate the functionality of the workflow, we simulated 1,000 mosaic single nucleotide polymorphisms using TweakVar within whole genome sequencing files of different coverages. MVs were called with Illuminas DRAGEN and compared to the ground truth. Our results show 150x-445x coverage performed comparably, with a true-positive rate between 50.4% (300x) and 54.9% (150x) and no false-positives detected. Across all samples, increasing variant allele frequency had a significant positive effect on call success. Additionally, we observed that call rates for variants in lower complexity regions improved with increasing read depth. We did not find significant effects attributable to specific mutation patterns or mean read map quality. MosaicSim fills a critical unmet need by providing representative, customizable ground truth datasets for MV benchmarking, enabling systematic evaluation and optimization of variant calling methods.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Whole-genome long-read sequencing downsampling and its effect on variant calling precision and recall 97%
- A Complete Pedigree-Based Graph Workflow for Rare Candidate Variant Analysis 96%
- HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads 95%
Similar papers in this journal
- Concerning the eXclusion in human genomics: The choice of sex chromosome representation in the human genome drastically affects number of identified variants 96%
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 95%
- WeavePop: A bioinformatics workflow to explore and analyze genomic variants of eukaryotic populations 93%
Similar papers in this journal
Similar papers in this journal
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 97%
- Towards a better understanding of the low recall of insertion variants with short-read based variant callers 93%
- Benchmarking long-read variant calling in diploid and polyploid genomes: insights from human and plants 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.