Benchmarking, detection, and genotyping of structural variants in a population of whole-genome assemblies using the SVGAP pipeline
Hu, M.; Wan, P.; Chen, C.; Tang, S.; Chen, J.; Wang, L.; Chakraborty, M.; Zhou, Y.; Chen, J.; Gaut, B. S.; Emerson, J. J.; Yi, L.
Show abstract
Comparisons of complete genome assemblies offer a direct procedure for characterizing all genetic differences among them. However, existing tools are often limited to specific aligners or optimized for specific organisms, narrowing their applicability, particularly for large and repetitive plant genomes. Here, we introduce SVGAP, a pipeline for structural variant (SV) discovery, genotyping, and annotation from high-quality genome assemblies at the population level. Through extensive benchmarks using simulated SV datasets at individual, population, and phylogenetic contexts, we demonstrate that SVGAP performs favorably relative to existing tools in SV discovery. Additionally, SVGAP is one of the few tools to address the challenge of genotyping SVs within large assembled genome samples, and it generates fully genotyped VCF files. Applying SVGAP to 26 maize genomes revealed hidden genomic diversity in centromeres, driven by abundant insertions of centromere-specific LTR-retrotransposons. The output of SVGAP is well-suited for pan-genome construction and facilitates the interpretation of previously unexplored genomic regions.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Detection of simple and complex de novo mutations without, with, or with multiple reference sequences 95%
- Differences in activity and stability drive transposable element variation in tropical and temperate maize. 94%
- Transcriptional activity and epigenetic regulation of transposable elements in the symbiotic fungus Rhizophagus irregularis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.