Unveiling Assembly Errors in Immunoglobulin Loci: A Comprehensive Evaluation of Long-read Genome Assemblies Across Vertebrates
Zhu, Y.; Watson, C. T.; Safonova, Y.; Pennell, M.; Bankevich, A.
Show abstract
Long-read sequencing technologies have revolutionized genome assembly producing near-complete chromosome assemblies for numerous organisms, which are invaluable to research in many fields. However, regions with complex repetitive structure continue to represent a challenge for genome assembly algorithms, particularly in areas with high heterozygosity. Robust and comprehensive solutions for the assessment of assembly accuracy and completeness in these regions do not exist. In this study we focus on the assembly of biomedically important antibody-encoding immunoglobulin (IG) loci, which are characterized by complex duplications and repeat structures. High-quality full-length assemblies for these loci are critical for resolving haplotype-level annotations of IG genes, without which, functional and evolutionary studies of antibody immunity across vertebrates are not tractable. To address these challenges, we developed a pipeline, "CloseRead", that generates multiple assembly verification metrics for analysis and visualization. These metrics expand upon those of existing quality assessment tools and specifically target complex and highly heterozygous regions. Using CloseRead, we systematically assessed the accuracy and completeness of IG loci in publicly available assemblies of 74 vertebrate species, identifying problematic regions. We also demonstrated that inspecting assembly graphs for problematic regions can both identify the root cause of assembly errors and illuminate solutions for improving erroneous assemblies. For a subset of species, we were able to correct assembly errors through targeted reassembly. Together, our analysis demonstrated the utility of assembly assessment in improving the completeness and accuracy of IG loci across species.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Structural variation of the malaria-associated human glycophorin A-B-E region 92%
- Dual indexed design of in-Drop single-cell RNA-seq libraries improves sequencing quality and throughput 92%
- Refining the transcriptome of the human malaria parasite Plasmodium falciparum using amplification-free RNA-seq 92%
Similar papers in this journal
- TeloSearchLR: an algorithm to detect novel telomere repeat motifs using long sequencing reads 92%
- A High-quality Oxford Nanopore Assembly of the Hourglass Dolphin (Lagenorhynchus cruciger) Genome 92%
- Host adaptation and genome evolution of the broad host range fungal rust pathogen, Austropuccinia psidii 92%
Similar papers in this journal
- IMGT(R) Analysis of the Human IGH Locus: Unveiling Novel Polymorphisms and Copy Number Variations in Genome Assemblies from Diverse Ancestral Backgrounds 95%
- Genome assemblies of Indian desi cattle reveals hotspots of rearrangements and immune-related genetic diversity 93%
- Benchmarking computational methods for B-cell receptor reconstruction from single-cell RNA-seq data 93%
Similar papers in this journal
Similar papers in this journal
- Pushing the limits of HiFi assemblies reveals centromere diversity between two Arabidopsis thaliana genomes 94%
- End resection and telomere healing of DNA double-strand breaks during nematode programmed DNA elimination 94%
- Enhanced Detection and Genotyping of Disease-Associated Tandem Repeats Using HMMSTR and Targeted Long-Read Sequencing 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.