The impact of FASTQ and alignment read order on structural variation calling from long-read sequencing data
Lesack, K.; Wasmuth, J.
Show abstract
BackgroundStructural variation (SV) calling from DNA sequencing data has been challenging due to several factors, such as the ambiguity of short-read alignments, multiple complex SVs in the same genomic region, and the lack of "truth" datasets for benchmarking. Additionally, caller choice, parameter settings, and alignment method are known to affect SV calling. However, the impact of FASTQ read order on SV calling has not been explored for long-read data. ResultsIn this study, we used PacBio DNA sequencing data from 15 Caenorhabditis elegans isolates to evaluate the dependence of different SV callers on FASTQ read order. Comparisons of variant call format (VCF) files generated from the original and permutated FASTQ files demonstrated that the order of input data had a large impact on SV prediction, particularly for pbsv. The overall differences were lowest for Sniffles, regardless of the aligner used. The type of variant most affected by read order varied by caller. For pbsv, most differences occurred for deletions and duplications, while for Sniffles, permutating the read order had a stronger impact on insertions. For SVIM, inversions and deletions accounted for most differences. ConclusionThe results of this study highlight the dependence of SV calling on the order of reads encoded in FASTQ files, which has not been recognized in long-read approaches. These findings have implications for the replication of SV studies and the development of consistent SV calling protocols. Our study suggests that researchers should pay attention to the order of reads when analyzing long-read sequencing data for SV calling.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance analysis of conventional and AI-based variant callers using short and long reads 93%
- HapSolo: An optimization approach for removing secondary haplotigs during diploid genome assembly and scaffolding. 93%
- GenErode: a bioinformatics pipeline to investigate genome erosion in endangered and extinct species 92%
Similar papers in this journal
- Towards a better understanding of the low recall of insertion variants with short-read based variant callers 94%
- Flexible, Production-Scale, Human Whole Genome Sequencing On A Benchtop Sequencer 93%
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 93%
Similar papers in this journal
- ConsensuSV-ONT - a modern method for accurate structural variant calling 94%
- SENSV: Detecting Structural Variations with Precise Breakpoints using Low-Depth WGS Data from a Single Oxford Nanopore MinION Flowcell 93%
- GeneToCN: An Alignment-Free Method for Gene Copy Number Estimation Directly from Next-Generation Sequencing Reads 93%
Similar papers in this journal
- Read trimming has minimal effect on bacterial SNP calling accuracy 92%
- Evaluation of the accuracy of bacterial genome reconstruction with Oxford Nanopore R10.4.1 long-read-only sequencing 92%
- Analysis of the ARTIC V4 and V4.1 SARS-CoV-2 primers and their impact on the detection of Omicron BA.1 and BA.2 lineage defining mutations 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.