ConnectedReads: machine-learning optimized long-range genome analysis workflow for next-generation sequencing
Su, C.-T.; Weng, S.; Li, Y.-L.; Chang, M.-T.
Show abstract
Current human genome sequencing assays in both clinical and research settings primarily utilize short-read sequencing and apply resequencing pipelines to detect genetic variants. However, theses mapping-based data analysis pipelines remains a considerable challenge due to an incomplete reference genome, mapping errors and high sequence divergence. To overcome this challenge, we propose an efficient and effective whole-read assembly workflow with unsupervised graph mining algorithms on an Apache Spark large-scale data processing platform called ConnectedReads. By fully utilizing short-read data information, ConnectedReads is able to generate assembled contigs and then benefit downstream pipelines to provide higher-resolution SV discovery than that provided by other methods, especially in high diversity against reference and N-gap regions of reference. Furthermore, we demonstrate a cost-effective approach by leveraging ConnectedReads to investigate all spectra of genetic changes in population-scale studies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- 3rd-ChimeraMiner: A pipeline for integrated analysis of whole genome amplification generated chimeric sequences using long-read sequencing 96%
- LDBlockShow: a fast and convenient tool for visualizing linkage disequilibrium and haplotype blocks based on variant call format files 96%
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 95%
Similar papers in this journal
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 98%
- Detecting genomic deletions from high-throughput sequence data with unsupervised learning 98%
- SLR-superscaffolder: a de novo scaffolding tool for synthetic long reads using a top-to-bottom scheme 97%
Similar papers in this journal
- FM3VCF: A Software Library for Accelerating the Loading of Large VCF Files in Genotype Data Analyses 94%
- NGSpop: A desktop software that supports population studies by identifying sequence variations from next-generation sequencing data 94%
- BD5: an open HDF5-based data format to represent quantitative biological dynamics data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.