Back

ConnectedReads: machine-learning optimized long-range genome analysis workflow for next-generation sequencing

Su, C.-T.; Weng, S.; Li, Y.-L.; Chang, M.-T.

2019-09-20 bioinformatics
10.1101/776807 bioRxiv
Show abstract

Current human genome sequencing assays in both clinical and research settings primarily utilize short-read sequencing and apply resequencing pipelines to detect genetic variants. However, theses mapping-based data analysis pipelines remains a considerable challenge due to an incomplete reference genome, mapping errors and high sequence divergence. To overcome this challenge, we propose an efficient and effective whole-read assembly workflow with unsupervised graph mining algorithms on an Apache Spark large-scale data processing platform called ConnectedReads. By fully utilizing short-read data information, ConnectedReads is able to generate assembled contigs and then benefit downstream pipelines to provide higher-resolution SV discovery than that provided by other methods, especially in high diversity against reference and N-gap regions of reference. Furthermore, we demonstrate a cost-effective approach by leveraging ConnectedReads to investigate all spectra of genetic changes in population-scale studies.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.