Accelerating RepeatClassifier Based on Spark and Greedy Algorithm with Dynamic Upper Boundary
Hu, K.; Liao, X.; Zou, Y.; Wang, J.
Show abstract
1Transposable elements (TEs) represent quantitatively important components of genome sequences (e.g. 90% of the wheat genome), and play important roles in genome organization and evolution. The promotion of unsupervised annotation of transposable elements is of great significance. Classification is an important step in TE annotation, which summarize the information about the type or mechanism for the raw repetitive sequences. RepeatClassifier is a basic homology-based classification tool which compares the TE families to both the Repeat Protein Database (DB) and libraries of RepeatMasker. Unfortunately, RepeatClassifier is inefficient and takes a few days to classify the repetitive sequences of large genomes. Hence, we proposed Spark-based RepeatClassifier (SRC) which uses Greedy Algorithm with Dynamic Upper Boundary (GDUB) for data division and load balancing, and Spark to improve the parallelism of RepeatClassifier. Experimental results show that SRC can not only ensure the same level of accuracy as that of RepeatClassifier, but also achieve 42-88 times of acceleration compared to RepeatClassifier. At the same time, SRC shows excellent parallel performance when dealing with input datasets with unbalanced length distribution. SRC is publicly available at https://github.com/BioinformaticsCSU/SRC.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 96%
- TrieDedup: A fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing 96%
- Detecting genomic deletions from high-throughput sequence data with unsupervised learning 96%
Similar papers in this journal
- Deep6mA: a deep learning framework for exploring similar patterns in DNA N6-methyladenine sites across different species 95%
- GCNCDA: A New Method for Predicting CircRNA-Disease Associations Based on Graph Convolutional Network Algorithm 94%
- Real-time resolution of short-read assembly graph using ONT long reads 94%
Similar papers in this journal
- SC-JNMF: Single-cell clustering integrating multiple quantification methods based on joint non-negative matrix factorization 92%
- NGScloud2: optimized bioinformatic analysis using Amazon Web Services 92%
- Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors 92%
Similar papers in this journal
- MetaLogo: a heterogeneity-aware sequence logo generator and aligner 95%
- LDBlockShow: a fast and convenient tool for visualizing linkage disequilibrium and haplotype blocks based on variant call format files 95%
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.