Fast-Part: Fast and Accurate Data Partitioning for Biological Sequence Analysis
Ahmed, S.; Emon, M. I.; Moumi, N. A.; ZHANG, L.
Show abstract
Developing effective machine learning models for classifications of biological sequences depends heavily on the quality of the training and test datasets split. Existing tools are either computationally expensive, unable to maintain the desired level of similarity between the training and test datasets, or unable to retain training-test ratio stratification. Here, we present Fast-Part, a fast and accurate sequence data partitioning tool that ensures strict homology separation between the training and test datasets and the best possible training: test stratification ratio, and at the same time, is computationally fast. Fast-Part demonstrates rapid and accurate partitioning performance across diverse protein sequence datasets and maintains strict partitioning compared to the existing tools. Fast-Part can handle massive datasets and maintain strict homology partitioning.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- \"GPress: a framework for querying General Feature Format (GFF) files and feature expression files in a compressed form\" 96%
- Compact and evenly distributed k-mer binning for genomic sequences 96%
- ganon: precise metagenomics classification against large and up-to-date sets of reference sequences 96%
Similar papers in this journal
- Sequence Compression Benchmark (SCB) database - a comprehensive evaluation of reference-free compressors for FASTA-formatted sequences 97%
- Smash++: an alignment-free and memory-efficient tool to find genomic rearrangements 96%
- CoCoPyE: feature engineering for learning and prediction of genome quality indices 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.