PlasmidHunter: Accurate and fast prediction of plasmid sequences using gene content profile and machine learning
Tian, R.; Imanian, B.
Show abstract
Plasmids are extrachromosomal DNA found in microorganisms. They often carry beneficial genes that help bacteria adapt to harsh conditions, but they can also carry genes that make bacteria harmful to humans. Plasmids are also important tools in genetic engineering, gene therapy, and drug production. However, it can be difficult to identify plasmid sequences from chromosomal sequences in genomic and metagenomic data. Here, we have developed a new tool called PlasmidHunter, which uses machine learning to predict plasmid sequences based on gene content profile. PlasmidHunter achieved high accuracies (up to 96.7%) and fast speeds in benchmark tests, outperforming other existing tools.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PlasForest: a homology-based random forest classifier for plasmid detection in genomic datasets 96%
- GenAPI: a tool for gene absence-presence identification in fragmented bacterial genome sequences 94%
- PtWAVE: A High-Sensitive deconvolution software of sequencing trace for the Detection of Large Indels in Genome Editing 93%
Similar papers in this journal
- MOSTPLAS: A Self-correction Multi-label Learning Model for Plasmid Host Range Prediction 95%
- Plasmid Permissiveness of Wastewater Microbiomes can be Predicted from 16S rDNA sequences by Machine Learning 94%
- 3CAC: improving the classification of phages and plasmids in metagenomic assemblies using assembly graphs 93%
Similar papers in this journal
- BacTermFinder: A Comprehensive and General Bacterial Terminator Finder using a CNN Ensemble 95%
- Estimating Assembly Base Errors Using K-mer Abundance Difference (KAD) Between Short Reads and Genome Assembled Sequences 94%
- Life at the extremes: Maximally divergent microbes with similar genomic signatures linked to extreme environments 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.