Precise detection of Acrs in prokaryotes using only six features
Dong, C.; Pu, D.-K.; Ma, C.; Wang, X.; Wen, Q.-F.; Zeng, Z.; Guo, F.-B.
Show abstract
Anti-CRISPR proteins (Acrs) can suppress the activity of CRISPR-Cas systems. Some viruses depend on Acrs to expand their genetic materials into the host genome which can promote species diversity. Therefore, the identification and determination of Acrs are of vital importance. In this work we developed a random forest tree-based tool, AcrDetector, to identify Acrs in the whole genomescale using merely six features. AcrDetector can achieve a mean accuracy of 99.65%, a mean recall of 75.84%, a mean precision of 99.24% and a mean F1 score of 85.97%; in multi-round, 5-fold cross-validation (30 different random states). To demonstrate that AcrDetector can identify real Acrs precisely at the whole genome-scale we performed a cross-species validation which resulted in 71.43% of real Acrs being ranked in the top 10. We applied AcrDetector to detect Acrs in the latest data. It can accurately identify 3 Acrs, which have previously been verified experimentally. A standalone version of AcrDetector is available at https://github.com/RiversDong/AcrDetector. Additionally, our result showed that most of the Acrs are transferred into their host genomes in a recent stage rather than early.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SweepCluster: A SNP clustering tool for detecting gene-specific sweeps in prokaryotes 95%
- SLR-superscaffolder: a de novo scaffolding tool for synthetic long reads using a top-to-bottom scheme 94%
- PtWAVE: A High-Sensitive deconvolution software of sequencing trace for the Detection of Large Indels in Genome Editing 93%
Similar papers in this journal
- CRISPR-Detector: Fast and Accurate Detection, Visualization, and Annotation of Genome-Wide Mutations Induced by Gene Editing Events 93%
- CDCP: a visualization and analyzing platform for single-cell datasets 92%
- NanoTrans: an integrated computational framework for comprehensive transcriptome analysis with Nanopore direct-RNA sequencing 92%
Similar papers in this journal
- Estimating Assembly Base Errors Using K-mer Abundance Difference (KAD) Between Short Reads and Genome Assembled Sequences 94%
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 93%
- GeneMark-EP and -EP+: eukaryotic gene prediction with self-training in the space of genes and proteins 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.