StartLink+: Prediction of Gene Starts in Prokaryotic Genomes by an Algorithm Integrating Independent Sources of Evidence
Gemayel, K.; Lomsadze, A.; Borodovsky, M.
Show abstract
Algorithms of ab initio gene finding were shown to make sufficiently accurate predictions in prokaryotic genomes. Nonetheless, for up to 15-25% of genes per genome the gene start predictions might differ even when made by the supposedly most accurate tools. To address this discrepancy, we have introduced StartLink+, an approach combining ab initio and multiple sequence alignment based methods. StartLink+ makes predictions for a majority of genes per genome (73% on average); in tests on sets of genes with experimentally verified starts the StartLink+ accuracy was shown to be 98-99%. When StartLink+ predictions made for a large set of prokaryotic genomes were compared with the database annotations we observed that on average the gene start annotations deviated from the predictions for ~5% of genes in AT-rich genomes and for 10-15% of genes in GC-rich genomes.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GeneMark-EP and -EP+: eukaryotic gene prediction with self-training in the space of genes and proteins 95%
- BacTermFinder: A Comprehensive and General Bacterial Terminator Finder using a CNN Ensemble 95%
- iLoci: Robust evaluation of genome content and organization for provisional and mature genome assemblies 95%
Similar papers in this journal
Similar papers in this journal
- Read-SpaM: assembly-free and alignment-free comparison of bacterial genomes with low sequencing coverage 96%
- AmpliDiff: An Optimized Amplicon Sequencing Approach to Estimating Lineage Abundances in Viral Metagenomes 94%
- PoMeLo: a systematic computational approach to predicting metabolic loss in pathogen genomes 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.