CodAn: predictive models for the characterization of mRNA transcripts in Eukaryotes
Nachtigall, P. G.; Kashiwabara, A. Y.; Durham, A. M.
Show abstract
Characterization of the coding sequences (CDSs) is an essential step on transcriptome annotation. Incorrect characterization of CDSs can lead to the prediction of non-existent proteins that can eventually compromise knowledge if databases are populated with similar incorrect predictions made in different genomes. Even though some recent methods have succeeded in correctly prediction of the stop codon position in strand-specific sequences, prediction of the complete CDS is still far from a gold standard. More importantly, prediction in strand-blind sequences and in partial sequences is deficient, presenting very low accuracy. Here, we present CodAn, a new computational approach to predict CDS and UTR, that significantly pushes the boundaries of CDS prediction in strand-blind and in partial sequences, increases strand-specific full-CDS predictions and matches or surpasses gold-standard results in strand-specific stop codon predictions. CodAn is freely available for download at https://github.com/pedronachtigall/CodAn.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 96%
- Tailored machine learning models for functional RNA detection in genome-wide screens 96%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 95%
Similar papers in this journal
- Finding differentially expressed sRNA-Seq regions with srnadiff 95%
- Short k-mer Abundance Profiles Yield Robust Machine Learning Features and Accurate Classifiers for RNA Viruses 93%
- On taming the effect of transcript level intra-condition count variation during differential expression analysis: a story of dogs, foxes and wolves 93%
Similar papers in this journal
- Illuminating the dark side of the human transcriptome with TAMA Iso-Seq analysis 94%
- ENNGene: an Easy Neural Network model building tool for Genomics 94%
- Comprehensive genome-wide identification of angiosperm upstream ORFs with peptide sequences conserved in various taxonomic ranges using a novel pipeline, ESUCA 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.