Back

TranSuite: a software suite for accurate translation and characterization of transcripts

Entizne, J. C.; Guo, W.; Calixto, C. P. G.; Spensley, M.; Tzioutziou, N. A.; Zhang, R.; Brown, J. W.

2020-12-16 bioinformatics
10.1101/2020.12.15.422989 bioRxiv
Show abstract

Protein translation programs often select the longest open reading frame (ORF) in a transcript leading to numerous inaccurate and mis-annotated ORFs in databases. Unproductive transcript isoforms containing premature termination codons (PTCs) are potential substrates for nonsense-mediated decay (NMD). These transcripts often contain truncated ORFs but are incorrectly annotated due to selection of a long ORF beginning at an AUG downstream of the PTC despite the transcript containing the authentic translation start AUG. In gene expression and alternative splicing analyses, it is important to identify transcript isoforms which code for different protein variants and to distinguish these from potential NMD substrates. Here, we present TranSuite, a pipeline of bioinformatics tools that address these challenges by performing accurate translations, characterizing alternative ORFs and identifying NMD and other features of transcripts in newly assembled and existing transcriptomes. Directly comparing ORFs defined by TranSuite and TransDecoder for the Arabidopsis transcriptome AtRTD2 identified ORF mis-calling in over 16k (27%) of transcripts by TransDecoder.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.