Back

Start right to end right: authentic open reading frame selection matters

Bagherian, M.; Harris, G.; Sathishkumar, P.; Lloyd, J. P. B.

2025-07-12 genomics
10.1101/2025.06.10.658836 bioRxiv
Show abstract

Accurate annotation of open reading frames (ORFs) is fundamental for understanding gene function and post-transcriptional regulation. A critical but often overlooked aspect of transcriptome annotation is the selection of authentic translation start sites. Many genome annotation pipelines identify the longest possible ORF in alternatively spliced transcripts, using internal methionine codons as putative start sites. However, this computational approach ignores the biological reality that ribosomes select start codons based on sequence context, not ORF length. Here, we demonstrate that this practice leads to systematic misannotation of nonsense-mediated decay (NMD) targets in the Arabidopsis thaliana Araport11 reference transcriptome. Using TranSuite software to identify authentic start codons, we reanalyzed transcriptomic data from an NMD-deficient mutant and found that correct ORF annotation more than doubles the number of identifiable NMD targets with premature termination codons followed by downstream exon junctions, from 203 to 426 transcripts. Furthermore, we show that incorrect ORF annotations can lead to erroneous protein structure predictions, potentially introducing computational artifacts into protein databases. Our findings underscore the importance of biologically informed ORF annotation for accurate assessment of post-transcriptional regulation and proteome prediction, with implications for all eukaryotic genome annotation projects.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.