Back

Prediction and evaluation of Split-ORFs using Ribo-seq data

Kalk, C.; Murtagh, J.; Despic, V.; Mueller-McNicoll, M.; Schulz, M.

2026-05-26 bioinformatics
10.64898/2026.05.22.727176 bioRxiv
Show abstract

Split Open Reading frames (Split-ORFs) occur in transcripts containing at least two open reading frames, each encoding a part of the same full-length protein. These multiple open reading frames arise from alternatively spliced transcript isoforms. Split-ORFs have been described in the SR protein family of splicing factors, where the resulting protein halves play important autoregulatory roles. Here, we present the Split-ORF pipeline, a computational tool that predicts Split-ORFs from transcripts sequences and identifies regions unique to the predicted Split-ORF products. Using this pipeline, we predicted more than 14,000 Split-ORF transcripts from alternatively spliced human transcripts containing premature termination codons or retained introns. Hundreds of the Split-ORF unique regions show significant Ribo-seq coverage across diverse cell types and diseases. The candidate Split-ORF genes with significant Ribo-seq coverage are enriched for RNA-binding and RNA-processing functions and the majority of them encodes RNA-binding proteins. Together, these results suggest that Split-ORFs are more widespread than previously assumed and are expressed across diverse cellular contexts. This work paves the road for future studies of the Split-ORF candidates, the mechanisms of their biogenesis and their functions within the RNA-binding protein class.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
RNA
189 papers in training set
Top 0.1%
21.5%
2
Nucleic Acids Research
1281 papers in training set
Top 0.8%
18.2%
3
Genome Research
468 papers in training set
Top 0.8%
6.6%
4
NAR Genomics and Bioinformatics
242 papers in training set
Top 0.6%
6.1%
50% of probability mass above
5
PLOS Computational Biology
1863 papers in training set
Top 7%
5.4%
6
Genome Biology
637 papers in training set
Top 2%
5.1%
7
Nature Communications
5641 papers in training set
Top 29%
5.1%
8
BMC Genomics
406 papers in training set
Top 2%
4.0%
9
PLOS ONE
5266 papers in training set
Top 45%
2.1%
10
Scientific Reports
3612 papers in training set
Top 55%
1.7%
11
PLOS Genetics
862 papers in training set
Top 8%
1.5%
12
Bioinformatics
1204 papers in training set
Top 7%
1.4%
13
Frontiers in Genetics
230 papers in training set
Top 4%
1.3%
14
eLife
5828 papers in training set
Top 58%
1.1%
15
BMC Bioinformatics
457 papers in training set
Top 5%
1.1%
16
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.1%
1.1%
17
RNA Biology
78 papers in training set
Top 1.0%
1.1%
18
Genes
144 papers in training set
Top 3%
1.1%
19
Communications Biology
993 papers in training set
Top 26%
1.0%
20
International Journal of Molecular Sciences
494 papers in training set
Top 14%
0.9%
21
Life Science Alliance
285 papers in training set
Top 7%
0.9%
22
Bioinformatics Advances
203 papers in training set
Top 5%
0.8%
23
Cell Reports
1498 papers in training set
Top 30%
0.6%
24
Computational and Structural Biotechnology Journal
242 papers in training set
Top 8%
0.6%
25
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.6%
26
npj Genomic Medicine
36 papers in training set
Top 1%
0.6%