Back

Sequence models conditioned on splicing factor expression predict splicing in unseen tissues

Reddy, A. J.; Sudmant, P. H.; Ioannidis, N. M.

2026-01-21 genomics
10.64898/2026.01.20.700496 bioRxiv
Show abstract

Predicting how RNA splicing varies across tissues is important for understanding the impact of genetic variation and identifying splicing-based disease mechanisms. Although many sequence-based deep learning models have been developed to predict splicing, most predict splice sites rather than full splicing events, are restricted to tissues seen during training, or do not account for trans-regulatory variation such as differences in splicing factor expression. Here, we present Splice Ninja, a sequence-based deep learning model that predicts percent spliced-in (PSI) values for individual splicing events across tissues by conditioning on the expression levels of 301 splicing factors. Trained on PSI measurements from many different human tissues and cell types, Splice Ninja is evaluated on three entirely held-out tissues. Despite not seeing these tissues during training, it makes accurate PSI predictions and can identify a substantial fraction of splicing events with high tissue-specificity. Its performance is comparable to Pangolin [1], which is trained directly on the test tissues, but falls short of TrASPr [2], a substantially larger model also trained on the test tissues. Splice Ninja demonstrates that integrating trans-regulatory context into sequence-based splicing models enables generalization to new cellular environments. This framework offers a promising direction for building robust, context-aware predictors of alternative splicing. Our code is available at https://github.com/anikethjr/splice_ninja.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.