Back

ExoShorkie: Predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning

Mandl, J.; Orenstein, Y.

2026-01-27 bioinformatics
10.64898/2026.01.25.701486 bioRxiv
Show abstract

MotivationPredicting the RNA-seq coverage of native and exogenous sequences is central to many endogenous and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over the challenging out-of-distribution exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. ResultsHere, we introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method in predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation beyond native genomic contexts. Availability and implementationExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.