Back

Interpretable Distillation Reveals that Deep-learning-based Splicing Models Suffer from Pervasive Confounders and Blind Spots

Liu, S.; Zhang, W.; Regev, O.

2025-12-01 genomics
10.1101/2025.11.28.691252 bioRxiv
Show abstract

Despite their growing popularity, genomic deep-learning-based models function largely as black boxes, raising concerns about their trustworthiness. Here we develop a framework to explain model prediction logic using interpretable distillation. Applying our framework, we find that RNA splicing prediction models suffer from pervasive confounders and blind spots, leading to poor performance on non-reference sequences. Our findings illuminate fundamental limitations of training models on genomic sequences and suggest ways to overcome them.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.