Interpretable Distillation Reveals that Deep-learning-based Splicing Models Suffer from Pervasive Confounders and Blind Spots
Liu, S.; Zhang, W.; Regev, O.
Show abstract
Despite their growing popularity, genomic deep-learning-based models function largely as black boxes, raising concerns about their trustworthiness. Here we develop a framework to explain model prediction logic using interpretable distillation. Applying our framework, we find that RNA splicing prediction models suffer from pervasive confounders and blind spots, leading to poor performance on non-reference sequences. Our findings illuminate fundamental limitations of training models on genomic sequences and suggest ways to overcome them.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Detection of aberrant splicing events in RNA-seq data with FRASER 98%
- Empirical prediction of variant-associated cryptic-donors with 87% sensitivity and 95% specificity 97%
- G4mer: An RNA language model for transcriptome-wide identification of G-quadruplexes and disease variants from population-scale genetic data 97%
Similar papers in this journal
- The SpliZ generalizes "Percent Spliced In" to reveal regulated splicing at single-cell resolution 96%
- Multiplexed transcriptome discovery of RNA binding protein binding sites by antibody-barcode eCLIP 95%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.