Back

Quick and effective approximation of in silico saturation mutagenesis experiments with first-order Taylor expansion

Sasse, A.; Chikina, M.; Mostafavi, S.

2023-11-14 bioinformatics
10.1101/2023.11.10.566588 bioRxiv
Show abstract

To understand the decision process of genomic sequence-to-function models, various explainable AI algorithms have been proposed. These methods determine the importance of each nucleotide in a given input sequence to the models predictions, and enable discovery of cis regulatory motif grammar for gene regulation. The most commonly applied method is in silico saturation mutagenesis (ISM) because its per-nucleotide importance scores can be intuitively understood as the computational counterpart to in vivo saturation mutagenesis experiments. While ISM is highly interpretable, it is computationally challenging to perform, because it requires computing three forward passes for every nucleotide in the given input sequence; these computations add up when analyzing a large number of sequences, and become prohibitive as the length of the input sequences and size of the model grows. Here, we show how to use the first-order Taylor approximation to compute ISM, which reduces its computation cost to a single forward pass for an input sequence. We use our theoretical derivation to connect ISM with the gradient of the model and show how this approximation is related to a recently suggested correction of the models gradients for genomic sequence analysis. We show that the Taylor ISM (TISM) approximation is robust across different model ablations, random initializations, training parameters, and data set sizes.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.