Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics
Tseng, A. M.; Shrikumar, A.; Kundaje, A.
Show abstract
Deep learning models can accurately map genomic DNA sequences to associated functional molecular readouts such as protein-DNA binding data. Base-resolution importance (i.e. "attribution") scores inferred from these models can highlight predictive sequence motifs and syntax. Unfortunately, these models are prone to overfitting and are sensitive to random initializations, often resulting in noisy and irreproducible attributions that obfuscate underlying motifs. To address these shortcomings, we propose a novel attribution prior, where the Fourier transform of input-level attribution scores are computed at training-time, and high-frequency components of the Fourier spectrum are penalized. We evaluate different model architectures with and without attribution priors trained on genome-wide binary or continuous molecular profiles. We show that our attribution prior dramatically improves models stability, interpretability, and performance on held-out data, especially when training data is severely limited. Our attribution prior also allows models to identify biologically meaningful sequence motifs more sensitively and precisely within individual regulatory elements. The prior is agnostic to the model architecture or predicted experimental assay, yet provides similar gains across all experiments. This work represents an important advancement in improving the reliability of deep learning models for deciphering the regulatory code of the genome.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NetTIME: a multitask and base-pair resolution framework for improved transcription factor binding site prediction 97%
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 96%
- Tomtom-lite: Accelerating Tomtom enables large-scale and real-time motif similarity scoring 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Deep Mendelian Randomization: Investigating the causal knowledge of genomic deep learning models 96%
- Epigenetics is all you need: A Transformer to decode chromatin structural compartments from the epigenome 95%
- Representation Learning of Genomic Sequence Motifs with Convolutional Neural Networks 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.