PISA: a versatile interpretation tool for visualizing cis-regulatory rules in genomic data
Zeitlinger, J.; McAnany, C. E.; Weilert, M.; Kamulegeya, F.; Mehta, G.; Gardner, J. M.; Kundaje, A.
Show abstract
Sequence-to-function neural networks learn cis-regulatory sequence rules driving many types of genomic data. Interpreting these models to relate the sequence rules to underlying biological processes remains challenging, especially for complex genomic readouts such as MNase-seq, which maps nucleosome occupancy but is confounded by experimental bias. We introduce pairwise influence by sequence attribution (PISA), an interpretation tool that combinatorially decodes which bases contributed to the readout at a specific genomic coordinate. PISA visualizes the effects of transcription factor motifs, detects undiscovered motifs with complex contribution patterns, and reveals experimental biases. By learning the bias for MNase-seq, PISA enables unprecedented nucleosome prediction models, allowing the de novo discovery of nucleosome-positioning motifs and their longrange chromatin effects, as well as the design of sequences with altered nucleosome configurations. These results show that PISA is a versatile tool that expands our ability to train and interpret sequence-to-function neural networks on genomics data and understand the underlying cis-regulatory code.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Connecting high-resolution 3D chromatin organization with epigenomics 97%
- Hi-C-LSTM: Learning representations of chromatin contacts using a recurrent neural network identifies genomic drivers of conformation 97%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 98%
- Deciphering the 3D genome organization across species from Hi-C data 96%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.