Oyster: a neural network for modelling genomic sequences that enables exact position-specific k-mer contributions
Abdulnabi, H.; Westwood, J. T.
Show abstract
Genomic functions arise from nucleotide sequences and their overlapping k-mers - subsequences whose contributions depend on their composition, position and associations. Understanding these contributions requires computing a k-mer contribution function that may or may not consider k-mer associations. Neural networks that model associations yield powerful predictors but are notoriously hard to interpret; conversely, models that ignore associations deliver exact, position-specific contributions yet might underperform. We introduce Oyster, the first convolutional architecture that can be toggled between Exact (ignoring associations) and non-Exact modes. Exact Oysters yield closed-form k-mer contributions directly from their weights, without post-hoc attribution. By letting users choose between interpretability and complexity, Oyster provides a unified framework for transparent, high-performance sequence-to-function modeling. We apply Oyster to predict intensities of YY1-DNA interactions in human K562 cells from 500-nt DNA windows and intensities of eleven histone post-translational modifications. Exact and non-Exact variants achieved statistically indistinguishable performance, highlighting that k-mer associations are not necessarily important for all biological phenomena. Modelling of YY1-DNA interactions is dependent on the YY1 motif which is expected but is also dependent on several histone post-translational modifications including H3K9ac and H2AFZ.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improved data sets and evaluation methods for the automatic prediction of DNA-binding proteins 96%
- iPromoter-BnCNN: a novel branched CNN based predictor for identifying and classifying sigma promoters 95%
- A Novel Approach to T-Cell Receptor Beta Chain (TCRB) Repertoire Encoding Using Lossless String Compression 95%
Similar papers in this journal
Similar papers in this journal
- Identifying promoter sequence architectures via a chunking-based algorithm using non-negative matrix factorisation 96%
- Beam search decoder for enhancing sequence decoding speed in single-molecule peptide sequencing data 95%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 94%
Similar papers in this journal
- HiC-GNN: A Generalizable Model for 3D Chromosome Reconstruction Using Graph Convolutional Neural Networks 94%
- Thresholding Gini Variable Importance with a single trained Random Forest: An Empirical Bayes Approach 93%
- Multi-scale phase separation by explosive percolation with single chromatin loop resolution 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.