Hidden State Genomics: Graph-Based Analysis of Sparse Auto-Encoder Feature Activity in Genomic Language Models
Kmiec, E.; O'Brien, S.; McCoy, M.
Show abstract
Pre-trained genomic language model (gLM) representations have been anticipated to enable enhanced deep learning predictions on several genomics tasks, but current benchmarking has led to questions over what they actually encode. We studied this with mechanistic interpretability on InstaDeeps Nucleotide Transformer v2 (500M), training sparse autoencoders across all 24 encoder layers to probe latent features. Correlation-based annotation against reference regulatory tracks was inconsistent across layers and insufficient for causal interpretation. We therefore built typed sequence-to-feature knowledge graphs to explore the SAE feature space and compared cisplatin-binding versus non-binding genomic DNA sequence communities by PageRank centrality, validating candidate features with decoder-based interventions and a CNN binding classifier. Interventions showed asymmetric effects: suppressive features could collapse predictive signal, while binding-associated features shifted predictions cumulatively with the presence of other binding-associated signals. Dependency maps further indicated strong local feature sensitivity within sequences. Together, these results provide evidence that gLM representations encode highly granular sequence syntax and conservation patterns, aligning more strongly with tightly coupled molecular interactions and local biophysical constraints than with complex, distributed regulatory logic. Within the scope of our intervention setting, this pattern is consistent with stronger performance on selected molecular tasks and weaker performance on broader regulatory inference, motivating scalable methods for causal feature annotation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 96%
- Predicting Mean Ribosome Load for 5'UTR of any length using Deep Learning 95%
- CoVar: A generalizable machine learning approach to identify the coordinated regulators driving variational gene expression 94%
Similar papers in this journal
Similar papers in this journal
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 97%
- Graph Convolutional Networks for Epigenetic State Prediction Using Both Sequence and 3D Genome Data 96%
- A Novel Approach to T-Cell Receptor Beta Chain (TCRB) Repertoire Encoding Using Lossless String Compression 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.