Decode-gLM: Tools to Interpret, Audit, and Steer GenomicLanguage Models
Maiwald, A.; Crook, O. M.; Jedryszek, P.; Draye, F.; Morris, G. M.
Show abstract
While genomic language models are enabling the de novo design of entire genomes, they remain challenging to interpret, limiting their trustworthiness. Here, we show that sparse autoencoders (SAEs) trained on Nucleotide Transformer activations decompose hidden representations into interpretable biological features without supervision. Across layers and model sizes, SAEs identified over 60 diverse functional annotations encoded in the models activations. This included viral regulatory elements such as the CMV enhancer, despite viral genomes being excluded from training data. Tracing this signal revealed contamination in reference databases, demonstrating that interpretability methods can audit training data and identify hidden data leakage. We then show that Meta-SAEs, trained on the decoder weights of another SAE, can identify conceptual hierarchies encoded in the model, including a more abstract feature related to multiple HIV annotations. We confirmed that the features identified by our SAEs were learned during pretraining through probing a randomly initialised model. Finally, we demonstrate that our SAEs allow us to steer model predictions in biologically meaningful ways, showing that we can use an antibiotic-resistance SAE-feature to steer the model toward the A1408G aminoglycoside-resistance mutation in the ribosomal gene 16S rRNA. Together, these results establish SAEs as a method for both discovery and auditing, providing a toolkit for interpretable and trustworthy genomic foundation models. Readers can explore our findings at https://interpretglm.netlify.app/.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Private information leakage from functional genomics data: Quantification with calibration experiments and reduction via data sanitization protocols 94%
- The Tolman-Eichenbaum Machine: Unifying space and relational memory through generalisation in the hippocampal formation 94%
- SPLASH: a statistical, reference-free genomic algorithm unifies biological discovery 94%
Similar papers in this journal
- Deep generative model embedding of single-cell RNA-Seq profiles on hyperspheres and hyperbolic spaces 95%
- Hi-C-LSTM: Learning representations of chromatin contacts using a recurrent neural network identifies genomic drivers of conformation 95%
- scPRINT: pre-training on 50 million cells allows robust gene network predictions 95%
Similar papers in this journal
- Representation Learning of Genomic Sequence Motifs with Convolutional Neural Networks 96%
- Improving deep models of protein-coding potential with a Fourier-transform architecture and machine translation task 95%
- Predicting drug polypharmacology from cell morphology readouts using variational autoencoder latent space arithmetic 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.