Back

When Less Is Enough: Low-Rank Structure in DNA Sequence-to-Function Models

Gilfeather, E.; Chikina, M.; Kostka, D.

2026-01-23 bioinformatics
10.64898/2026.01.21.700827 bioRxiv
Show abstract

MotivationThe rapid success of deep learning sequence-to-function (S2F) models has driven a trend toward ever larger architectures for regulatory genomics. While these models achieve meaningful predictive performance, their growing size and computational cost have made them increasingly opaque, difficult to deploy, and impractical for widespread use. As S2F models mature, an open question is how much model capacity is truly necessary to represent regulatory sequence function. ResultsWe show that, within linear layers, much of the predictive power of state-of-the-art S2F models is concentrated in low-rank structure. Using post hoc singular value decomposition, we construct low-rank approximations of the Sei model that reduce model size by up to 90% while preserving over 90% correlation with full-model predictions. Strikingly, extremely low-rank models -- including rank 1 -- can match or exceed full-model performance on several variant-effect benchmarks, indicating that dominant regulatory signals lie in a surprisingly low-dimensional subspace. Combining low-rank approximation with static model quantization enables practical CPU inference, yielding over 5x speedups relative to the full Sei model and more than 100x faster inference than larger S2F models on comparable tasks. Applying the same approach to Borzoi and Enformer reveals consistent behavior, demonstrating that low-rank structure is a general property of modern S2F architectures rather than a model-specific artifact. Availability and ImplementationLow-rank Sei is available at https://github.com/kostkalab/seillra. Contactmchikina@pitt.edu, kostka@pitt.edu

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.