Back

Safeguarding open-weight genomic foundation models through weight locking

Karatzikos, A.; Vasilopoulou, A.; Chan, C.; Mouratidis, I.; Georgakopoulos-Soares, I.

2026-07-10 bioinformatics
10.64898/2026.07.07.736795 bioRxiv
Show abstract

BackgroundGenomic foundation models can dramatically accelerate biological research by learning general-purpose representations of genomic data that transfer across tasks, enabling researchers to predict variant effects, regulatory elements, and molecular function, among others. To safeguard against potential biosecurity threats and malicious misuse of open-weight models, a common strategy involves excluding human-infecting viral genomes from the models training corpora. This strategy, however, can be easily circumvented by fine-tuning models on abundantly available viral data. Weight-locking with spectral deformation has been proposed as a potential method to prevent fine-tuning of neural networks, but has not been systematically evaluated in biological AI models. MethodsWe applied spectral deformation locking to the Evo-1-8k-base genomic foundation model and evaluated a panel of attack configurations spanning naive fine-tuning, low-rank adaptation (LoRA), a simple inserted-layer bypass baseline, and a white-box singular value decomposition (SVD)-chain factorisation at chain lengths k [isin] {2, 3, 5}. Recovered virological capability was quantified on three Human Virome Understanding Evaluation (HVUE) tasks. ResultsThe lock defended against the naive attacker by either standard pipeline. Naive full fine-tuning under the strong lock drove downstream virological capability significantly below the pretrained baseline on pathogenicity and host tropism, converting the attack into a capability loss rather than a gain, while naive low-rank adaptation neither moved held-out perplexity (PPL) nor recovered downstream capability above pretrained. Thus, we conclude that by neither route does the naive attacker reach the gain achieved by fine-tuning an unlocked model. Consistent with previous results in non-biological models, an informed attacker who implements the SVD-chain construction does recover capability on pathogenicity prediction, at the cost of increased computational requirements for the fine-tuning process. Availabilityhttps://github.com/Georgakopoulos-Soares-lab/glm-locking.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
12.6%
2
Nature Machine Intelligence
70 papers in training set
Top 0.1%
10.8%
3
Briefings in Bioinformatics
354 papers in training set
Top 1%
6.6%
4
Patterns
78 papers in training set
Top 0.1%
6.6%
5
PLOS Computational Biology
1863 papers in training set
Top 6%
6.1%
6
Nature Communications
5641 papers in training set
Top 26%
6.1%
7
Bioinformatics Advances
203 papers in training set
Top 1.0%
5.4%
50% of probability mass above
8
Nature Methods
385 papers in training set
Top 2%
4.7%
9
GigaScience
212 papers in training set
Top 1%
3.4%
10
Cell Systems
201 papers in training set
Top 2%
2.7%
11
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.6%
12
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 23%
2.3%
13
BMC Bioinformatics
457 papers in training set
Top 3%
2.3%
14
Genome Biology
637 papers in training set
Top 5%
2.3%
15
Scientific Reports
3612 papers in training set
Top 51%
1.9%
16
Frontiers in Genetics
230 papers in training set
Top 3%
1.7%
17
Nature
645 papers in training set
Top 7%
1.7%
18
iScience
1154 papers in training set
Top 27%
1.1%
19
Nature Biotechnology
172 papers in training set
Top 4%
1.0%
20
Advanced Science
286 papers in training set
Top 8%
1.0%
21
BioData Mining
22 papers in training set
Top 0.6%
1.0%
22
Nucleic Acids Research
1281 papers in training set
Top 13%
0.9%
23
PLOS ONE
5266 papers in training set
Top 63%
0.8%
24
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.8%
25
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.7%
0.8%
26
Journal of Chemical Information and Modeling
238 papers in training set
Top 3%
0.6%