Back

Vibe Coding Specificity Foundation Models

Reddy, S. T.

2026-06-04 synthetic biology
10.64898/2026.06.04.730134 bioRxiv
Show abstract

Molecular recognition -- the determination of which agent binds which target -- governs adaptive immunity, gene regulation, signal transduction, RNA silencing, enzyme catalysis, and the selectivity of therapeutics. Determining binding specificity remains dependent on experimental screening or domain-specific computational tools that do not generalize across binding modalities. Transformer softmax attention is mathematically identical to the Boltzmann distribution governing molecular binding1. This identity, together with five conditions of molecular recognition systems, prescribes a single neural network architecture for cross-modal binding prediction: dual sequence encoders, symmetric contrastive learning, and a learned physical temperature2. A Specificity Foundation Model (SFM) is an instance of this physics-derived, sequence-to-sequence architecture that maps any agent-target sequence pair to a binding compatibility score, enabling bidirectional retrieval across molecular recognition domains without requiring structural information. The first SFM for antibody-antigen binding demonstrated [~]100,000-fold greater data efficiency than comparable vision-language models3. Here we report six SFMs across six molecular recognition domains -- transcription factor-DNA, enzyme-substrate, peptide-MHC, CRISPR gRNA-off-target genomic DNA, microRNA-mRNA target, and small molecule drug-target protein -- using the identical architecture without modification and trained using publicly available data only. Evaluated by cross-modal retrieval from pools of 512 candidates (random baseline 0.2%), in-distribution R@1 ranges from 27.7% to 98.0% across the six domains. mir-SFM retrieves miRNA targets at 98.0% R@1, including the [~]80% of validated interactions that seed-matching tools cannot find. mhcSFM achieves 95.4% R@1 on held-out rare HLA alleles absent from training. Applying crisprSFM to CRISPR off-target prediction improves precision to 94.0% compared to 33.2% from Hamming distance alone. All six SFMs were built by a domain expert with no programming experience using vibe coding -- natural-language-directed AI coding agents -- with numerical claims independently verified by an orthogonal AI auditor. These results establish SFMs as a physics-derived, sequence-native class of model that augments experimental and computational workflows across molecular recognition domains.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.1%
25.9%
2
Nature Communications
5641 papers in training set
Top 23%
7.1%
3
Nature
645 papers in training set
Top 3%
5.3%
4
ACS Synthetic Biology
287 papers in training set
Top 0.7%
5.3%
5
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
3.9%
6
Bioinformatics
1204 papers in training set
Top 5%
3.9%
50% of probability mass above
7
Nature Methods
385 papers in training set
Top 2%
3.9%
8
Nature Biotechnology
172 papers in training set
Top 1%
3.1%
9
Science
477 papers in training set
Top 3%
2.6%
10
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.6%
11
Nucleic Acids Research
1281 papers in training set
Top 7%
2.3%
12
Molecular Systems Biology
162 papers in training set
Top 1%
2.3%
13
Nature Computational Science
55 papers in training set
Top 0.4%
2.1%
14
Cell Systems
201 papers in training set
Top 2%
1.9%
15
mAbs
32 papers in training set
Top 0.3%
1.7%
16
Trends in Biotechnology
12 papers in training set
Top 0.1%
1.6%
17
npj Antimicrobials and Resistance
11 papers in training set
Top 0.1%
1.6%
18
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.4%
19
Advanced Intelligent Systems
11 papers in training set
Top 0.2%
1.1%
20
PLOS Computational Biology
1863 papers in training set
Top 18%
1.1%
21
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 36%
1.1%
22
eLife
5828 papers in training set
Top 59%
1.1%
23
Science Advances
1243 papers in training set
Top 26%
1.1%
24
Scientific Reports
3612 papers in training set
Top 67%
1.1%
25
npj Systems Biology and Applications
125 papers in training set
Top 2%
1.0%
26
Neuron
337 papers in training set
Top 5%
1.0%
27
Advanced Science
286 papers in training set
Top 9%
0.9%
28
Chemical Science
73 papers in training set
Top 2%
0.8%
29
Journal of Molecular Biology
232 papers in training set
Top 5%
0.6%