Back

A geometric atlas of how ESM3 organizes modalities across depth

Steenwyk, J. L.

2026-07-12 bioinformatics
10.64898/2026.07.08.737319 bioRxiv
Show abstract

Protein language models learn general-purpose representations from large collections of protein sequences and structures, and have advanced the prediction of protein structure and function. ESM3 is a multimodal protein language model that ingests a protein through several channels at once, including amino-acid sequence, three-dimensional structure, secondary structure (SS8), solvent accessibility (SASA), and discrete functional annotations, summing their embeddings into a single residual stream. Little is known about whether these modalities occupy separate subspaces and the depth at which they fuse. The present analysis examines ESM3 (esm3-sm-open-v1; 1.4 billion parameters; 48 transformer layers) once per modality in isolation and applies representational-similarity analysis across all 48 layers. The four physical modalities (sequence, structure, SS8, SASA) begin in distinct subspaces, remain maximally separated through roughly the first half of layers, and then fuse into a shared low-dimensional subspace between layers 25 and 35. The fusion is ordered. The structure-derived modalities (structure, SS8, SASA) are mutually aligned from the input, whereas sequence joins last, after layer 28. The functional-annotation modality never fuses; instead, it remains representationally orthogonal to the physical modalities at every layer, and this orthogonality holds whether the annotation is supplied as whole-protein or per-residue, suggesting that it is content-driven rather than a tokenization arti-fact. The fusion is a learned property, absent in a randomly initialized model of the same architecture, holds at the residue level below the mean-pool, and reorganizes variance, converting between-condition variance into within-condition variance while the stream never approaches isotropy. Fusion depth is independent of protein length but is delayed by structural disorder. The phenomenon is universal across diverse organisms. Across 5,555 proteins from 12 organisms spanning eukaryota, bacteria, and archaea, every superkingdom (and every individual organism) reaches peak modality fusion at the same network depth (layer 35).

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.1%
19.6%
2
Nature Communications
5641 papers in training set
Top 7%
19.6%
3
Nature Methods
385 papers in training set
Top 0.8%
11.2%
50% of probability mass above
4
Nature
645 papers in training set
Top 3%
5.8%
5
Bioinformatics
1204 papers in training set
Top 5%
4.3%
6
Cell Systems
201 papers in training set
Top 1%
3.4%
7
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 15%
3.4%
8
Scientific Reports
3612 papers in training set
Top 46%
2.2%
9
Patterns
78 papers in training set
Top 1%
2.0%
10
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.0%
11
Nature Biotechnology
172 papers in training set
Top 2%
2.0%
12
Genome Biology
637 papers in training set
Top 5%
1.8%
13
eLife
5828 papers in training set
Top 47%
1.8%
14
Nature Computational Science
55 papers in training set
Top 0.7%
1.5%
15
PLOS Computational Biology
1863 papers in training set
Top 16%
1.4%
16
Science
477 papers in training set
Top 6%
1.2%
17
Cell
431 papers in training set
Top 8%
1.2%
18
Communications Biology
993 papers in training set
Top 28%
0.9%
19
iScience
1154 papers in training set
Top 32%
0.9%
20
Bioinformatics Advances
203 papers in training set
Top 4%
0.9%
21
PRX Life
42 papers in training set
Top 1%
0.6%
22
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
0.6%
23
Genome Research
468 papers in training set
Top 7%
0.6%
24
Science Advances
1243 papers in training set
Top 35%
0.5%
25
PNAS Nexus
159 papers in training set
Top 5%
0.5%
26
Nucleic Acids Research
1281 papers in training set
Top 16%
0.5%
27
Nature Genetics
286 papers in training set
Top 6%
0.5%