Back

Chemically informed representations of amino acids enable learning beyond the canonical protein alphabet

Christiansen, J. C.; Gonzalez-Valdes Tejero, M.; Hembo, C. S.; Li, Y.; Barra, C.

2026-03-16 bioinformatics
10.64898/2026.03.12.711352 bioRxiv
Show abstract

Computational models of proteins typically represent sequences using a fixed twenty-letter alphabet describing canonical amino acids. Although this symbolic representation underlies most machine learning approaches to protein analysis, it abstracts away the chemical structure of residues and cannot naturally encode post-translational modifications (PTMs). As a result, current models struggle to incorporate chemical variation beyond the canonical amino acid alphabet. Here we introduce a chemically informed representation of peptides based on two-dimensional depictions of amino acid structures. Peptides are encoded as mosaics of residue depictions and embedded using a convolutional autoencoder, allowing machine learning models to learn physicochemical features directly from molecular structure. Because these representations capture chemical properties rather than symbolic residue identities, they enable learning across structurally related residues and support generalization to modified amino acids not explicitly observed during training. Applied to Major Histocompatibility Complex class I binding prediction, these embeddings achieve competitive performance while enabling chemically interpretable attribution of the molecular features driving predictions.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.1%
34.4%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.6%
7.9%
3
Nature Communications
5641 papers in training set
Top 21%
7.9%
50% of probability mass above
4
Bioinformatics
1204 papers in training set
Top 3%
6.7%
5
Cell Systems
201 papers in training set
Top 1%
4.3%
6
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 15%
3.4%
7
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.2%
8
PLOS Computational Biology
1863 papers in training set
Top 12%
2.4%
9
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.7%
2.1%
10
Nature Methods
385 papers in training set
Top 4%
1.7%
11
eLife
5828 papers in training set
Top 48%
1.7%
12
Scientific Reports
3612 papers in training set
Top 58%
1.5%
13
mAbs
32 papers in training set
Top 0.3%
1.4%
14
PRX Life
42 papers in training set
Top 0.6%
1.3%
15
Advanced Science
286 papers in training set
Top 6%
1.3%
16
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
17
Communications Chemistry
48 papers in training set
Top 1.0%
1.1%
18
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
19
Patterns
78 papers in training set
Top 2%
1.0%
20
Cell Reports Methods
165 papers in training set
Top 4%
0.8%
21
Molecular Systems Biology
162 papers in training set
Top 3%
0.8%
22
PLOS ONE
5266 papers in training set
Top 64%
0.6%
23
Journal of Cheminformatics
29 papers in training set
Top 0.8%
0.6%
24
Protein Science
246 papers in training set
Top 4%
0.6%