Chemically informed representations of amino acids enable learning beyond the canonical protein alphabet
Christiansen, J. C.; Gonzalez-Valdes Tejero, M.; Hembo, C. S.; Li, Y.; Barra, C.
Show abstract
Computational models of proteins typically represent sequences using a fixed twenty-letter alphabet describing canonical amino acids. Although this symbolic representation underlies most machine learning approaches to protein analysis, it abstracts away the chemical structure of residues and cannot naturally encode post-translational modifications (PTMs). As a result, current models struggle to incorporate chemical variation beyond the canonical amino acid alphabet. Here we introduce a chemically informed representation of peptides based on two-dimensional depictions of amino acid structures. Peptides are encoded as mosaics of residue depictions and embedded using a convolutional autoencoder, allowing machine learning models to learn physicochemical features directly from molecular structure. Because these representations capture chemical properties rather than symbolic residue identities, they enable learning across structurally related residues and support generalization to modified amino acids not explicitly observed during training. Applied to Major Histocompatibility Complex class I binding prediction, these embeddings achieve competitive performance while enabling chemically interpretable attribution of the molecular features driving predictions.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.