Back

Chasing the Bigfoot: shared molecular patterns among pMHC-I differentiate self from non self peptides

Silveira, E. d. S.; Vian, B. B.; Antonio, E. C.; Tarabini, R. F.; Vieira, G. F.

2025-09-12 immunology
10.1101/2025.09.08.674516 bioRxiv
Show abstract

The ability of T cells to discriminate self from non-self peptides is a cornerstone of adaptive immunity, yet the structural principles guiding this process remain poorly defined. Here, we present a proof-of-concept study using physicochemical descriptors derived from images of peptide-MHC (pMHC) surfaces to explore patterns of immunogenicity. We modeled 353 pMHC complexes restricted to HLA-A*02:01, comprising 106 self-derived peptides from the HLA Ligand Atlas and 247 viral epitopes from the Immune Epitope Database. Electrostatic potentials of TCR-facing interfaces were computed and transformed into quantitative features across 92 immunologically relevant regions of interest. Unsupervised hierarchical clustering uncovered subsets of pMHCs forming self-only, viral-only, and mixed clusters, revealing conserved structural fingerprints linked to tolerance or immunogenicity. Heatmap analysis identified close viral-self proximities, highlighting candidate pairs potentially involved in molecular mimicry and autoimmune triggers. Building on these findings, we trained an XGBoost classifier that achieved strong performance (F1 score 0.87; AUC 0.84) in distinguishing self from viral peptides. To our knowledge, this is the first demonstration that structural and physicochemical fingerprints of pMHC complexes are sufficient to discriminate self from non-self. These results establish a framework for computational immunology and provide hypotheses for validation, with implications for autoimmunity, vaccine design, and immunotherapy.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.