Back

Improving generalizability for MHC-I binding peptide predictions through structure-based geometric deep learning

Marzella, D. F.; Crocioni, G.; Radusinovic, T.; Lepikhov, D.; Severin, H.; Bodor, D. L.; Rademaker, D. T.; Lin, C.; Georgievska, S.; Renaud, N.; Kessler, A. L.; Lopez-Tarifa, P.; Buschow, S.; Bekkers, E.; Xue, L. C.

2024-01-30 bioinformatics
10.1101/2023.12.04.569776 bioRxiv
Show abstract

The interaction between peptides and major histocompatibility complex (MHC) molecules is pivotal in autoimmunity, pathogen recognition and tumor immunity. Recent advances in cancer immunotherapies demand for more accurate computational prediction of MHC-bound peptides. We address the generalizability challenge of MHC-bound peptide predictions, revealing limitations in current sequence-based approaches. Our structure-based methods leveraging geometric deep learning (GDL) demonstrated promising improvement in generalizability across unseen MHC alleles. Further, we tackle data efficiency by introducing a self-supervised learning approach on structures (3D-SSL). Without being exposed to any binding affinity data, our 3D-SSL outperforms sequence-based methods trained on [~]90 times more datapoints. Finally, we demonstrate the resilience of structure-based GDL methods to biases in binding data on an Hepatitis B virus vaccine immunopeptidomics case study. This proof-of-concept study highlights structure-based methods potential to enhance generalizability and data efficiency, with important implications for data-intensive fields like T-cell receptor specificity predictions, paving the way for enhanced comprehension and manipulation of immune responses.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.