STRUMP-I: Structure-based machine learning approach to pMHC-I binding prediction using force field energy features
Voshall, A.; Chae, J.; Li, H.; Ko, J.; Park, W.-Y.; Lee, E. A.; Choi, Y.
Show abstract
The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on Major Histocompatibility Complex class I (MHC-I) molecules, collectively termed peptide-MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting neoantigens in cancers, the ability to accurately predict which peptides will bind to the diverse population of MHC alleles is critically important. Current computational methods for pMHC-I prediction fall broadly into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and energetics of pMHC binding. While sequence-based methods have been popularly used, their performance is dependent on the size and quality of training data. On the other hands, while structure-based approaches can generalize better across diverse MHC alleles, they traditionally depend on identifying a single global minimum energy conformation, an assumption that often fails due to the inherent binding promiscuity of MHC-I molecules. To address these limitations, we developed a STRUMP-I (STRUcture-based pMHC Prediction (for class I)), a novel pMHC binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine-learning features. STRUMP-I achieves performance comparable to state-of-the-art sequence-based models while significantly outperforming them on MHC alleles with limited representation in training data. Furthermore, STRUMP-I demonstrates strong synergy when integrated with sequence-based methods, notably enhancing prediction precision. The robustness and generalizability of STRUMP-I were confirmed by evaluating its predictive performance on independent, previously unseen datasets, including an experimentally validated cancer neoantigen dataset. This combined approach advances our capability to reliably identify clinically relevant neoantigen targets. The source code and trained models are available at https://github.com/yoonjoolab/STRUMP-I
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Stitchr: stitching coding TCR nucleotide sequences from V/J/CDR3 information 95%
- Simulation of adaptive immune receptors and repertoires with complex immune information to guide the development and benchmarking of AIRR machine learning 95%
- SingleCellSignalR: Inference of intercellular networks from single-cell transcriptomics 94%
Similar papers in this journal
- PANDORA v2.0: Benchmarking peptide-MHC II models and software improvements 96%
- Computationally profiling peptide:MHC recognition by T-cell receptors and T-cell receptor-mimetic antibodies 95%
- Regulatory approved monoclonal antibodies contain framework mutations predicted from human antibody repertoires 95%
Similar papers in this journal
- THLANet: A Deep Learning Framework for Predicting TCR-pHLA Binding in Immunotherapy Applications 95%
- Positional SHAP (PoSHAP) for Interpretation of Machine Learning Models Trained from Biological Sequences 95%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 94%
Similar papers in this journal
- Current challenges for epitope-agnostic TCR interaction prediction and a new perspective derived from image classification 95%
- A simple workflow to identify novel Small Linear Motif (SLiM)-mediated interactions with AlphaFold 95%
- DeepImmuno: Deep learning-empowered prediction and generation of immunogenic peptides for T cell immunity 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.