Extracting and Interpreting the Effects of Higher Order Sequence Features on Peptide MHC Binding
Dai, Z.; Huisman, B. D.; Birnbaum, M. E.; Gifford, D. K.
Show abstract
Understanding the factors contributing to peptide MHC (pMHC) affinity is critical for the study of immune responses and the development of novel therapeutics. Developments in yeast display platforms have enabled the collection of pMHC binding data for vast libraries of peptides. However, methods for interpreting this data are still at an early stage. In this work we propose an approach for extracting peptide sequence features that affect pMHC binding from such datasets. In the process we develop the theoretical framework for fitting and interpreting these features. We demonstrate that these features accurately capture the kinetics underlying pMHC binding, and can be used to predict pMHC binding well enough to rival the current state of the art. We then analyze the extracted factors and show that they correlate with our current structural understanding of MHC molecules. Finally, we discuss the implication these factors have on the complexity of peptide engineering.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable Pairwise Distillations for Generative Protein Sequence Models 96%
- Predicting Affinity Through Homology (PATH): Interpretable Binding Affinity Prediction with Persistent Homology 96%
- A linear discriminant analysis model of imbalanced associative learning in the mushroom body compartment 95%
Similar papers in this journal
Similar papers in this journal
- PALLAS: Penalized mAximum LikeLihood and pArticle Swarms for inference of gene regulatory networks from time series data 94%
- Improved DNA-versus-Protein Homology Search for Protein Fossils 94%
- Small-sample estimation of the mutational support and the distribution of mutations in the SARS-Cov-2 genome 93%
Similar papers in this journal
- Gene prioritization based on random walks with restarts and absorbing states, to define gene sets regulating drug pharmacodynamics from single-cell analyses 95%
- Theoretical properties of nearest-neighbor distance distributions and novel metrics for high dimensional bioinformatics data 95%
- Time Series Experimental Design Under One-Shot Sampling: The Importance of Condition Diversity 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.