Do Pseudosequences Matter in Neoantigen Prediction?
Valeria, A.; Karchin, R.
Show abstract
Computational prediction of neoantigens that elicit T cell responses is central to the development of personalized cancer vaccines. Many current predictors represent MHC class I alleles using selected subsets of residues, known as pseudosequences, yet the extent to which pseudosequence choice and encoding strategy influence predictive performance has not been systematically examined. This study addresses that gap by evaluating a range of MHC representations within the BigMHC EL framework. We compared pseudosequence definitions based on protein structure and evolutionary diversity, a randomly sampled pseudosequence baseline, pseudosequences of varying lengths, embeddings generated using the ESM-2 protein language model, and a graph-based annotation embedding derived from allele groupings. Models using biologically informed pseudosequences consistently outperformed the random baseline, underscoring the importance of residue selection. Protein structure and evolutionary diversity pseudosequences showed similar performance, likely reflecting overlap in residues near the peptide-binding groove. We also found that pseudosequences of approximately 30 to 35 residues produced the strongest performance. Lastly, ESM-2 and annotation-based embeddings outperformed the random baseline but did not surpass curated pseudosequences under the current setup. Together, these findings indicate that curated pseudosequences remain efficient representations of MHC alleles in neoantigen prediction models, while alternative encodings can approximate but not yet replace residue-level sequence information.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BERTrand - peptide:TCR binding prediction using Bidirectional Encoder Representations from Transformers augmented with random TCR pairing 97%
- BERTMHC: Improves MHC-peptide class II interaction prediction with transformer and multiple instance learning 96%
- EPIC-TRACE: predicting TCR binding to unseen epitopes using attention and contextualized embeddings 96%
Similar papers in this journal
- THLANet: A Deep Learning Framework for Predicting TCR-pHLA Binding in Immunotherapy Applications 95%
- Positional SHAP (PoSHAP) for Interpretation of Machine Learning Models Trained from Biological Sequences 94%
- An Integrated Approach to the Characterization of Immune Repertoires Using AIMS: An Automated Immune Molecule Separator 94%
Similar papers in this journal
- Interpretable deep learning to uncover the molecular binding patterns determining TCR-epitope interactions 95%
- Do Domain-Specific Protein Language Models Outperform General Models on Immunology-Related Tasks? 95%
- Guiding a language-model based protein design method towards MHC Class-I immune-visibility profiles for vaccines and therapeutics 95%
Similar papers in this journal
Similar papers in this journal
- Current challenges for epitope-agnostic TCR interaction prediction and a new perspective derived from image classification 96%
- DeepImmuno: Deep learning-empowered prediction and generation of immunogenic peptides for T cell immunity 95%
- TEINet: a deep learning framework for prediction of TCR-epitope binding specificity 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.