MHChron: diversity-balanced dataset design for robust peptide-MHC binding prediction across MHC class I and II
Chronowska, M.; Shrimpton-Phoenix, E.; Kluonis, T.
Show abstract
Accurate prediction of peptide-MHC (pMHC) binding is central to immunogenicity assessment, yet many existing predictors are trained and evaluated on narrow allele sets and restricted peptide-lengths. Here, we present MHChron, a unified pMHC binding prediction framework predicated on systematic data curation, meticulous engineering of dataset balance and diversity, and rigorous evaluation through careful splits controlling for data leakage. We assemble one of the most diverse pMHC training dataset reported to date, integrating publicly available binding data across a broad allele coverage (class I n=214, class II n=98) and peptide length range (from 8 to 36 residues). Using a focused and carefully sampled subset of this dataset, we train complementary sequence-based and structure-aware models and test them under increasingly stringent generalisation regimes. Both models achieve consistently strong performance, outperforming the evaluated state-of-the-art predictors despite being trained on numerically fewer data points. Notably, the structure-aware model did not consistently surpass the sequence-based model, except under the most demanding setting of extrapolation to unseen allele clusters, suggesting that performance gains stem primarily from dataset diversity and rigorous evaluation rather than architectural complexity. Sequence-based MHChron is released with reproducible installation and an automated whole-protein screening pipeline, enabling broad and practical use.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep learning assessment of nativeness and pairing likelihood for antibody and nanobody design with AbNatiV2 96%
- AlphaBind, a Domain-Specific Model to Predict and Optimize Antibody-Antigen Binding Affinity 94%
- Prediction of protein biophysical traits from limited data: a case study on nanobody thermostability through NanoMelt 94%
Similar papers in this journal
- BindPred: A Framework for Predicting Protein-Protein Binding Affinity from Language Model Embeddings 95%
- Deep Local Analysis deconstructs protein-protein interfaces and accurately estimates binding affinity changes upon mutation 94%
- Deep Local Analysis evaluates protein docking conformations with locally oriented cubes 93%
Similar papers in this journal
- DynamicGT: a dynamic-aware geometric transformer model to predict protein binding interfaces in flexible and disordered regions 94%
- FlowDesign: Improved Design of Antibody CDRs Through Flow Matching and Better Prior Distributions 93%
- AlphaFold2 enables accurate deorphanization of ligands to single-pass receptors 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.