Back

Measuring peptide-MHC generalization to unseen alleles across both HLA classes

Mysore, V.

2026-06-23 bioinformatics
10.64898/2026.06.18.733075 bioRxiv
Show abstract

Reported peptide-MHC (pMHC) AUROCs of 0.85-0.95 overstate generalization to unseen alleles: because immunopeptidome data are dense on a few well-studied alleles and sparse on the rest, training and test sets come to share near-identical alleles, so the numbers partly reflect interpolation rather than extrapolation to new MHC grooves. This is a property of the data, not of any one method. We assembled an open, harmonized corpus of 5.8 million experimental measurements across both HLA classes and use it to control the leakage explicitly: alleles held out at the sequence and cluster level, peptide-disjoint splits, and provenance-matched negatives. On strictly novel alleles, generalization is in the high 0.7s rather than the 0.9s a conventional split returns. Against this benchmark we trained a predictor that spans both classes in one model and factors presentation into a peptide-only ligand-likeness term and an allele-specific term; it exceeds eight published predictors by per-allele {Delta}AUROC = +0.22 to +0.37 (p < 10-9), most on the least-studied genes. Corpus, benchmark, and model are released. Author summaryOur immune cells display protein fragments on the cell surface, held by molecules (the human leukocyte antigens, or HLAs) that vary from person to person. Predicting which fragments a given HLA displays matters for cancer vaccines, transplant matching, and the safety of engineered therapies, and many computational tools now do it well. Most available data come from a few common HLAs, so test cases tend to resemble training cases, and the published accuracy looks better than it really is for the rare HLAs that matter most in the clinic. We assembled a large, openly shared collection of experimental measurements across both major HLA classes and used it to test prediction more directly, holding out HLAs that are sequence-distant from those in training. Accuracy on these is measurable but lower than the usual figures suggest. We also built a predictor that handles both HLA classes in one model and gains most relative to existing tools on the rare HLAs where they are weakest. The data, benchmark, and model are available for the same test.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 14%
12.7%
2
Cell Systems
201 papers in training set
Top 0.2%
11.7%
3
Bioinformatics
1204 papers in training set
Top 2%
10.4%
4
eLife
5828 papers in training set
Top 22%
5.4%
5
PLOS Computational Biology
1863 papers in training set
Top 8%
4.8%
6
Nature Biotechnology
172 papers in training set
Top 1.0%
4.2%
7
mAbs
32 papers in training set
Top 0.2%
3.2%
50% of probability mass above
8
Nature Machine Intelligence
70 papers in training set
Top 0.9%
3.2%
9
Nature Methods
385 papers in training set
Top 3%
3.2%
10
Patterns
78 papers in training set
Top 0.6%
3.1%
11
Frontiers in Immunology
638 papers in training set
Top 5%
2.4%
12
Nature
645 papers in training set
Top 6%
2.1%
13
Cell Genomics
172 papers in training set
Top 2%
2.1%
14
Molecular Systems Biology
162 papers in training set
Top 1%
1.9%
15
Genome Biology
637 papers in training set
Top 5%
1.9%
16
Cell Reports Methods
165 papers in training set
Top 2%
1.9%
17
Science
477 papers in training set
Top 5%
1.7%
18
Scientific Reports
3612 papers in training set
Top 55%
1.7%
19
iScience
1154 papers in training set
Top 22%
1.4%
20
Communications Biology
993 papers in training set
Top 19%
1.3%
21
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 33%
1.3%
22
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.1%
23
Cell
431 papers in training set
Top 8%
1.1%
24
ImmunoInformatics
12 papers in training set
Top 0.1%
1.1%
25
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
26
Nature Genetics
286 papers in training set
Top 5%
1.0%
27
Bioinformatics Advances
203 papers in training set
Top 5%
0.8%
28
PLOS ONE
5266 papers in training set
Top 62%
0.8%
29
Nucleic Acids Research
1281 papers in training set
Top 15%
0.6%
30
The American Journal of Human Genetics
234 papers in training set
Top 3%
0.6%