Geographically Biased Composition of NetMHCpan Training Datasets and Evaluation of MHC-Peptide Binding Prediction Accuracy on Novel Alleles
Atkins, T. K.; Solanki, A.; Cornette, J.; Vasmatzis, G.; Riedel, M.
Show abstract
Bias in neural network model training datasets has been observed to decrease prediction accuracy for groups underrepresented in training data. Thus, investigating the composition of training datasets used in machine learning models with health-care applications is vital to ensure equity. Two such machine learning models are NetMHCpan-4.1 and NetMHCIIpan-4.0, used to predict antigen binding scores to major histocompatibility complex class I and II molecules, respectively. As antigen presentation is a critical step in mounting the adaptive immune response, previous work has used these or similar predictions models in a broad array of applications, from explaining asymptomatic viral infection to cancer neoantigen prediction. However, these models have also been shown to be biased toward hydrophobic peptides, suggesting the network could also contain other sources of bias. Here, we report the composition of the networks training datasets are heavily biased toward European Caucasian individuals and against Asian and Pacific Islander individuals. We test the ability of NetMHCpan-4.1 and NetMHCpan-4.0 to distinguish true binders from randomly generated peptides on alleles not included in the training datasets. Unexpectedly, we fail to find evidence that the disparities in training data lead to a meaningful difference in prediction quality for alleles not present in the training data. We attempt to explain this result by mapping the HLA sequence space to determine the sequence diversity of the training dataset. Furthermore, we link the residues which have the greatest impact on NetMHCpan predictions to structural features for three alleles (HLA-A*34:01, HLA-C*04:03, HLA-DRB1*12:02).
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Integrated Approach to the Characterization of Immune Repertoires Using AIMS: An Automated Immune Molecule Separator 95%
- Neural Network Models for Sequence-Based TCR and HLA Association Prediction 95%
- THLANet: A Deep Learning Framework for Predicting TCR-pHLA Binding in Immunotherapy Applications 94%
Similar papers in this journal
- Machine learning optimization of peptides for presentation by class II MHCs 96%
- BERTrand - peptide:TCR binding prediction using Bidirectional Encoder Representations from Transformers augmented with random TCR pairing 95%
- EPIC-TRACE: predicting TCR binding to unseen epitopes using attention and contextualized embeddings 95%
Similar papers in this journal
Similar papers in this journal
- VaxOptiML: Leveraging Machine Learning for Accurate Prediction of MHC-I & II Epitopes for Optimized Cancer Immunotherapy 92%
- CDR3 and V- genes show distinct reconstitution patterns in T-cell repertoire post allogenic bone marrow transplantation 91%
- MHC genotyping from rhesus macaque exome sequences 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.