Discovering and overcoming the bias in neoantigen identification by unified machine learning models
Zhang, Z.; Wu, W.; Wei, L.; Wang, X.
Show abstract
Neoantigens play a crucial role in tumor immune process and precisely identifying them can greatly contribute to tumor immunotherapy design. There are three main steps in the neoantigen immune process, i.e., binding with MHCs, extracellular presentation, and immunogenicity induction. Various computational methods have been developed, but the overall accuracy of neoantigen identification remains relatively low. Here, we established a unified transformer-based framework ImmuBPI that comprised three tasks. Cross-task model interpretation discovered a counterfactual pattern learned by the immunogenicity prediction model overlooked in previous studies. We demonstrated that this model bias arose from training data imbalance and the neglect of bias control would not only mislead the models but also hinder existing benchmarks from fair evaluation. We further designed a mutual information-based debiasing strategy that showed preliminary potential in alleviating the bias. We believe these observations will provide insightful perspectives for future neoantigen prediction by high-lighting the necessity of bias control.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Personalized deep learning of individual immunopeptidomes to identify neoantigens for cancer vaccines 95%
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 94%
- Inferring spatial single-cell-level interactions through interpreting cell state and niche correlations learned by self-supervised graph transformer 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.