FairTCR: Equity-Aware TCR--pMHC Binding Prediction\\Across HLA Alleles and Cohort Strata
Nowak, P.; Kowalski, J.; Lewandowski, T.
Show abstract
Public TCR-pMHC binding databases are heavily skewed toward a handful of well-studied HLA alleles-- most prominently HLA-A*02:01, which covers [~]45% of curated records--and toward patients from European-ancestry cohorts. Standard empirical risk minimization (ERM) trained on such data achieves strong pooled accuracy but routinely underperforms on rare alleles and underrepresented cohorts, creating systematic disparities that are invisible in single-metric benchmarks. We introduce FairTCR, a group distributionally robust optimization (GDRO) framework that minimizes worstgroup loss across HLA supertypes and cohort strata via online exponentiated gradient updates. FairTCR reduces the average- worst-group AUPRC disparity from 0.190 (ERM) to 0.098 on a curated VDJdb-IEDB benchmark, achieving a 48.4% disparity reduction while maintaining competitive average AUPRC (0.432 vs. 0.431 for ERM). Per-HLA analysis shows that rare allele groups (B*08:01, B*44:02) gain up to 0.062 AUPRC points, directly improving the equity of computational pre-screening for underrepresented patient populations.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EPIC-TRACE: predicting TCR binding to unseen epitopes using attention and contextualized embeddings 95%
- Consensus Label Propagation with Graph Convolutional Networks for Single-Cell RNA Sequencing Cell Type Annotation 94%
- LSMMD-MA: Scaling multimodal data integration for single-cell genomics data analysis 94%
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- Neural Network Models for Sequence-Based TCR and HLA Association Prediction 94%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 94%
Similar papers in this journal
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 94%
- Synthetic observations from deep generative models and binary omics data with limited sample size 94%
- Graph Contrastive Learning as a Versatile Foundation for Advanced scRNA-seq Data Analysis 92%
Similar papers in this journal
- Designing meaningful continuous representations of T cell receptor sequences with deep generative models 94%
- Sparse Epistatic Regularization of Deep Neural Networks for Inferring Fitness Functions 93%
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.