Back

HLA Alleles Imprint Distinct Biases in the Usage Preferences of TCR Vβ segments

Castorina, L. V.; Noakes, M. T.; Pisani, L.; Greissl, J.; Robins, H.; Chen-Harris, H.; Zahid, H. J.

2026-03-19 immunology
10.64898/2026.02.19.706729 bioRxiv
Show abstract

T cells must co-recognize peptide and HLA, yet the extent to which this specificity is shaped by germline-encoded TCR-HLA contacts versus selection during thymic development has remained difficult to quantify. Leveraging population-scale TCR{beta} repertoires linked to donor HLA genotypes, we construct allele-specific V{beta} usage profiles and normalize them to repertoire-wide baselines to derive interpretable HLA-TCR V{beta} preference vectors. We demonstrate that different HLA alleles imprint distinct biases in the usage frequency of TCR V{beta} gene segments among the public TCRs that engage those alleles; consistent with germline-encoded TCR-HLA contact preferences, certain V{beta} genes are over-represented for particular HLA alleles. Similarities in HLA amino-acid sequence predict similarities in both their V{beta} preferences and peptide-binding motifs; a residue-level analysis disentangles HLA positions primarily associated with TCR engagement from those associated with peptide motifs. The TCR-associated HLA positions localize to canonical TCR-facing helices, whereas peptide-associated HLA positions track binding pockets, revealing distinct molecular routes by which HLA polymorphism shapes the TCR and peptide sides of recognition. CMV exposure stratifications confirmed these patterns are not explained by a few dominant infections. Together, these data support that germline-encoded constraints set the landscape of TCR-HLA compatibility, while thymic and peptide-driven forces tune the realized repertoire. These HLA-specific V{beta} biases are a biological prior that should provide a baseline for better understanding of TCR-pHLA specificity as a whole and should be accounted for in future evaluation of any TCR-pHLA specificity prediction methods.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.