Back

Quantifying the Recoverability of V and J Genes from TCR CDR3 Sequences Using Generative Repertoire Models

Huang, S. J.; Baras, A. S.

2026-08-25 immunology
10.64898/2026.08.24.746073 bioRxiv
Show abstract

Introduction: How much of the variable (V) and joining (J) gene identity of a T-cell receptor is recoverable from its third complementarity-determining region (CDR3) amino-acid sequence alone? Immune repertoire studies often report the CDR3 with V and J annotation that is missing, low-confidence, or inconsistent, so what the CDR3 alone can and cannot fix is both a basic question about the receptor and a practical one for reading those repertoires. Methods: For each of 118,096 pooled human rearrangements (37,687 and 80,409 {beta}) we computed the posterior distribution over candidate genes under a generative model of V(D)J recombination and under its post-selection counterpart, and measured recoverability by conditional entropy, the candidate-list size needed to contain the annotated gene, the fraction of sequences admitting a high-confidence single-gene call, and the structure of gene-by-gene confusion. Results: The J gene was nearly determined by the CDR3 in both chains. The V gene was only partially recoverable, and behaved as a group rather than a gene: junctional trimming and non-templated insertion, together with the loss of synonymous codon information in translation, leave sets of mutually confusable V genes whose grouping departs sharply from germline family nomenclature (adjusted Rand index 0.05 for and 0.21 for {beta}). Selection sharpened the V posterior modestly (usage-controlled entropy shift -0.06 nats for and -0.28 for {beta}) and redistributed which V gene was most probable, a locus-scale rewrite in {beta} against a mild reweight in . Both the recoverability measurements and the confusion grouping reproduced in two held-out tumor cohorts. Discussion: V identity is an emergent, system-level property of the repertoire, set jointly by recombination and selection and invisible in any single rearrangement, so it should be reported as a calibrated group rather than a single gene. We also release the pipeline with a computational tool which can output a set of candidate genes with confidence values given a CDR3 sequence.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.