Quantifying the Recoverability of V and J Genes from TCR CDR3 Sequences Using Generative Repertoire Models
Huang, S. J.; Baras, A. S.
Show abstract
Introduction: How much of the variable (V) and joining (J) gene identity of a T-cell receptor is recoverable from its third complementarity-determining region (CDR3) amino-acid sequence alone? Immune repertoire studies often report the CDR3 with V and J annotation that is missing, low-confidence, or inconsistent, so what the CDR3 alone can and cannot fix is both a basic question about the receptor and a practical one for reading those repertoires. Methods: For each of 118,096 pooled human rearrangements (37,687 and 80,409 {beta}) we computed the posterior distribution over candidate genes under a generative model of V(D)J recombination and under its post-selection counterpart, and measured recoverability by conditional entropy, the candidate-list size needed to contain the annotated gene, the fraction of sequences admitting a high-confidence single-gene call, and the structure of gene-by-gene confusion. Results: The J gene was nearly determined by the CDR3 in both chains. The V gene was only partially recoverable, and behaved as a group rather than a gene: junctional trimming and non-templated insertion, together with the loss of synonymous codon information in translation, leave sets of mutually confusable V genes whose grouping departs sharply from germline family nomenclature (adjusted Rand index 0.05 for and 0.21 for {beta}). Selection sharpened the V posterior modestly (usage-controlled entropy shift -0.06 nats for and -0.28 for {beta}) and redistributed which V gene was most probable, a locus-scale rewrite in {beta} against a mild reweight in . Both the recoverability measurements and the confusion grouping reproduced in two held-out tumor cohorts. Discussion: V identity is an emergent, system-level property of the repertoire, set jointly by recombination and selection and invisible in any single rearrangement, so it should be reported as a calibrated group rather than a single gene. We also release the pipeline with a computational tool which can output a set of candidate genes with confidence values given a CDR3 sequence.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Population variability in the generation and thymic selection of T-cell repertoires 94%
- TCR2HLA: calibrated inference of HLA genotypes from TCR repertoires enables identification of immunologically relevant metaclonotypes 93%
- Nucleotide context models outperform protein language models for predicting antibody affinity maturation 92%
Similar papers in this journal
- TCR meta-clonotypes for biomarker discovery with tcrdist3: identification of public, HLA-restricted SARS-CoV-2 associated TCR features 93%
- Characterisation of the immune repertoire of a humanised transgenic mouse through immunophenotyping and high-throughput sequencing 93%
- Population based selection shapes the T cell receptor repertoire during thymic development 93%
Similar papers in this journal
Similar papers in this journal
- The Observed T cell receptor Space database enables paired-chain repertoire mining, coherence analysis and language modelling 93%
- Dynamics of B-cell repertoires and emergence of cross-reactive responses in COVID-19 patients with different disease severity 92%
- A comprehensive map of the dendritic cell transcriptional network engaged upon innate sensing of HIV 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.