Back

UcTCRp: a TCRβ-based framework for quantitative MAIT- and iNKT-associated repertoire-state profiling

Chen, L.; Li, Y.; Shan, S.; Wang, K.; Feng, C.; Dou, Y.; Xu, Q.; Cai, L.; Wang, H.; Wang, H.; Bo, X.; Zhang, J.

2026-05-28 bioinformatics
10.64898/2026.05.25.727598 bioRxiv
Show abstract

MAIT and iNKT cells are conventionally identified using invariant or semi-invariant TCR chains, antigen-loaded tetramers, or transcriptomic phenotypes. These requirements limit their detection in public and clinical immune-repertoire datasets that contain only TCR{beta} sequences. Here we present UcTCRp, a TCR{beta}-only framework for profiling MAIT- and iNKT-associated repertoire states in bulk immune repertoires. UcTCRp integrates V-gene context and CDR3{beta} sequence features using a transformer-based representation pretrained on more than one million TCR{beta} sequences and supervised with curated cross-species MAIT, iNKT and conventional T cell references. The framework defines conserved model-informative TCR{beta} features, uses V-matched negative sampling to reduce germline-segment shortcuts, and generalizes across independent human and mouse datasets. In paired scRNA-seq/scTCR-seq datasets, UcTCRp recovered transcriptome-defined MAIT and iNKT cells and identified additional MAIT-like candidates supported by receptor evidence but missed by expression-only annotation. Bulk calibration against paired single-cell references and synthetic spike-in experiments established operating characteristics for repertoire-level abundance estimation. These results establish unpaired TCR{beta} repertoires as an actionable substrate for reconstructing unconventional T cell-associated immune states, enabling archived repertoire resources to be repurposed for systems-level studies of tissue immunity, disease and therapeutic response.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.