Back

BenchRep-T: A Systematic Evaluation of T-Cell Repertoire-Based Disease Diagnostics

Im, C.; Cohen-Lavi, L.; Buendia, A.; Kundaje, A.; Boyd, S. D.

2026-06-10 immunology
10.64898/2026.06.09.727013 bioRxiv
Show abstract

Adaptive immune receptor repertoire sequencing data has emerged as a promising potential modality for disease diagnosis, relying on computational methods to analyze T-cell receptor (TCR) sequences from an individuals blood sample. Published methods rely on different cohorts, data preprocessing pipelines, and evaluation metrics, making direct comparison across methods challenging. We present BenchRep-T, a unified benchmark that standardizes multiple publicly available TCR repertoire datasets and evaluates nine computational approaches, spanning statistical enrichment of shared sequences, feature-engineered ensembles, deep learning, and sequence clustering. BenchRep-T evaluates methods on four tasks: disease classification across conditions, performance scaling under restricted sequence-sampling depth, recovery of known antigen-specific driver sequences, and evaluation of sensitivity to demographic confounding. Under controlled evaluation, simple baselines prove competitive, with tree-based models trained on V- and J-gene usage and short sequence motifs approaching the classification performance of more complex methods. Our findings underscore the complexity of modeling TCR repertoire data, and show that no single method dominates across all tasks. BenchRep-T provides a framework for rigorous and reproducible evaluation of TCR repertoire classification methods to accelerate the development of immune repertoire-based diagnostics.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.