Back

Single-cell characterization and machine learning-based classification reveal the transcriptional identity and antigen-experienced features of peripheral CD4+CD8+ double-positive T cells

Shin, E.; Yun, S. G.; Cho, Y.

2026-01-02 immunology
10.64898/2026.01.01.697324 bioRxiv
Show abstract

T cell mediated immunity depends on diverse subsets with distinct regulatory and cytotoxic roles. While CD4 and CD8 T cells have traditionally been viewed as separate lineages, double-positive T (DPT) cells coexpressing both markers have emerged as a rare subset with potential roles in immune modulation, cytotoxicity, and memory. Here, we integrated single-cell RNA sequencing, single-cell TCR sequencing, and supervised machine learning to characterize transcriptionally defined DPT cells in human peripheral blood. DPT cells formed a reproducible transcriptional state distinct from conventional single-positive T cells and exhibited gene expression programs associated with cytotoxic effector potential, antigen experience, and immune cell migration. Clonal repertoire analysis revealed enrichment of expanded clonotypes and sharing of identical TCR clonotypes with both CD4 and CD8 T cell populations, indicating shared antigen-driven clonal histories rather than a distinct developmental lineage. To enable robust identification of this rare state, we developed a machine learning classifier trained on sorted CD4 and CD8 T cells, which outperformed marker-based approaches. Application of this framework to public COVID-19 single-cell datasets demonstrated its generalizability under immune perturbation. Together, these findings establish DPT cells as a reproducible, antigen-experienced transcriptional state with cytotoxicity- and migration-associated programs, and provide a framework for systematic identification of rare T cell states, expanding our understanding of T cell diversity.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.