Back

Single-cell foundation models predict durable CAR T response despite imperfect cell annotation

Shen, L.; Bai, Z.; Yang, M.; Li, N.; Fan, R.

2026-07-23 systems biology
10.64898/2026.07.22.740224 bioRxiv
Show abstract

CD19-targeted chimeric antigen receptor (CAR) T cell therapy achieves high initial response rates in B-cell acute lymphoblastic leukemia (B-ALL), yet half of patients relapse within one year. Pre-infusion product composition decoded by single-cell RNA sequencing (scRNA-seq) carries information predictive of long-term CAR T persistence, but extracting this information from individual patients typically requires highly sophisticated bioinformatics expert annotation, limiting clinical translation. Here, we evaluate whether single-cell foundation models (scFMs) can extract clinically actionable information from engineered CAR T products. We applied four scFMs (scGPT, scFoundation, CellPLM and UCE), including fine-tuned versions of scGPT and scFoundation, to paired basal and CD19-stimulated pre-infusion CAR T products from 33 pediatric patients with B-ALL. Although annotation accuracy declined relative to healthy peripheral blood references, scFM-derived cell composition stratified patients with long-duration B-cell aplasia with a leave-one-out cross-validated area under the receiver operating characteristic curve of 0.879 (95% confidence interval, 0.742-0.986). Notably, foundation-model-identified cell proportion analysis matched or exceeded expert annotations for several predictive features, demonstrating that accurate clinical prediction may not require perfect per-cell annotation to begin with. CD8+XCL1/2+ cells were further identified as the biomarker consistently associated with durable CAR T persistence across models under CD19 stimulation, whereas other candidate populations showed limited reproducibility. Finally, we translate these findings into a locally deployable decision-support AI agent that predicts the probability of sustained CAR T persistence from pre-infusion CAR T scRNA-seq data.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.