Back

Identifying immune signatures of common exposures through co-occurrence of T-cell receptors in tens of thousands of donors

May, D. H.; Woodhouse, S.; Zahid, H. J.; Elyanow, R.; Doroschak, K.; Noakes, M. T.; Taniguchi, R.; Yang, Z.; Grino, J.; Byron, R.; Oaks, J.; Sherwood, A.; Greissl, J.; Chen-Harris, H.; Howie, B.; Robins, H. S.

2024-03-27 immunology
10.1101/2024.03.26.583354 bioRxiv
Show abstract

Memory T cells are records of clonal expansion from prior immune exposures, such as infections, vaccines and chronic diseases like cancer. A subset of the receptors of these expanded T cells in a typical immune repertoire are highly public, i.e., present in many individuals exposed to the same exposure. For the most part, the exposures associated with these public T cells are unknown. To identify public T-cell receptor signatures of immune exposures, we mined the immunosequencing repertoires of tens of thousands of donors to define clusters of co-occurring T cells. We first built co-occurrence clusters of T cells responding to antigens presented by the same Human Leukocyte Antigen (HLA) and then combined those clusters across HLAs. Each cross-HLA cluster putatively represents the public T-cell signature of a single prevalent exposure. Using repertoires from donors with known serological status for 7 prevalent exposures (HSV-1, HSV-2, EBV, Parvovirus, Toxoplasma gondii, Cytomegalovirus and SARS-CoV-2), we identified a single T-cell cluster strongly associated with each exposure and used it to construct a highly sensitive and specific diagnostic model for the exposure. These T-cell clusters constitute the public immune responses to prevalent exposures, 7 known and many others unknown. By learning the exposure associations for more T-cell clusters, this approach could be used to derive a ledger of a persons past and present immune exposures.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.