Back

T-cell receptor specific protein language model for prediction and interpretation of epitope binding (ProtLM.TCR)

Essaghir, A.; Sathiyamoorthy, N. K.; Smyth, P.; Postelnicu, A.; Ghiviriga, S.; Ghita, A.; Singh, A.; Kapil, S.; Phogat, S.; Singh, G.

2022-11-29 bioinformatics
10.1101/2022.11.28.518167 bioRxiv
Show abstract

The cellular adaptive immune response relies on epitope recognition by T-cell receptors (TCRs). We used a language model for TCRs (ProtLM.TCR) to predict TCR-epitope binding. This model was pre-trained on a large set of TCR sequences (~62.106) before being fine-tuned to predict TCR-epitope bindings across multiple human leukocyte antigen (HLA) of class-I types. We then tested ProtLM.TCR on a balanced set of binders and non-binders for each epitope, avoiding model shortcuts like HLA categories. We compared pan-HLA versus HLA-specific models, and our results show that while computational prediction of novel TCR-epitope binding probability is feasible, more epitopes and diverse training datasets are required to achieve a better generalized performances in de novo epitope binding prediction tasks. We also show that ProtLM.TCR embeddings outperform BLOSUM scores and hand-crafted embeddings. Finally, we have used the LIME framework to examine the interpretability of these predictions.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
ImmunoInformatics
12 papers in training set
Top 0.1%
18.3%
2
Briefings in Bioinformatics
354 papers in training set
Top 0.6%
9.7%
3
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.3%
6.6%
4
Bioinformatics
1204 papers in training set
Top 3%
6.6%
5
Scientific Reports
3612 papers in training set
Top 13%
6.2%
6
PLOS Computational Biology
1863 papers in training set
Top 6%
6.2%
50% of probability mass above
7
mAbs
32 papers in training set
Top 0.1%
4.3%
8
Frontiers in Immunology
638 papers in training set
Top 3%
4.3%
9
Bioinformatics Advances
203 papers in training set
Top 2%
3.4%
10
BMC Bioinformatics
457 papers in training set
Top 3%
3.2%
11
Communications Biology
993 papers in training set
Top 6%
3.2%
12
iScience
1154 papers in training set
Top 9%
2.8%
13
Protein Science
246 papers in training set
Top 2%
2.1%
14
Frontiers in Bioinformatics
49 papers in training set
Top 0.3%
2.1%
15
Nature Machine Intelligence
70 papers in training set
Top 1%
1.7%
16
Nature Communications
5641 papers in training set
Top 45%
1.7%
17
PLOS ONE
5266 papers in training set
Top 58%
1.0%
18
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
19
Computers in Biology and Medicine
128 papers in training set
Top 4%
0.8%
20
Journal of Proteome Research
234 papers in training set
Top 2%
0.8%
21
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 1%
0.8%
22
Advanced Science
286 papers in training set
Top 11%
0.6%
23
International Journal of Molecular Sciences
494 papers in training set
Top 18%
0.6%
24
Journal of Chemical Information and Modeling
238 papers in training set
Top 3%
0.6%
25
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 1%
0.6%
26
GigaScience
212 papers in training set
Top 5%
0.6%