Back

AntiCP3: Prediction of Anticancer Proteins Using Evolutionary Information from Protein Language Models

Gupta, A.; Chauhan, M.; Tomer, R.; Raghava, G. P. S.

2025-05-03 bioinformatics
10.1101/2025.04.29.651196 bioRxiv
Show abstract

A number of computational methods have been developed in the past for predicting anticancer peptides, including AntiCP and AntiCP2 from our group. While these tools have been widely used by the scientific community, they are not suitable for predicting anticancer proteins. In this study, we present AntiCP3, the first dedicated method for the prediction of anticancer proteins. All models were trained using five-fold cross-validation and evaluated on an independent dataset not used during training. Our initial analysis revealed distinct compositional differences between anticancer peptides and proteins, justifying the need for a separate prediction framework. We first implemented similarity-based approaches, which yielded moderate performance. Subsequently, we developed machine learning and deep learning models using conventional protein features, achieving a maximum AUC of 0.72. The performance improved to an AUC of 0.79 with the incorporation of evolutionary information through PSSM profiles. Further enhancement was observed when embeddings from a fine-tuned protein language model ESM-t33 were used, leading to a best AUC of 0.90. Finally, a hybrid approach combining BLAST with our machine learning model achieved an AUC of 0.91. To facilitate the scientific community, we have implemented AntiCP3 as both a web server and standalone software for the prediction of anticancer proteins (https://webs.iiitd.edu.in/raghava/anticp3/). We have also deployed our model at hugging face https://huggingface.co/raghavagps-group/anticp3. Highlights Existing methods have been developed for predicting anticancer peptides. AntiCP3 is specifically optimised for predicting anticancer proteins. PSI-BLAST used to obtain evolutionary information in form of PSSM profile. Hybrid methods developed using alignment free and alignment based approach. A web server and standalone tool have been developed to assist the community.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
Briefings in Bioinformatics
354 papers in training set
Top 0.7%
9.5%
2
Computational Biology and Chemistry
28 papers in training set
Top 0.1%
7.8%
3
Computers in Biology and Medicine
128 papers in training set
Top 0.4%
6.6%
4
BMC Bioinformatics
457 papers in training set
Top 1%
6.6%
5
Bioinformatics
1204 papers in training set
Top 4%
4.8%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.8%
4.8%
7
PROTEOMICS
43 papers in training set
Top 0.2%
4.0%
8
Scientific Reports
3612 papers in training set
Top 35%
3.2%
9
PLOS ONE
5266 papers in training set
Top 38%
3.2%
50% of probability mass above
10
Journal of Biomolecular Structure and Dynamics
43 papers in training set
Top 0.4%
3.2%
11
ACS Omega
105 papers in training set
Top 0.6%
3.2%
12
Frontiers in Bioinformatics
49 papers in training set
Top 0.2%
2.7%
13
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
2.4%
14
Bioinformatics Advances
203 papers in training set
Top 3%
1.9%
15
BioData Mining
22 papers in training set
Top 0.2%
1.9%
16
PeerJ
308 papers in training set
Top 5%
1.9%
17
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.6%
1.7%
18
PLOS Computational Biology
1863 papers in training set
Top 16%
1.4%
19
Frontiers in Pharmacology
111 papers in training set
Top 2%
1.1%
20
Pharmaceuticals
34 papers in training set
Top 0.8%
1.1%
21
F1000Research
88 papers in training set
Top 3%
1.1%
22
BioMed Research International
28 papers in training set
Top 2%
1.1%
23
Genomics
64 papers in training set
Top 1%
1.1%
24
Informatics in Medicine Unlocked
22 papers in training set
Top 1.0%
1.0%
25
International Journal of Molecular Sciences
494 papers in training set
Top 13%
1.0%
26
Journal of Computational Biology
48 papers in training set
Top 0.9%
1.0%
27
Biology Methods and Protocols
61 papers in training set
Top 2%
1.0%
28
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 1%
1.0%
29
Protein Science
246 papers in training set
Top 3%
0.9%
30
Database
61 papers in training set
Top 1%
0.8%