Back

Benchmarking of Quantum SVM and Classical ML Algorithms for Prediction of Therapeutic Proteins

Tijare, P.; Mehta, N. K.; Raghava, G. P. S.

2025-05-07 bioinformatics
10.1101/2025.04.30.651419 bioRxiv
Show abstract

Over the past decade, quantum machine learning, particularly quantum support vector machines (QSVMs), has emerged as an optimistic alternative to classical machine learning (CML) techniques. This study rigorously benchmarks the performance of QSVM and CML-based models across four diverse datasets relevant to therapeutic proteins and peptides. Specifically, we evaluated these approaches for the prediction of B-cell epitopes (CLBtope), exosomal proteins (ExoPropred), hemolytic peptides (HemoPI), and toxic peptides (Toxinpred3). The maximum area under the receiver operating characteristic curve (AUC) for the CLBtope dataset achieved was 0.68 for QSVM and 0.82 for CML models. For the ExoPropred dataset, the maximum AUCs were 0.66 (QSVM) and 0.72 (CML). In contrast, both QSVM and CML models demonstrated high performance on the HemoPI dataset, yielding maximum AUCs of 0.95 and 0.98, respectively. Similarly, for the Toxinpred3 dataset, the maximum AUCs were 0.84 (QSVM) and 0.94 (CML). All models were evaluated using independent validation datasets not used during training. These results suggest that although CML currently demonstrates superior predictive capability for these tasks, the similar progression in performance indicates potential for future advancements in QSVM. HighlightsO_LIComparative study of QSVM and CML models on four bioinformatics datasets C_LIO_LIQSVM performance tries to approach CML in tasks involving hemolytic and toxic peptide prediction C_LIO_LIIndependent validation confirms robustness of performance metrics C_LIO_LIResults highlight the potential of QSVMs as real-world quantum hardware continues to matures C_LI

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.