Back

How to select the best zero-shot model for the viral proteins?

Yu, Y.; Jiang, F.; Zhong, B.; Hong, L.; Li, M.

2024-10-06 bioinformatics
10.1101/2024.10.06.616860 bioRxiv
Show abstract

Predicting the fitness of viral proteins holds notable implications for understanding viral evolution, advancing fundamental biological research, and informing drug discovery. However, the considerable variability and evolution of viral proteins make predicting mutant fitness a major challenge. This study introduces the ProPEC, a Perplexity-based Ensemble Model, aimed at improving the performance of zero-shot predictions for protein fitness across diverse viral datasets. We selected five representative pretrained language models (PLMs) as base models. ProPEC, which integrates perplexity-weighted scores from these PLMs with GEMME, demonstrates superior performance compared to individual models. Through parameter sensitivity analysis, we highlight the robustness of perplexity-based model selection in ProPEC. Additionally, a case study on T7 RNA polymerase activity dataset underscores ProPECs predictive capabilities. These findings suggest that ProPEC offers an effective approach for advancing viral protein fitness prediction, providing valuable insights for virology research and therapeutic development. TOC Graphic O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/616860v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@30be8forg.highwire.dtl.DTLVardef@2ebf35org.highwire.dtl.DTLVardef@10b517eorg.highwire.dtl.DTLVardef@134681_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.