Phenotypic Bioactivity Prediction as Open-set Biological Assay Querying
Sun, Y.; Zhang, X.; Zheng, Q.; Li, H.; Zhang, J.; Hong, L.; Wang, Y.; Zhang, Y.; Xie, W.
Show abstract
The traditional drug discovery pipeline is severely bottlenecked by the need to design and execute bespoke biological assays for every new target and compound--a process that is both time-consuming and prohibitively expensive. While machine learning has accelerated virtual screening, current models remain confined to "closed-set" paradigms, unable to generalize to entirely novel biological assays without target-specific experimental data. Here, we present OpenPheno, a groundbreaking multimodal foundation model that fundamentally redefines bioactivity prediction as an open-set, visual-language question-answering (QA) task. By integrating chemical structures (SMILES), universal phenotypic profiles (Cell Painting images), and natural language descriptions of biological assays, OpenPheno unlocks the highly coveted "profile once, predict many" paradigm. Instead of conducting countless target-specific wet-lab experiments, researchers only need to capture a single, low-cost Cell Painting image of a novel compound. OpenPheno then evaluates this universal phenotypic "fingerprint" against the text-based description of any unseen assay, predicting bioactivity in a zero-shot manner. On 54 entirely unseen assays, it achieves strong zero-shot performance (mean AUROC 0.75), exceeding supervised baselines trained with full labeled data, and few-shot adaptation further improves predictions. In the most stringent setting where both compounds and assays are novel, OpenPheno maintains robust generalization (mean AUROC 0.66), opening up a new paradigm for a highly scalable, cost-effective, and universal engine for next-generation drug discovery.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TrustAffinity: accurate, reliable and scalable out-of-distribution protein-ligand binding affinity prediction using trustworthy deep learning 95%
- Annotating metabolite mass spectra with domain-inspired chemical formula transformers 94%
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 94%
Similar papers in this journal
Similar papers in this journal
- ProtNote: a multimodal method for protein-function annotation 95%
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 94%
- Deep Local Analysis deconstructs protein-protein interfaces and accurately estimates binding affinity changes upon mutation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.