A unified language model bridging de novo and fragment-based 3D molecule design delivers potent CBL-B inhibitors for cancer treatment
Wang, H.; Sun, G.; Zhang, B.; Wang, Y.; Xi, B.; Yang, M.; Liu, C.; Ge, Y.; Fan, F.; Feng, W.; Zhu, Y.; Xiao, Y.; Wang, Y.; Liu, Z.; Jiang, D.; Wang, H.; Zhou, W.; Huang, B.
Show abstract
The rational design of small molecules is central to drug discovery, yet current artificial intelligence (AI) methodologies for generating three-dimensional (3D) molecules are often siloed, focusing on either de novo design or fragment-based design. The lack of a holistic framework limits AIs application across the complex and multi-step pipeline spanning from novel scaffold identification to lead compound optimization, and prevents AI from effectively learning from the entire process. Here, we introduce UniLingo3DMol, a language model for 3D molecular generation, empowered by fragment permutation-capable molecular representation alongside multi-stage and multi-task training strategy. This integrated design enables UniLingo3DMol to seamlessly span both de novo and fragment-retained molecular design, demonstrating superior performance over existing generation models in in silico evaluations across more than 100 diverse biological targets. We further leveraged UniLingo3DMol in the design of inhibitors targeting CBL-B, a crucial immune E3 ubiquitin ligase and attractive immunotherapy target. This strategy led to a lead compound demonstrating excellent in vitro activity and robust in vivo anti-tumor efficacy. Our findings establish UniLingo3DMol as a generalized and powerful platform, showing the strong potential to advance AI-driven drug discovery.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Efficient Generation of Protein Pockets with PocketGen 97%
- PSICHIC: physicochemical graph neural network for learning protein-ligand interaction fingerprints from sequence data 96%
- TrustAffinity: accurate, reliable and scalable out-of-distribution protein-ligand binding affinity prediction using trustworthy deep learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.