Back

CLigOpt: Controllable Ligand Design through Target-Specific Optimisation

Li, Y.; da Costa Avelar, P. H.; Chen, X.; Zhang, L.; Wu, M.; Tsoka, S.

2024-03-17 bioinformatics
10.1101/2024.03.15.585255 bioRxiv
Show abstract

AO_SCPCAPBSTRACTC_SCPCAPO_ST_ABSMotivationC_ST_ABSKey challenge in deep generative models for molecular design is to navigate random sampling of the vast molecular space, and produce promising molecules that compromise property controls across multiple chemical criteria. Fragment-based drug design (FBDD), using fragments as starting points, is an effective way to constrain chemical space and improve generation of biologically active molecules. Furthermore, optimisation approaches are often implemented with generative models to search through chemical space, and identify promising samples which satisfy specific properties. Controllable FBDD has promising potential in efficient target-specific ligand design. ResultsWe propose a controllable FBDD model, CLigOpt, which can generate molecules with desired properties from a given fragment pair. CLigOpt is a Variational AutoEncoder-based model which utilises co-embeddings of node and edge features to fully mine information from molecular graphs, as well as a multi-objective Controllable Generation Module to generate molecules under property controls. CLigOpt achieves consistently strong performance in generating structurally and chemically valid molecules, as evaluated across six metrics. Applicability is illustrated through ligand candidates for hDHFR and it is shown that the proportion of feasible active molecules from the generated set is increased by 10%. Molecular docking and synthesisability prediction tasks are conducted to prioritise generated molecules to derive potential lead compounds. Availability and ImplementationThe source code is available via https://github.com/yutongLi1997/CLigOpt-Controllable-Ligand-Design-through-Target-Specific-Optimisation.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.