Can SMILES be fragmented into a concatenable ordered sequence of retrosynthetically interesting string block ?
Reboul, E.; Prabakaran, H.; Baaden, M.; Waldispuhl, J.; Taly, A.
Show abstract
Molecules generated by deep learning models are often difficult to synthesize. Their synthetic accessibility can be improved with automated retrosynthetic analysis, which allows for identifying synthons. However, synthons in a SMILES can be scattered throughout the string depending on the path taken through the molecular graph used to generate the SMILES. We tested whether the ensemble of possible SMILES for a molecule can be used to generate a concatenable ordered sequence of string fragments (blocks) from SMILES that match potential synthons obtained through automated retrosynthetic analysis. We found that exhaustively sampling the SMILES space of a molecule improves the coverage of retrosynthetic breaks. We achieved full coverage of retrosynthetic bonds in string form for 85\% of the 1.9 million molecules in the MOSES dataset. Doing so allowed us to test our block SMILES in an unconditional de novo drug design test case with MolGPT and Monte Carlo Tree Search (MCTS). We found that using blocks as an LLM's token did degrade MolGPT performance due to the curse of dimensionality. However, using the SMILES selected by our blocking algorithm with the default SMILES tokenizer improved the reproduction of physico-chemical properties of samples and also improved uniqueness, novelty, and validity. The MCTS outperforms our MolGPT models in terms of validity and novelty. However, samples generated by the MCTS had physico-chemical properties that were further away from the MOSES baseline than the samples produced by molGPT, with an improved distribution of quantitative estimation of drug-likeness (QED).
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- DrugDiff - small molecule diffusion model with flexible guidance towards molecular properties 97%
- Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data 95%
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 94%
Similar papers in this journal
- Rational Discovery of Dual-Action Multi-Target Kinase Inhibitor for Precision Anti-Cancer Therapy Using Structural Systems Pharmacology 95%
- Controlling astrocyte-mediated synaptic pruning signals for schizophrenia drug repurposing with Deep Graph Networks 95%
- Predicting kinase inhibitors using bioactivity matrix derived informer sets 94%
Similar papers in this journal
Similar papers in this journal
- G-PLIP: Knowledge graph neural network for structure-free protein-ligand bioactivity prediction 94%
- Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo) 92%
- HerbComb: an integrated database for the discovery of novel combinational therapies from herbal medicines 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.