Journal of Chemical Information and Modeling
● American Chemical Society (ACS)
All preprints, ranked by how well they match Journal of Chemical Information and Modeling's content profile, based on 238 papers previously published here. The average preprint has a 0.19% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Herron, L.; Dakka, J.; Yao, K.; Shi, D.; Zhang, Y.; Jerome, S. V.
Show abstract
Recent years have seen a rise in applications of deep learning to problems in the molecular sciences. Among them, the diffusion model DiffDock stands out as a method for docking small molecules into protein binding sites. But DiffDock struggles to compete with conventional docking methods, especially for targets outside its training set. We develop a hybrid model called DiffDock-Glide which addresses some shortcomings of deep learning docking methods: it uses a modified generative process to generate samples within a binding pocket and the confidence model is replaced with Glides post-docking minimization pipeline. We evaluate DiffDock-Glide on the Posebusters dataset and show improved sampling of near-native poses, especially for sequences without homologues in the training set. We also evaluate DiffDock-Glides performance in virtual screening compounds from the DUD-E dataset against receptor structures generated by AlphaFold2 and report enrichment values that broadly surpass those from traditional Glide.
Surendran, A.; Zsigmond, K.; Quintana, R. A. M.
Show abstract
The visualization of high-dimensional chemical space is a critical tool for under-standing molecular diversity, structure-property relationships, and for guiding compound selection. However, the performance of non-linear dimensionality reduction (DR) techniques like t-Stochastic Neighborhood Embedding (t-SNE), Uniform Man-ifold Approximation and Projection (UMAP), and Generative Topographic Mapping (GTM) are often susceptible to the choice of hyperparameters, along with the high cost of their training for large datasets. In this study, we investigated the effect of undersampling methods on the choice of hyperparameter selection for these non-linear dimensionality reduction methods. Our results demonstrate that selecting small representative subsets of chemical data not only reduces computational costs associated with hyperparameter training but also serves as an innovative means to train non-linear DR methods, leading to projections that better preserve the local structure within the chemical space.
Zhang, H.; Zheng, J.; Li, H.; Wan, S.; Guan, G.; Li, B.; Liu, L.; He, W.
Show abstract
Traditional computational drug discovery approaches struggle to accurately evaluate the biological authenticity of protein-ligand binding conformations due to inherent limitations in empirical scoring functions and force field approximations. This study proposes MFPLI - a deep learning framework integrating multimodal physicochemical features to systematically assess the biological authenticity alignment between molecular docking poses and true co-crystal structures. By establishing a continuous surface characterization system for protein-ligand interfaces, we concurrently incorporate geometric curvature features (radius, shape index) and chemical interaction fields (electrostatic potential, hydrogen-bond networks, hydrophobicity gradients). A contrastive learning architecture based on Siamese equivariant graph neural networks was developed to enable discriminative analysis between co-crystal conformations and parameter-perturbed pseudo-conformations generated through inverse docking. The five-channel fusion model demonstrates robust performance on the time-split PoseBuster validation set (AUC=0.91), with predicted Euclidean distance deviation ({Delta}E) effectively distinguishing native co-crystal conformations from aberrant docking poses in 80% of samples. Notably, 71% of {Delta}E-negative samples concentrate within the [-0.3, 0] interval, reflecting physical consistency between model predictions and conformational transition processes. This framework establishes a novel paradigm for biological authenticity assessment in virtual screening for computer-aided drug discovery through synergistic modeling of surface topology and interaction chemistry.
Wang, H.; Zhang, Z.; Gong, H.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWRecent advancements in deep learning have greatly prompted the de novo design of drugs and materials. Previous studies have shown that a well-designed molecular representation is critical for improving the accuracy of deep-learning-based molecular property prediction methods. However, the lack of large-scale data enriched with detailed physicochemical information hinders effective learning of an informative molecular representation. To fill this data gap, we introduce qcMol, a dataset consisting of 1.2 million molecules from 95 datasets with high-quality quantum chemical annotations, to facilitate molecular representation learning as well as downstream molecular property prediction. Chemicals in this dataset include drug-like compounds, metabolites and molecules with matched experimental data, covering 247,448 kinds of scaffolds and a broad spectrum of molecular sizes. Each compound in qcMol is annotated with detailed quantum chemical information, obtained through reliable quantum chemical calculations based on B3LYP-D3/def2-SV(P)//GFN2-xTB as well as the follow-up wave function post-analysis. These features are organized into multiple formats, allowing for flexible integration into diversified molecular representation learning frameworks. The broad data distribution, comprehensive quantum chemical annotations and flexible data formats jointly enable qcMol to serve as the pre-training resource as well as the benchmark test set for deep learning models, benefiting the practical in silico drug discovery. qcMol is freely accessible from https://structpred.life.tsinghua.edu.cn/qcmol/.
Zhao, H.; Nittinger, E.; Tyrchan, C.
Show abstract
Chemical space exploration has gained significant interest with the increase in available building blocks, which enables the creation of ultra-large virtual libraries containing billions or even trillions of compounds. However, the challenge of selecting most suitable compounds for synthesis arises, and one such challenge is hit expansion. Recently, Thompson sampling, a probabilistic search approach, has been proposed by Walters et al. to achieve efficiency gains by operating in the reagent space rather than the product space. Here, we aim to address some of its shortcomings and propose optimizations. We introduce a warmup routine to ensure that initial probabilities are set for all reagents with a minimum number of molecules evaluated. Additionally, a roulette wheel selection is proposed with adapted stop criteria to improve sampling efficiency, and belief distributions of reagents are only updated when they appear in new molecules. We demonstrate that a 100% recovery rate can be achieved by sampling 0.1% of the fully enumerated library, showcasing the effectiveness of our proposed optimizations.
Thompson, T. D.; Miao, Y.
Show abstract
G protein-coupled receptor (GPCR) allosteric modulators (AMs) offer significant therapeutic advantages over orthosteric drugs, yet structure-based virtual screening lacks validated protocols accounting for the conformational complexity of GPCR allosteric sites. We benchmark docking protocols using PDB experimental structures and structural ensembles derived from Gaussian accelerated Molecular Dynamics (GaMD) simulations across four Class A GPCRs (including the muscarinic M2 and M4 receptors, the {beta}2-adrenergic receptor, and the C-C chemokine receptor type 2) with four programs (Glide HTVS, AutoDock Vina, DOCK3.8, and Boltz-2) against experimentally validated modulator libraries and property-matched decoys. GaMD ensemble docking improved early AM enrichment across all four targets under at least one program. Glide ensemble docking was the only protocol to consistently improve early AM recovery across all four targets, ranking known actives almost exclusively within the top 0.5% of compounds at CCR2 and improving M2R active recovery nearly 9-fold relative to the PDB structure. GaMD free-energy landscape topology governed ensemble re-ranking strategy selection: population-skewed landscapes favored top binding energy ranking (BEmin) while flat, multi-populated landscapes favored average binding energy ranking (BEavg), and at targets with dominant low-energy states, a single GaMD cluster matched or exceeded full ensemble or PDB performance. Taking the union of top percentile hits identified by both ensemble re-ranking methods, BEmin / BEavg, maximizes chemical diversity at the earliest percentiles. Program-specific scaffold recovery biases further motivated a consensus BEmin / BEavg approach to maximize hit diversity. The Boltz-2 deep-learning program showed minimal sensitivity to GaMD templates and underperformed conventional docking, suggesting its affinity predictions complement rather than replace physics- and empirical-based docking approaches for GPCR AM screening.
Jung, N.; Park, H.; Yang, J.; Seok, C.
Show abstract
Virtual screening has long been a central computational tool for rational ligand discovery, enabling the systematic prioritization of candidate molecules from large chemical libraries. Although docking and related approaches that explicitly account for receptor-ligand interactions have been developed and refined over several decades, achieving both reliable receptor-aware interaction modeling and computational scalability remains an open challenge, particularly for ultra-large chemical spaces. Ligand-based methods are fast and robust but do not explicitly incorporate receptor structure, whereas docking-based approaches model receptor-ligand interactions more directly at substantially higher computational cost. Here, we present G-screen, a freely available and scalable receptor-aware virtual screening framework designed for cases in which a reference protein-ligand complex structure is available. Instead of performing full docking, G-screen rapidly aligns candidate ligands to the reference ligand using a flexible global alignment algorithm (G-align) and evaluates receptor-aware pharmacophore interactions derived from the reference complex, thereby combining the efficiency of ligand-based alignment with explicit atomic-level interaction analysis. Benchmarking on DUD-E, LIT-PCBA, and MUV datasets demonstrates that G-screen achieves competitive discrimination and early enrichment relative to representative ligand-based and docking-based methods, while maintaining millisecond-scale per-molecule runtimes under multi-threaded execution. These results position G-screen as a practical and scalable receptor-aware screening strategy for efficiently filtering large chemical libraries when a reference complex structure is available. Scientific ContributionWe have developed a scalable virtual screening framework for efficiently filtering ultra-large chemical libraries using a flexible global alignment algorithm combined with receptor-aware pharmacophore evaluations. Despite explicitly capturing atomic-level interactions, the screening process using this method is highly efficient, maintaining millisecond-scale per-molecule runtimes under parallel execution. It achieves competitive discrimination and early enrichment, successfully bridging the speed of ligand-based approaches with the structural context of traditional docking.
Yang, B.; Xu, Y.; Xiang, C.; Zhu, Y.; Li, T.; Sinitskiy, A.; Li, J.
Show abstract
Lead optimization plays an important role in preclinical drug discovery. While deep learning has accelerated this process, structure-based approaches that leverage 3D protein-ligand information remain underexplored. Existing models could improve predicted affinity but often yield synthetically inaccessible compounds, whereas screening-based methods limit chemical novelty by relying on fixed fragment libraries. To bridge the gap, we introduce Slogen--a Structure-based Lead Optimization algorithm unifying fragment Generation and screENing. To achieve this, Slogen integrates a transformer-based variational autoencoder, pretrained on the BindingNet v2 dataset, with an E(3)-equivariant graph neural network that models 3D protein-fragment interactions. This unified framework enables both fragment generation and similarity-based screening, simultaneously addressing synthetic tractability and structural diversity. Benchmarking study shows that Slogen matches or surpasses state-of-the-art methods while exploring broader chemical space. Case studies on the Smoothened and D1 dopamine receptors demonstrate its capacity to design high-affinity, drug-like molecules, providing a practical method for structure-guided lead optimization.
Kim, J.; Ryu, S.; Park, H.; Seok, C.
Show abstract
Designing drug-like molecules that satisfy multiple objectives--such as high binding affinity, synthesizability, and drug-likeness--poses a complex global optimization problem over an astronomically large chemical space. Existing deep learning-based molecular generative models often treat this task as distribution modeling, relying on atom-level autoregressive actions with less consideration of explicit optimization feedback. Consequently, they frequently generate invalid structures, converge to local optima, or produce synthetically infeasible candidates. Here, we introduce SHARP (Synthesizable Hierarchical Action-space Reinforcement learning for Pareto optimization), a molecular generator that addresses these limitations via a fragment-based hierarchical action space and reinforcement learning. SHARP ensures synthetic accessibility by applying action masks guided by a pretrained Synthesizability Estimation Model (SEM). The reinforcement learning (RL) policy is trained using a composite reward function integrating docking scores, pharmacophore matching, and solvent accessibility to generate functionally relevant and experimentally tractable molecules. Furthermore, across four lead optimization tasks--fragment growing, linker design, scaffold hopping, and sidechain decoration--on a diverse receptor set, SHARP consistently outperforms prior methods in producing molecules at high affinity and synthesizability. These results demonstrate that reinforcement learning with a chemically intuitive action space design can be an effective solution to the optimization challenges in AI-driven drug discovery, offering a robust framework for rational molecular design in structure-based applications.
Wang, J.; Dokholyan, N. V.
Show abstract
The design of molecules for flexible protein pockets represents a significant challenge in structure-based drug discovery, as proteins often undergo conformational changes upon ligand binding. While deep learning-based approaches have shown promise in molecular generation, they typically treat protein pockets as rigid structures, limiting their ability to capture the dynamic nature of protein-ligand interactions. Here, we introduce YuelDesign, a novel diffusion-based framework specifically developed to address this challenge. YuelDesign employs a new protein encoding scheme with a fully connected graph representation to encode protein pocket flexibility, a systematic denoising process that refines both atomic properties and coordinates, and a specialized bond reconstruction module tailored for de novo generated molecules. Our results demonstrate that YuelDesign generates molecules with favorable drug-likeness and low synthetic complexity. The generated molecules also exhibit diverse chemical functional groups, including some not even present in the training set. Redocking analysis reveals that the generated molecules exhibit docking energies comparable to native ligands. Additionally, a detailed analysis of the denoising process shows how the model systematically refines molecular structures through atom type transitions, bond dynamics, and conformational adjustments. Overall, YuelDesign presents a versatile framework for generating novel molecules tailored to flexible protein pockets, with promising implications for drug discovery applications.
Xie, E.; Hasegawa, K.; Kementzidis, G.; Papadopoulos, E.; Aktas, B. H.; Deng, Y.
Show abstract
Alzheimers Disease (AD) is a progressive neurodegenerative disorder that affects over 51 million individuals globally. The {beta}-secretase (BACE1) enzyme is responsible for the production of amyloid beta (A{beta}) plaques in the brain. The accumulation of A{beta} plaques leads to neuronal death and the impairment of cognitive abilities, both of which are fundamental symptoms of AD. Thus, BACE1 has emerged as a promising therapeutic target for AD. Previous BACE1 inhibitors have faced various issues related to molecular size and blood-brain barrier permeability, preventing any of them from maturing into FDA-approved AD drugs. In this work, a generative AI framework is developed as the first AI application to the de novo generation of BACE1 inhibitors. Through a simple, robust, and accurate molecular representation, a Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP), and a Genetic Algorithm (GA), the framework generates and optimizes over 1,000,000 candidate inhibitors that improve upon the bioactive and pharmacological properties of current BACE1 inhibitors. Then, the molecular docking simulation models the candidate inhibitors and identifies 14 candidate drugs that exhibit stronger binding interactions to the BACE1 active site than previous candidate BACE1 drugs from clinical trials. Overall, the framework successfully discovers BACE1 inhibitors and candidate AD drugs, accelerating the developmental process for a novel AD treatment.
Menezes, F.; Wahida, A.; Popowicz, G. M.
Show abstract
The modern AI models promise decoding of the genomic landscape that holds, in principle, the information required for rational therapeutic design. Genes encode proteins whose functions are mediated by their three-dimensional structures via bonded and non-bonded interactions. Since the late 1970s, the advent of macromolecular crystallography inspired the notion that structural knowledge alone could enable a "lock-and-key" approach to drug design. However, this framework has failed to catalyze a step-change in generating new drugs. Drug discovery continues to depend on costly, resource-intensive, and largely serendipitous screening campaigns that probe only an infinitesimal fraction of the drug-like chemical space. Despite some success cases, our understanding of, and reasoning from, non-bonded interaction chemistry is still too limited for generalized applicability. Furthermore, though structural databases contain hundreds of thousands of entries, a strong historical bias pervades protein-drug structures, hindering reliable advances through data science. Here, we show how a simple machine learning model successfully infers the principles of non-bonded chemical interactions in the drug-receptor space. A reductionist approach to the training data led to a model generalizing drug-target interactions, minimizing memorization frequently seen in large, structural models. We show how the model can infer complex interactions incomprehensible for classical physics-based force field models and approach a quantum level of understanding. The model was validated by retrospective and prospective, real-life problems in drug optimization. when targeting a challenging protein-protein interface. Our approach offers a simple, interpretable and explainable way to steer drug optimization and condition complex generative models to greatly accelerate, diversify and enhance drug discovery.
Kadukova, M.; Chupin, V.; Grudinin, S.
Show abstract
Virtual screening is an essential part of the modern drug design pipeline, which significantly accelerates the discovery of new drug candidates. Structure-based virtual screening involves ligand conformational sampling, which is often followed by re-scoring of docking poses. A great variety of scoring functions have been designed for this purpose. The advent of structural and affinity databases and the progress in machine-learning methods have recently boosted scoring function performance. Nonetheless, the most successful scoring functions are typically designed for specific tasks or systems. All-purpose scoring functions still perform poorly on the virtual screening tests, compared to precision with which they are able to predict co-crystal binding poses. Another limitation is the low interpretability of the heuristics being used. We analyzed scoring functions performance in the CASF benchmarks and discovered that the vast majority of them have a strong bias towards predicting larger binding interfaces. This motivated us to develop a physical model with additional entropic terms with the aim of penalizing such a preference. We parameterized the new model using affinity and structural data, solving a classification problem followed by regression. The new model, called Convex-PLR, demonstrated high-quality results on multiple tests and a substantial improvement over its predecessor Convex-PL. Convex-PLR can be used for molecular docking together with VinaCPL, our version of AutoDock Vina, with Convex-PL integrated as a scoring function. Convex-PLR, Convex-PL, and VinaCPL are available at https://team.inria.fr/nano-d/convex-pl/.
Karatzas, P.; Brotzakis, Z. F.; Sarimveis, H.
Show abstract
Partially disordered proteins can contain both stable and unstable secondary structure segments and are involved in various (mis)functions in the cell. The extensive conformational dynamics of partially disordered proteins scaling with extent of disorder and length of the protein hampers the efficiency of traditional experimental and in-silico structure-based drug discovery approaches. Therefore new efficient paradigms in drug discovery taking into account conformational ensembles of proteins need to emerge. In this study, using as a test case the AR-V7 transcription factor splicing variant related to prostate cancer, we present an automated methodology that can accelerate the screening of small molecule binders targeting partially disordered proteins. By swiftly identifying the conformational ensemble of AR-V7, and reducing the dimension of binding-sites by a factor of 90 by applying appropriate physicochemical filters, we combine physics based molecular docking and multi-objective classification machine learning models that speed up the screening of thousands of compounds targeting AR-V7 multiple binding sites. Our method not only identifies previously known binding sites of AR-V7, but also discovers new ones, as well as increases the multi-binding site hit-rate of small molecules by a factor of 10 compared to naive physics-based molecular docking.
Scantlebury, J.; Vost, L.; Carbery, A.; Hadfield, T. E.; Turnbull, O. M.; Brown, N.; Chenthamarakshan, V.; Das, P.; Grosjean, H.; von Delft, F.; Deane, C. M.
Show abstract
Over the last few years, many machine learning-based scoring functions for predicting the binding of small molecules to proteins have been developed. Their objective is to approximate the distribution which takes two molecules as input and outputs the energy of their interaction. Only a scoring function that accounts for the interatomic interactions involved in binding can accurately predict binding affinity on unseen molecules. However, many scoring functions make predictions based on dataset biases rather than an understanding of the physics of binding. These scoring functions perform well when tested on similar targets to those in the training set, but fail to generalise to dissimilar targets. To test what a machine learning-based scoring function has learnt, input attribution--a technique for learning which features are important to a model when making a prediction on a particular data point--can be applied. If a model successfully learns something beyond dataset biases, attribution should give insight into the important binding interactions that are taking place. We built a machine learning-based scoring function that aimed to avoid the influence of bias via thorough train and test dataset filtering, and show that it achieves comparable performance on the CASF-2016 benchmark to other leading methods. We then use the CASF-2016 test set to perform attribution, and find that the bonds identified as important by PointVS, unlike those extracted from other scoring functions, have a high correlation with those found by a distance-based interaction profiler. We then show that attribution can be used to extract important binding pharmacophores from a given protein target when supplied with a number of bound structures. We use this information to perform fragment elaboration, and see improvements in docking scores compared to using structural information from a traditional, data-based approach. This not only provides definitive proof that the scoring function has learnt to identify some important binding interactions, but also constitutes the first deep learning-based method for extracting structural information from a target for molecule design.
Manrique, P. D.; Leus, I.; Lopez, C. A.; Mehla, J.; Malloci, G.; Gervasoni, S.; Vargiu, A.; Kinthada, R.; Herndon, L.; Hengartner, N.; Walker, J. K.; Rybenkov, V.; Ruggerone, P.; Zgurskaya, H.; Gnanakaran, S.
Show abstract
The ability of Gram-negative pathogens to adapt and protect themselves against antibiotics is a growing threat to public health. The low permeability of the outer membrane (OM) in combination with effective multidrug efflux pumps, constitute the two main antibiotic resistance mechanisms. Though much efforts have been devoted to discover new antibiotics that can bypass these defense mechanisms, no new antibiotic classes have been introduced into clinics in the last 35 years. Models that identify specific descriptors of molecular properties and predict the likelihood that a given compound is capable of successfully permeate the OM and inhibit bacterial growth while avoiding efflux could facilitate the discovery of novel classes of antibiotics. Here we evaluate 174 molecular descriptors of 1260 antimicrobial compounds and study their correlations with antibacterial activity in Gram-negative Pseudomonas aeruginosa. While part of these descriptors are computed using traditional approaches based on the physicochemical properties intrinsic to the compounds, ensemble docking and all-atom molecular dynamics (MD) simulations are used to derive additional bacterium-specific mechanistic properties. Descriptors of compound permeation across the OM were calculated using all-atom MD simulations of the compounds in different subregions of the OM model. Descriptors of interactions with efflux pumps were calculated from ensemble docking of compounds targeting specific binding pockets of MexB, the major efflux transporter of P. aeruginosa. Using these descriptors and the measured antibacterial inhibitory concentrations of compounds, we design and implement a statistical protocol to identify a subset of the molecular properties that are predictive of whether a given compound is a strong or weak permeator across the Gram-negative OM. Our results indicate that 88.4% of the compounds that show measurable antibacterial activity, follow very consistent rules of permeation, which highlight the critical role that the interaction between the compound and the OM have at predicting permeation. The remaining 11.6% of the compounds, although less predictive, are characterized by distinctive structural markers that can be used to minimize classification errors. An implementation of the permeation rules and the structural markers uncovered in our study is shown, and it demonstrates the accuracy of our approach in a set of previously unseen compounds. Taken together, our analysis sheds new light on the key molecular properties that drug candidates should have in order to be effective at OM permeation/inhibition of P. aeruginosa, and opens the gate to similar data-driven studies in other Gram-negative pathogens.
Lope Perez, K.; Miranda Quintana, R. A.
Show abstract
BitBIRCH is a novel clustering algorithm that enables the analysis of extremely large molecular libraries; however, its performance can be hindered by an excessive number of singletons or the formation of disproportionately large clusters. Here, we present a data-driven strategy to identify optimal BitBIRCH parameters that mitigate these limitations. Using the ChEMBL34 library as a case study (with additional datasets reported in the Supporting Information), we show that similarity thresholds between three and four standard deviations above the global mean provide a balanced trade-off between cluster count and medoid similarity. These values are efficiently approximate with the iSIM and iSIM-sigma frameworks. For the branching factor, values as high as computationally feasible are recommended, as increasing it to 1024 substantially reduced the number of singletons. We further introduce an iterative re-clustering procedure wherein the similarity threshold can be adjusted to merge related subclusters and singletons from the initial clustering, providing user-defined control over the extent of cluster fusion. This work provides practical guidelines to enhance the robustness and usability of BitBIRCH for large-scale molecular clustering.
Kong, L.; Suh, D.; Im, W.
Show abstract
Covalent inhibitor research is an emerging topic in drug discovery due to its superior performance in specificity and inhibition effects. While molecular docking is a popular strategy in prediction and assessment of ligand conformations or poses in receptor proteins, covalent ligand docking requires nontrivial preparation efforts, as the ligand structure changes during the covalent complex formation. In order to facilitate molecular docking for covalent ligands, we have developed CHARMM-GUI Covalent Ligand Docker (CGUI-CLD), a new module for covalent ligand docking supported by AutoDock4. CGUI-CLD automates ligand preparation, supports ligand modification, implements docking simulation, and presents results through an intuitive user interface. A knowledge-based library built in CGUI-CLD currently supports 66 warheads and 8 amino acids, which can be used to automate the covalent ligand transformation from a pre-reaction to a post-reaction adduct form seamlessly. Moreover, CHARMM-GUI High-Throughput Simulator is integrated for rapid generation of multiple molecular dynamics simulation systems. CGUI-CLD is expected to significantly reduce a massive workload of covalent ligand docking and advance covalent ligand research. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/738313v1_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@1ccfefforg.highwire.dtl.DTLVardef@1794047org.highwire.dtl.DTLVardef@16b0368org.highwire.dtl.DTLVardef@acbc3e_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOAbstract TOCC_FLOATNO C_FIG
Sayyah, E.; Kurul, E.; Tunc, H.; DURDAGI, S.
Show abstract
Molecular representation determines which aspects of chemical structure can be learned, compared, and interpreted in computational drug discovery. Existing encodings typically emphasize either compact string description, as in SMILES and SELFIES, or efficient similarity search, as in circular fingerprints, but they may not simultaneously provide deterministic sequence structure, graph-level interpretability, pharmacophore annotation, and high-fidelity molecular reconstruction. Here, we introduce MolCodon, a codon-based molecular language that represents small molecules as deterministic sequences of fixed-width three-character tokens over a five-symbol alphabet, C, N, O, S, and X. Inspired by the triplet organization of the genetic code, MolCodon assigns chemically defined codon families to atoms, bonds, ring and branch topology, fused-ring references, pharmacophore features, bond mobility, charge, and stereochemistry. A deterministic graph traversal with ring-contiguity preservation produces sequences in which chemically meaningful substructures remain locally organized and traceable to the underlying molecular graph. Across around 2,9 million molecules from six commercial screening libraries, MolCodon achieved 98.93% InChIKey-level round-trip fidelity, supporting its use as a high-fidelity sequence representation for drug-like chemistry. MolCodon-derived sparse sequence and trace features further outperformed SELFIES and Group SELFIES across ten QSAR tasks and exceeded classical fingerprint baselines in six out of ten tasks. As an application of the representation, MolCodon BLAST similarity engine decomposes molecular similarity into ring topology, branch context, attachment architecture, and pharmacophore correspondence, enabling interpretable scaffold-hopping searches. In a PARP1 virtual screening study, MolCodon retrieved scaffold-diverse candidates to a known PARP-1 inhibitor Olaparib. Together, these results establish MolCodon as a new molecular representation paradigm that transforms chemical graphs into high-fidelity, interpretable, and alignment-compatible codon sequences, opening a direct path for bioinformatics-inspired analysis of small-molecule chemical space. The MolCodon encoder, decoder, and BLAST similarity engine are freely available as open-source software at https://github.com/DurdagiLab/MolCodon
Sayyah, E.; Ulug, M. E.; Tunc, H.; DURDAGI, S.
Show abstract
Structure-based docking often assumes a rigid receptor, obscuring induced-fit side-chain rearrangements that govern affinity and selectivity. We introduce DRGSCROLL, an open access docking platform and web server that jointly optimizes ligand pose and continuous receptor side-chain {chi} angles within a genetic-algorithm (GA) optimizer. DRGSCROLL seeds broad {chi}-angle populations, evaluates candidates with a dual-objective fitness that rewards low interaction energy profile while penalizing steric clashes, and uses per-residue crossover plus stochastic mutation to maintain physically realistic {chi} sets. To favor exploration over premature local minima, per-iteration minimization is deliberately omitted; final poses are selected by clash-free filtering and RMSD-based clustering. Across PDBbind protein-ligand complexes, DRGSCROLL showed generation-wise decreases in clash counts and improved docking scores, indicating convergence to sterically viable, low-energy pocket conformations rarely accessed by rigid-receptor protocols. Relative to known publicly available and commercial docking programs such as Vina, Glide, and IFD; DRGSCROLL sampled more favorable energy distributions with lower median scores, consistent with superior induced-fit capture. In prospective virtual-screening evaluations, DRGSCROLL enhanced actives versus inactives discrimination for RET kinase and PARP1 targets, while maintaining balanced precision-recall trade-offs. Additionally, DRGSCROLLs endurance in differentiating actives from property-matched decoys was validated by benchmarking against a decoy database (Directory of Useful Decoys, DUDE), which showed improved AUC, recall, and enrichment performance in comparison to Vina GPU 2.1 under the same screening settings. By embedding continuous side-chain flexibility directly into the search, rather than relying on post hoc rotamer tweaks or heavy local minimization, DRGSCROLL addresses key combinatorial and feasibility bottlenecks in flexible docking. The method provides a scalable, physics-grounded route to adaptive receptor-ligand modeling that improves pose accuracy and early enrichment for flexible targets, from fragment screening to triage of synthetically ready or AI-generated small molecule libraries. Availability: platform page, https://www.drgscroll.com/; academic/nonprofit web server, https://drgscroll.bau.edu.tr/