Mapping Chemical Diversity: Descriptor-Guided Clustering of Natural Products in the COCONUT Database
Shreyasree, G.; Dileep, A.; Namani, A.; Karunakar, P.
Show abstract
Natural products represent a major source of bioactive compounds for drug discovery, yet their exploration remains challenging due to extensive structural complexity and scaffold diversity. Using the COCONUT database, we developed a cluster-oriented framework to systematically map and characterize the natural product chemical space through feature engineering, molecular clustering, and representative-based analysis. Descriptor selection identified a greedy maximum coverage strategy with a 0.35-0.85 correlation threshold range and 20 descriptors as the optimal feature set, enriched in physicochemical and graph-topological properties. Comparative evaluation of clustering approaches identified UMAP-HDBSCAN as the best-performing pipeline, generating 1,683 clusters with silhouette scores of 0.42 before and 0.24 after noise reassignment. Cluster profiling revealed a highly heterogeneous scaffold landscape, with 67.56% of clusters exhibiting low scaffold dominance and only 15.21% representing highly scaffold-dominated regions, supporting a chemical space composed largely of interconnected transitional clusters. Descriptor analyses showed that natural product clusters were generally enriched in saturated, low-aromaticity chemotypes with moderate lipophilicity and constrained molecular flexibility. Representative-based analyses demonstrated that central representatives (medoid and centroid-closest molecules) closely captured cluster-average properties, whereas diverse representatives better reflected structural breadth, findings further supported through descriptor-based and docking-based validation. Collectively, the results reinforce the natural product chemical space as a continuous yet structured manifold and provide a representative-guided framework for its efficient exploration in drug discovery applications. The complete data can be accessed at: https://github.com/shrek-28/DescriptorClusteringNPSpace
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- EVOSYNTH: Enabling Multi-Target Drug Discovery through Latent Evolutionary Optimization and Synthesis-Aware Prioritization 93%
- Using macromolecular electron densities to improve the enrichment of active compounds in virtual screening 93%
- A small molecule enhances arrestin-3 binding to the β2-adrenergic receptor 91%
Similar papers in this journal
- Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo) 93%
- G-PLIP: Knowledge graph neural network for structure-free protein-ligand bioactivity prediction 93%
- Intense bitterness of molecules: machine learning for expediting drug discovery 92%
Similar papers in this journal
- ProT-Diff: A Modularized and Efficient Approach to De Novo Generation of Antimicrobial Peptide Sequences through Integration of Protein Language Model and Diffusion Model 91%
- Transfer Learning and Permutation-Invariance improving Predicting Genome-wide, Cell-Specific and Directional Interventions Effects of Complex Systems 91%
- Mechanism-driven screening of membrane-targeting and pore-forming antimicrobial peptides 90%
Similar papers in this journal
- Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data 95%
- BitterMatch: Recommendation systems for matching molecules with bitter taste receptors 94%
- Stereochemically-aware bioactivity descriptors for uncharacterized chemical compounds 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.