iSIM-sigma: efficient standard deviation calculation for molecular similarity
Perez, K. L.; Zhao, B.; Quintana, R. A. M.
Show abstract
AbstractThe average and variance of the molecular similarities in a set is high-value and useful information for cheminformatics tasks like chemical space exploration and subset selection. However, the calculation of the variance of the complete similarity matrix has a quadratic complexity, O(N2). As the sizes of molecular libraries constantly increase, this pairwise approach is unfeasible. In this work, we present an alternative to obtaining the exact standard deviation of the molecular similarities in a set (with N molecules and M features) for the Russell-Rao (RR) and Sokal-Michener (SM) similarity indexes in O(N M2) complexity. Additionally, we present a highly accurate approximation with linear complexity, O(N), based on the sampling of representative molecules from the set. The proposed approximation can be extended to other similarity indexes, including the popular Jaccard-Tanimoto (JT). With only the sampling of 50 molecules, the proposed method can estimate the standard deviation of the similarities in a set with RMSE lower than 0.01 for sets of up to 50,000 molecules. In comparison, random sampling does not warrant a good approximation as shown in our results.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Data Imbalance in Drug Response Prediction - Multi-Objective Optimization Approach in Deep Learning Setting 95%
- Interpretable and Generalizable Attention-Based Model for Predicting Drug-Target Interaction Using 3D Structure of Protein Binding Sites: SARS-CoV-2 Case Study and in-Lab Validation 94%
- A New Paradigm for Applying Deep Learning to Protein-Ligand Interaction Prediction 94%
Similar papers in this journal
- BinderSpace: A Package for Sequence Space Analyses for Datasets of Affinity-Selected Oligonucleotides and Peptide-Based Molecules 95%
- Enhanced conformational exploration of protein loops using a global parameterization of the backbone geometry 92%
- Investigating Protein-Protein Allosteric Network using Current-Flow Scheme 92%
Similar papers in this journal
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 95%
- PL-PatchSurfer3: Improved Structure-Based Virtual Screening for Structure Variation Using 3D Zernike Descriptors 94%
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.