MolJam: A Multidimensional Framework for Assessing Molecular Dataset Quality and Its Impact on Machine Learning
Wang, P.; Shi, Z.; Gao, X.; Zhou, R.
Show abstract
High-quality molecular datasets are essential for reliable machine learning in cheminformatics and bioinformatics, yet dataset quality is rarely assessed systematically and its relationship with downstream model performance remains poorly understood. Here, we present MolJam, an open-source framework for quantitative assessment of molecular dataset quality across five dimensions-structural integrity, data quality, experimental information quality, chemical space coverage, and data distribution-using 12 standardized metrics. Application of MolJam to 11 MoleculeNet and eight ChEMBL-derived datasets revealed widespread and heterogeneous quality issues, including undefined stereochemistry in up to 70.72% of molecules, inconsistent molecular representations, and contradictory labels. We next asked whether improving these quality metrics necessarily improves machine learning performance. Refinement of the ESOL and Lipophilicity datasets increased their MolJam quality scores but produced mixed effects on predictive performance, suggesting a competing influence of reduced dataset size. Controlled ablation experiments further demonstrated that both dataset quality and data quantity contribute to model performance and, notably, that retaining molecules with incomplete stereochemical information can outperform their removal when the resulting gain in data quantity offsets the quality penalty. Thus, molecular dataset curation cannot be reduced to maximizing data cleanliness alone but requires balancing multiple dimensions of data quality against information loss. MolJam provides a standardized framework for diagnosing molecular dataset limitations, comparing benchmark quality, and quantitatively evaluating how data curation decisions influence downstream machine learning.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity 94%
- Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data 94%
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 93%
Similar papers in this journal
- Revolutionizing GPCR-Ligand Predictions: DeepGPCR with experimental Validation for High-Precision Drug Discovery 93%
- Hybrid Deep Learning with Protein Language Models and Dual-Path Architecture for Predicting IDP Functions 93%
- Data-driven strategies for drug repurposing: insights, recommendations, and case studies 92%
Similar papers in this journal
- Repurposing Therapeutics for COVID-19: Rapid Prediction of Commercially available drugs through Machine Learning and Docking 93%
- TargetDB: A target information aggregation tool and tractability predictor 92%
- Distinguishing classes of neuroactive drugs based on computational physicochemical properties and experimental phenotypic profiling in planarians 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.