Chemical Dice Integrator (CDI): A Scalable Framework for Multimodal Molecular Representation Learning
Ahuja, G.; Kumar, S.; Solanki, S.; Gupta, M.; Mohanty, S. K.; Satija, S.; Chauhan, S.; Duari, S.; Sharma, A.; Gautam, V.; Arora, S.; Shome, R.; Sinha, S.; Sharma, A. K.; Mittal, A.; Sengupta, D.; Murugan, N. A.
Show abstract
The machine learning landscape for molecular property prediction is fragmented, with numerous Featurizers each capturing a narrow, specialized view of chemical structure. This heterogeneity forces a suboptimal choice of representation a priori, limiting model generalizability. We introduce the Chemical Dice Integrator (CDI), a hierarchical framework that unifies six orthogonal molecular representations, physicochemical (Mordred), topological (GROVER), visual (ImageMol), biological (Signaturizer), quantum-mechanical (MOPAC), and linguistic (ChemBERTa), into a single, coherent embedding. The framework consists of CDI-Basic, a two-tiered autoencoder that fuses these modalities, and CDI-Generalised, a Mamba State-Space Model (SSM) that learns a direct, efficient map from SMILES strings to the unified embedding space. Extensive benchmarking across 23 classification (171 tasks) and 10 regression datasets demonstrates that CDI embeddings consistently achieve superior predictive performance compared to individual Featurizers and standard feature aggregation methods. The CDI-Generalised model achieves this performance with exceptional computational efficiency, outperforming deep learning Featurizers in terms of speed and resource overhead. Furthermore, we demonstrate that the CDI embedding is chemically intuitive, allowing for the sensitive distinction of nuanced structural variants, such as chiral enantiomers and kekulized SMILES forms. By bridging multimodal chemical intelligence with scalable, sequence-based inference, CDI offers a strong foundation for molecular machine learning.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Adding stochastic negative examples into machine learning improves molecular bioactivity prediction 97%
- Learning Binding Affinities via Fine-tuning of Protein and Ligand Language Models 96%
- BOLD-GPCRs: A Transformer-Powered App for Predicting Ligand Bioactivity and Mutational Effects Across Class A GPCRs 96%
Similar papers in this journal
- EVOSYNTH: Enabling Multi-Target Drug Discovery through Latent Evolutionary Optimization and Synthesis-Aware Prioritization 96%
- A network medicine framework for multi-modal data integration in therapeutic target discovery 92%
- Deep generative modeling of temperature-dependent structural ensembles of proteins 91%
Similar papers in this journal
- ProT-Diff: A Modularized and Efficient Approach to De Novo Generation of Antimicrobial Peptide Sequences through Integration of Protein Language Model and Diffusion Model 93%
- A Multi-Property Optimizing Generative Adversarial Network for de novo Antimicrobial Peptide Design 92%
- Interpretable PROTAC degradation prediction with structure-informed deep ternary attention framework 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.