MolLM: A Unified Language Model to Integrate Biomedical Text with 2D and 3D Molecular Representations
Tang, X.; Tran, A.; Tan, J.; Gerstein, M.
Show abstract
MotivationThe current paradigm of deep learning models for the joint representation of molecules and text primarily relies on 1D or 2D molecular formats, neglecting significant 3D structural information that offers valuable physical insight. This narrow focus inhibits the models versatility and adaptability across a wide range of modalities. Conversely, the limited research focusing on explicit 3D representation tends to overlook textual data within the biomedical domain. ResultsWe present a unified pre-trained language model, MolLM, that concurrently captures 2D and 3D molecular information alongside biomedical text. MolLM consists of a text Transformer encoder and a molecular Transformer encoder, designed to encode both 2D and 3D molecular structures. To support MolLMs self-supervised pre-training, we constructed 160K molecule-text pairings. Employing contrastive learning as a supervisory signal for cross-modal information learning, MolLM demonstrates robust molecular representation capabilities across 4 downstream tasks, including cross-modality molecule and text matching, property prediction, captioning, and text-prompted molecular editing. Through ablation, we demonstrate that the inclusion of explicit 3D representations improves performance in these downstream tasks. Availability and implementationOur code, data, and pre-trained model weights are all available at https://github.com/gersteinlab/MolLM.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cross-Modality and Self-Supervised Protein Embedding for Compound-Protein Affinity and Contact Prediction 97%
- DTI-Voodoo: machine learning over interaction networks and ontology-based background knowledge predicts drug-target interactions 96%
- GraphDTA: Predicting drug-target binding affinity with graph neural networks 96%
Similar papers in this journal
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 96%
- All-Atom Protein Sequence Design using Discrete Diffusion Models 95%
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.