Learning a CoNCISE language for small-molecule binding
Erden, M.; Devkota, K.; Varghese, L.; Cowen, L.; Singh, R.
Show abstract
Rapid advances in deep learning have improved in silico methods for drug-target interaction (DTI) prediction. However, current methods do not scale to the massive catalogs that list millions or billions of commercially-available small molecules. Here, we introduce CoNCISE, a method that accelerates drug-target interaction (DTI) prediction by 2-3 orders of magnitude while maintaining high accuracy. CoNCISE uses a novel vector-quantized codebook approach and a residual-learning based training of hierarchical codes. Strikingly, we find that much of binding-specificity information in the small molecule space can be compressed into just 15 bits of information per compound, characterizing all small molecules into 32,768 hierarchically-organized binding categories. Our DTI architecture, which combines these compact ligand representations with fixed-length protein embeddings in a cross-attention framework, achieves state-of-the-art prediction accuracy at unprecedented speed. We demonstrate CoNCISEs practical utility by indexing 6.4 billion ligands in the Enamine dataset, enabling researchers to query vast chemical libraries against a protein target in seconds. A "CoNCISE + docking" pipeline screened Enamine to propose strong binders (predicted KD {approx} 10-20 {micro}M) of three difficult-to-drug targets, each within two hours. CoNCISEs advance could democratize access to largescale computational drug discovery, potentially enabling rapid identification of promising molecules for therapeutic targets and cellular perturbations.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model 96%
- Learning with uncertainty for biological discovery and design 96%
- DynamicGT: a dynamic-aware geometric transformer model to predict protein binding interfaces in flexible and disordered regions 95%
Similar papers in this journal
- Tokenized and Continuous Embedding Compressions of Protein Sequence and Structure 96%
- Hierarchical affinity landscape navigation through learning a shared pocket-ligand space 96%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.