Back

Learning a CoNCISE language for small-molecule binding

Erden, M.; Devkota, K.; Varghese, L.; Cowen, L.; Singh, R.

2025-01-13 bioinformatics
10.1101/2025.01.08.632039 bioRxiv
Show abstract

Rapid advances in deep learning have improved in silico methods for drug-target interaction (DTI) prediction. However, current methods do not scale to the massive catalogs that list millions or billions of commercially-available small molecules. Here, we introduce CoNCISE, a method that accelerates drug-target interaction (DTI) prediction by 2-3 orders of magnitude while maintaining high accuracy. CoNCISE uses a novel vector-quantized codebook approach and a residual-learning based training of hierarchical codes. Strikingly, we find that much of binding-specificity information in the small molecule space can be compressed into just 15 bits of information per compound, characterizing all small molecules into 32,768 hierarchically-organized binding categories. Our DTI architecture, which combines these compact ligand representations with fixed-length protein embeddings in a cross-attention framework, achieves state-of-the-art prediction accuracy at unprecedented speed. We demonstrate CoNCISEs practical utility by indexing 6.4 billion ligands in the Enamine dataset, enabling researchers to query vast chemical libraries against a protein target in seconds. A "CoNCISE + docking" pipeline screened Enamine to propose strong binders (predicted KD {approx} 10-20 {micro}M) of three difficult-to-drug targets, each within two hours. CoNCISEs advance could democratize access to largescale computational drug discovery, potentially enabling rapid identification of promising molecules for therapeutic targets and cellular perturbations.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.