Back

Harnessing Chemical Space Neural Networks to Systematically Annotate GPCR ligands

Hansson, F. G.; Madsen, N. G.; Hansen, L. G.; Jakociunas, T.; Lengger, B.; Keasling, J. D.; Jensen, M. K.; Acevedo-Rocha, C. G.; Jensen, E. D.

2024-04-01 bioinformatics
10.1101/2024.03.29.586957 bioRxiv
Show abstract

Machine learning (ML) has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors (hGPCRs) in FDA-approved drugs, exhaustive in-distribution drug-target interaction (DTI) testing across all pairs of hGPCRs and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution (OOD) exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks (CSNN) that leverages network homophily and training-free graph neural networks (GNNs) with Labels as Features (LaF). We show that CSNNs ability to make accurate predictions strongly correlates with network homophily. Thus, LaFs strongly increase a ML models capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 DTIs, 539 compounds, 7 hGPCRs) to discover novel DTIs for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.