Back

DrugLM: A Unified Framework to Enhance Drug-Target Interaction Predictions by Incorporating Textual Embeddings via Language Models

Li, T.; Fang, Z.; Zhang, X.; Tang, K.; Chen, H.; Jiang, Z.; Zhao, T.; Xu, R.; Cheng, F.; Li, X.; Li, J.

2025-07-11 bioinformatics
10.1101/2025.07.09.657250 bioRxiv
Show abstract

MotivationAccurate prediction of drug-target interactions (DTIs) is central to computational drug discovery, offering the potential to reduce experimental costs and accelerate development timelines. While existing deep learning approaches such as Graph Neural Networks and Transformers have shown promise, they often overlook the rich semantic information embedded in textual descriptions of drugs and targets. These descriptions encode critical biomedical knowledge, including mechanisms of action, biological pathways involved, and therapeutic effects of drugs, which can enhance DTI prediction performance. ResultsWe introduce DrugLM, a unified framework that integrates embeddings derived from large language models (LLMs) into DTI-specific model architectures. DrugLM leverages textual descriptions of drugs and targets to generate semantic embeddings using a range of pretrained LLMs. These embeddings can be seamlessly incorporated into existing DTI models. We systematically evaluate multiple LLMs on benchmark DTI datasets and demonstrate strong performance even without fine-tuning. Moreover, supervised parameter-efficient fine-tuning of the LLMs further improves embedding quality, leading to enhanced prediction accuracy. Notably, a simple multilayer perceptron (MLP) using only LLM-derived embeddings surpasses several established DTI methods, underscoring the power of semantic features. Our findings highlight the practical value of integrating LLMs into DTI pipelines and offer a straightforward recipe for improved drug discovery: LLM embeddings of drugs and targets are both effective and easy to use. AvailabilityOur code and dataset are available at https://github.com/ShPhoebus/DrugLM

Published in BioMedInformatics · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.