Rethinking Representation Complexity in Drug-Target Prediction via Supervised Vector Quantization
Chen, J.; Zhang, Y.-z.; Lu, L.; Wu, M.; Chen, Z.; Imoto, S.; Li, C.
Show abstract
Accurate prediction of drug-target interactions (DTIs) is crucial for computational drug discovery. Although pretrained language models offer richer molecular and protein representations, their increasing complexity does not always lead to better predictive performance. In many cases, the inclusion of redundant or irrelevant features may obscure biologically relevant patterns. In this study, we systematically evaluate the contribution of complex features in DTI prediction and demonstrate that only a portion of these features is truly informative. Based on this insight, we propose a Vector Quantization (VQ)-based module that functions as a plug-and-play feature selection layer within deep learning architectures. When combined with a simple fully connected classifier, this supervised VQ (SVQ) framework not only surpasses recent state-of-the-art DTI methods in performance, but also enhances interpretability through the learning of discriminative codewords. This work highlights the importance of input feature selection in deep learning and offers a new perspective for constructing robust and interpretable DTI prediction models.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 97%
- GexMolGen: Cross-modal Generation of Hit-like Molecules via Large Language Model Encoding of Gene Expression Signatures 97%
- Interpretable and Generalizable Attention-Based Model for Predicting Drug-Target Interaction Using 3D Structure of Protein Binding Sites: SARS-CoV-2 Case Study and in-Lab Validation 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.