F-CPI: Prediction of activity changes induced by fluorine substitution using multimodal deep learning
Zhang, Q.; Yin, W.; Chen, X.; Zhou, A.; Zhang, G.; Zhao, Z.; Li, Z.; Zhang, Y.; Shen, J.; Zhu, W.; Jiang, X.; Xu, Z.
Show abstract
There are a large number of fluorine (F)-containing compounds in approved drugs, and F substitution is a common method in drug discovery and development. However, F is difficult to form traditional hydrogen bonds and typical halogen bonds. As a result, accurate prediction of the activity after F substitution is still impossible using traditional drug design methods, whereas artificial intelligence driven activity prediction might offer a solution. Although more and more machine learning and deep learning models are being applied, there is currently no model specifically designed to study the effect of F on bioactivities. In this study, we developed a specialized deep learning model, F-CPI, to predict the effect of introducing F on drug activity, and tested its performance on a carefully constructed dataset. Comparison with traditional machine learning models and popular CPI task models demonstrated the superiority and necessity of F-CPI, achieving an accuracy of approximately 89% and a precision of approximately 67%. In the end, we utilized F-CPI for the structural optimization of hit compounds against SARS-CoV-2 3CLpro. Impressively, in one case, the introduction of only one F atom resulted in a more than 100-fold increase in activity (IC50: 22.99 nM vs. 28190 nM). Therefore, we believe that F-CPI is a helpful and effective tool in the context of drug discovery and design.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep learning integration of molecular and interactome data for protein-compound interaction prediction 98%
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 98%
- AdapToR: Adaptive Topological Regression for quantitative structure-activity relationship modeling 96%
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 95%
- In Silico Identification of Potential Inhibitors of Mycobacterium tuberculosis DNA Gyrase from Phytoconstituents of Indian Medicinal Plants 94%
- Combining Multi-Dimensional Molecular Fingerprints to Predict hERG Cardiotoxicity of Compounds 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.