Back

Machine Learning-Driven Drug Repurposing for KRAS G12C and KRAS G12D Inhibition

Fuschi, G.; Germain, J. S.; Bebensee, D.; Moawad, C.; Aladysheva, A.; Mohamed, A.; Elwakeel, E.; Brooks, B. R.; Amin, M.

2025-05-20 cancer biology
10.1101/2025.05.16.654410 bioRxiv
Show abstract

KRAS is a predominant oncogenic driver across multiple cancers, long considered "undruggable" due to its high nucleotide affinity and lack of classical binding pockets. Although recent advances have led to covalent inhibitors like Sotorasib and Adagrasib for the KRAS G12C mutation, effective therapies for other common variants--most notably KRAS G12D, which is highly prevalent in aggressive pancreatic cancers--remain limited. In this study, we employ machine learning to identify potential inhibitors for both KRAS G12D and G12C by screening FDA-approved compounds from the ChEMBL database. Random Forest and Neural Network models were trained on binding affinity data from three BindingDB datasets: wild-type KRAS GTPase, KRAS G12C, and KRAS G12D. Our models identified high-affinity candidates including anti-cancer kinase inhibitors (e.g., Cobimetinib, Gilteritinib) as well as drugs from other categories (e.g., Bromocriptine, Cefepime). By incorporating atomic hybridization as a feature, we aim to capture the effects of induced polarization, potentially improving prediction accuracy. Despite modest agreement across models, our approach highlights several promising candidates--such as Acalabrutinib, which has prior evidence of G12C activity--for further investigation. These results offer a data-driven foundation for experimental validation and the future development of targeted KRAS therapies.

Published in ACS Omega (predicted rank #11) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.