Back

Knowledge-guided machine-learning and reverse screening combined method to predict cancer cell line responses to cytotoxic molecules

Cuozzo, A.; Gomes, M. A. S. P.; Daina, A.; Zoete, V.

2025-12-19 bioinformatics
10.64898/2025.12.16.694721 bioRxiv
Show abstract

Estimating the cell line targets of cytotoxic small molecules is important for drug discovery and central for targeted therapy in oncology. Accurate prediction of sensitive cell lines enables early identification of efficacy and toxicity, optimization of drug selectivity, and can foster drug repurposing. While most bioactive compounds interact with multiple macromolecular targets, the cytotoxicity encompasses diverse complex biological and chemical mechanisms that could even not all be related to binding to macromolecules, making the prediction of cytotoxicity specificity particularly challenging. To address early-phase prediction of cancer cell line targets of cytotoxic compounds, we developed a method combining a machine-learning classification model with a ligand-based reverse screening procedure able to rank cell-line from the most probable to the least probable target of any cytotoxic molecule. The development focused on addressing the challenges related to the scarcity of available experimental data on non-cytotoxic compounds. A knowledge-guided generation of realistic alleged inactives allowed to train several binary logistic regression models. The most robust classification model was trained on 164,134 cytotoxic compounds extracted from ChEMBL to generate a score of predicted sensitivity of cell lines for any new cytotoxic molecule. The method demonstrated strong predictive ability, recovering at least one experimental target within the 15 most probable cell-lines for 71% of nearly 11,000 external cytotoxic compounds tested across 1018 cancer cell lines.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.