Back

Digital Twin Approaches for Interpretable Side Effect Prediction in Drug Discovery

Ecker, A.; Szabo, G.; Szalma, J.; Ficho, E.; Reguly, I.; Csikasz-Nagy, A.

2025-10-15 pharmacology and toxicology
10.1101/2025.10.14.682276 bioRxiv
Show abstract

Artificial intelligence plays an ever-greater role in preclinical drug development, ranging from target identification and molecule design to ADME-Tox prediction; however, predicting side effects before performing clinical trials is still lagging behind. The best performing side effect predictors in the literature use either ATC codes, which are expert-derived features not even available at early stages, or graph neural networks based on chemical similarity, which - although use readily available features - are "black boxes" that do not deliver actionable insights. We argue that a paradigm shift is needed. Instead of using the latest neural network architectures that have proved worthy in other domains with a plethora of available data, one could use the off-targets of the compounds and build simple and interpretable predictors of side effects. To add another layer of biological realism, intricate biophysical mechanisms within the cells could also be simulated and used as features for training. Although not outperforming the current methods by a great margin, this digital twin-based model has the benefit of being interpretable, i.e., it puts biology behind the predictions. We showcase, with real-world examples, how the side effects predicted by this model can be interpreted and traced back to off-target proteins, and the complexes and signaling pathways in which they partake. In this way, the proposed model not only provides actionable insights, but in the future, may contribute to the amendment of secondary pharmacology assays. HighlightsO_LINo standard tool is available to predict side effects in early-phase drug discovery. C_LIO_LIAs publicly available side effect data is scarce, only simple models should be trained. C_LIO_LISimple models trained on biorealistic features, such as off-target proteins, are interpretable. C_LIO_LIInterpreting predictions can highlight critical off-targets and are therefore actionable. C_LI

Published in Drug Discovery Today · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.