Back

Explainable AI Predicts Hematoxicity from Cancer Treatment Using Multimodal Real-World Data

Keyl, J.; Keyl, P.; Lenfers, T.; Hosch, R.; Kiermeyer, N.; Schallenberg, S.; Kim, M.; Bauer, S.; Bechrakis, N.; Forsting, M.; Fuehrer-Sakel, D.; Kebir, S.; Gruenwald, V.; Hadaschik, B.; Haubold, J.; Herrmann, K.; Kasper, S.; Kimmig, R.; Lang, S.; Rassaf, T.; Roesch, A.; Schadendorf, D.; Siveke, J. T.; Stuschke, M.; Sure, U.; Totzeck, M.; Welt, A.; Wiesweg, M.; Egger, J.; Hartmann, S.; Montavon, G.; Nensa, F.; Mueller, K.-R.; Schuler, M.; Kleesiek, J.; Klauschen, F.

2026-04-30 oncology
10.64898/2026.04.29.26352032 medRxiv
Show abstract

Adverse drug effects remain a major barrier to safe and effective cancer therapy, underscoring the need for tools that predict treatment-related toxicities. We analyzed multimodal real-world data from 14,596 cancer patients across 38 cancer entities, encompassing 330 clinical, tumor, and imaging characteristics, along with 89 anticancer agents. Hematological adverse events (HAE), defined by nadirs of hemoglobin, leukocyte, neutrophil, and platelet values within two months of treatment initiation, were highly prevalent (87.7%; 33.1% severe). We developed Toxix, an explainable artificial intelligence (xAI) framework modeling interactions between patient characteristics and drug combinations. Toxix achieved strong predictive performance for severe toxicities (median AUROC 0.85 for anemia; >0.76 for leukopenia, neutropenia, and thrombocytopenia) and was validated in an external cohort of 2,768 patients with non-small cell lung cancer. Model explainability enabled systematic characterization of drug-patient interactions underlying HAEs. Toxix provides a real-world informed framework for personalized and toxicity-aware cancer therapy planning.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.