Back

Ovarian cancer recurrence prediction: comparing confirmatory to real world predictors with machine learning

Katsimpokis, D.; van Odenhoven, A. E. C.; van Erp, M. A. J. M.; Wenzel, H. H. B.; van der Aa, M. A.; van Swieten, M. M. H.; Smedts, H. P. M.; Piek, J. M. J.

2025-03-12 epidemiology
10.1101/2025.03.04.25321571 medRxiv
Show abstract

IntroductionOvarian cancer is one of the deadliest cancers in women, with a 5-year survival rate of 17-28% in advanced stage (FIGO IIB-IV) disease and is often diagnosed at advanced stage. Machine learning (ML) has the potential to provide a better survival prognosis than traditional tools, and to shed further light on predictive factors. This study focuses on advanced stage ovarian cancer and contrasts expert-derived predictive factors with data-driven ones from the Netherlands Cancer Registry (NCR) to predict progression-free survival. MethodsA Delphi questionnaire was conducted to identify fourteen predictive factors which were included in the final analysis. ML models (regularized Cox regression, Random Survival Forests and XGBoost) were used to compare the Delphi expert-based set of variables to a real-world data (RWD) variable set derived from the NCR. A traditional, non-regularized, Cox model was used as the benchmark. ResultsWhile regularized Cox regression models with the RWD variable set outperformed the traditional Cox regression with the Delphi variables (c-index: 0.70 vs. 0.64 respectively), the XGBoost model showed the best performance overall (c-index: 0.75). The most predictive factors for recurrence were treatment types and outcomes as well as socioeconomic status, which were not identified as such by the Delphi questionnaire. ConclusionOur results highlight that ML algorithms have higher predictive power compared to the traditional Cox regression. Moreover, RWD from a cancer registry identified more predictive variables than a panel of experts. Overall, these results have important implications for AI-assisted clinical prognosis and provide insight into the differences between AI-driven and expert-based decision-making in survival prediction.

Published in ESMO Real World Data and Digital Oncology · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.