Back

A Generative Patient Digital Twin for Sequential Treatment of Oropharyngeal Squamous Carcinomas

Yu, Z.; Hirpara, H. J.; Zhang, X.; Canahuate, G.; Wentzel, A.; Wang, Y.; Floricel, C.; Attia, S. K.; Mohamed, A. S. R.; Naser, M. A.; Fuller, C. D.; Marai, G. E.

2025-10-13 health informatics
10.1101/2025.10.10.25337766 medRxiv
Show abstract

Real-world testing of treatment strategies is often infeasible, emphasizing the need for robust simulation frameworks that can model diverse patient characteristics and predict treatment outcomes. In this study, we present a generative simulator designed to synthesize patient profiles and forecast treatment results under hypothetical scenarios,with the goal of facilitating personalized treatment planning with what-if scenarios. The proposed simulator is able to predict disease progression and treatment outcomes based on synthesized profiles through application of Variational Autoencoder and XGBoost models. Simulations evaluated its ability of generating realistic baseline patient profiles, and the predictive accuracy of the combined framework. The proposed method outperformed rule-based approaches and multilayer perceptron models in predicting 22 out of 25 clinical variables, with performance measured by F1 scores for categorical variables and Mean Squared Error for numerical variables. Case studies of two patients drawn from ground truth data illustrate that the simulator framework can represent both short treatment courses with early relapse and prolonged multi-modal trajectories with recurrent disease. These results underscore the frameworks ability to capture complex relationships in clinical data and highlight its advantages over baseline methods. Although this work focuses on validating the simulators generative and predictive capabilities, it establishes a foundation for future research in personalized treatment planning, including what-if analyses and reinforcement learning studies.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.