Back

External Control Arm with Synthetic Real-world Data for Comparative Oncology using Single Trial Arm Evidence (ECLIPSE): A Case Study using Lung-MAP S1400I

Gupta, A.; Segars, L.; Singletary, D.; Hansen, J. L.; Geale, K.; Arora, A.; Gomes, M.; Ramagopalan, S.; Cheung, W.; Arora, P.

2024-09-11 oncology
10.1101/2024.09.10.24313417 medRxiv
Show abstract

2.Single-arm trials supplemented with external comparator arm(s) (ECA) derived from real-world data are sometimes used when randomized trials are infeasible. However, due to data sharing restrictions, privacy/security concerns, or for logistical reasons, patient-level real-world data may not be available to researchers for analysis. Instead, it may be possible to use generative models to construct synthetic data from the real-world dataset that can then be freely shared with researchers. Although the use of generative models and synthetic data is gaining prominence, the extent to which a synthetic data ECA can replace original data while preserving patient privacy in small samples is unclear. ObjectiveTo compare the efficacy of nivolumab + ipilimumab combination therapy ("experimental arm") versus nivolumab monotherapy ("control arm") in patients with metastatic non-small cell lung cancer (mNSCLC) using real-world data from two real-world databases ("original ECA"), and synthetic data versions of these datasets ("synthetic ECA"), with the aim of validating synthetic data for use in ECA analysis. Study designNon-randomized analyses of treatment efficacy comparing the experimental arm to the (i) original ECA and (ii) synthetic ECA, with baseline confounding adjustment. Data sourcesThe experimental arm is from the Lung-MAP no-match substudy S1400I (NCT02785952) provided by National Clinical Trials Network (NCTN) in the United States. The real-world data source for the ECA is data from population-based oncology data from the Canadian province of Alberta, and from Nordic countries in Europe, specifically Denmark and Norway.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.