Replicate Engineered Virtual Patient Populations as Surrogates for Real Patient-Level Data
Alenghat, F. J.
Show abstract
ObjectivesTo demonstrate a new method for generating virtual, individual-level data by testing it on a known clinical trial population. DesignVirtualization of aggregate data from a clinical trial. SettingVirtual Participants936,100 virtual patients InterventionsNone Main Outcomes MeasuresOdds ratios for adverse outcomes in virtual patient populations compared to clinical trial participants. MethodsThe replicate engineered virtual patient populations (RE-ViPPs) method, based on aggregate cross-tabulated categorical population data, does not require access to individual-level data. Using sequential regression combined with randomization, it generates virtual individual patients to comprise populations that, on average, closely resemble the real population in question. The method is validated by applying it to aggregated data from the seminal SPRINT trial, which compared intensive versus standard blood pressure treatment goals on major adverse cardiovascular events. ResultsThe method yields virtual populations, each with 9361 patients, faithfully mimicking the real SPRINT participants. Multiple logistic regression on 100 such populations shows that factors with the highest odds ratios for the primary event are, in descending order, past clinical cardiovascular disease, age [≥] 75, chronic kidney disease, high non-HDL, and smoking history. Intensive blood pressure treatment, the trials intervention, had an odds ratio of 0.74 [0.63-0.87]. On all these measures, the 100 RE-ViPPs mirrored the real SPRINT participants, including the intensive therapy result (actual SPRINT odds ratio: 0.74 [0.62-0.88]). ConclusionsClinical data dissemination has limitations. The most coveted data is descriptive at the individual level but comes with significant cost, effort, and time. There is potential for privacy breaches, and the open-data movement has progressed slowly due to data-ownership concerns. RE-ViPPs closely matched the true SPRINT population. Applied to trials, registries, and databases, RE-ViPPs could reduce open-data burdens by encouraging dissemination of aggregate cross-tabulated real data that allow investigators to generate and measure virtual patients.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Controlled evaLuation of Angiotensin Receptor Blockers for COVID-19 respIraTorY disease (CLARITY): Statistical analysis plan for a randomised controlled Bayesian adaptive sample size trial 92%
- Rationale and design of the Learning Implementation of Guideline-based decision support system for Hypertension Treatment (LIGHT) Trial and LIGHT-ACD Trial 92%
- Graphing and reporting heterogeneous treatment effects through reference classes 91%
Similar papers in this journal
- Analysis of clinical trial registry entry histories using the novel R package cthist 93%
- Actionable absolute risk prediction of atherosclerotic cardiovascular disease: a behavior-management approach based on data from 464,547 UK Biobank participants 93%
- Development and Validation of the Michigan Chronic Disease Simulation Model (MICROSIM) 92%
Similar papers in this journal
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 94%
- Aggregating Multiple Real-World Data Sources using a Patient-Centered Health Data Sharing Platform: an 8-week Cohort Study among Patients Undergoing Bariatric Surgery or Catheter Ablation of Atrial Fibrillation 94%
- Zero Shot Health Trajectory Prediction Using Transformer 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.