Extending Inferences From Sample To Target Populations: On The Generalizability Of A Real-World Clinico-Genomic Database Non-Small Cell Lung Cancer Cohort
Thomas, D. S.; Collin, S.; Berrocal-Almanza, L.; Stirnadel- Farrant, H.; Zhang, Y.; Sun, P.
Show abstract
PurposeThe study aimed to extend inferences on the Average Treatment Effects (ATEs) from a Flatiron Health Clinico-Genomic Database (CGDB) Non-Small Cell Cancer (NSCLC) sample to the target population represented by SEER cancer registrations. The work also demonstrates the potential for the non-random selection to cause bias through a quantitative framework that compares the marginal and joint distributions of effect modifiers through each non-random selection process. MethodsATEs for a binary treatment were estimated within the sample (SATE) and extended to the population (PATE) using combined information from SEER and a weighted estimator to standardize the joint distributions of effect modifiers. To understand potential biases through selection, the marginal and joint distributions of effect modifiers were compared through each stepwise process using two referent populations: SEER registrations and a superset of all NSCLC patients in the Flatiron Health network. ResultsWithin a subset of 1,166 stage III-IV NSCLCs receiving a binary treatment and combined information from 149,056 SEER registrations, the SATE & PATE for differences in survival at month 48 were -3.7 (-8.7, 1.6) & -3.4 (-8.7, 2.8) percentage points. Through each sequential selection, the joint distributions of effect modifiers were not discernibly different among the selected & unselected. Estimates of survival were unbiased by selection. ConclusionsCombined information from cancer registrations can be used to extend inferences from a selected sample to the target population. ATEs within a CGDB were an unbiased estimate of the population because the sequential selection did not differentially select effect modifiers causative of survival. Key PointsO_LICombined information from cancer registrations can be used to extend inferences from a selected sample to the target population C_LIO_LIWe outline a quantitative framework for determining the potential for non-random selection to cause bias, through comparing the marginal and joint distributions of effect modifiers through each sequential selection process C_LIO_LIAverage Treatment Effects within a highly selected genetic cohort were an unbiased estimate of the population because the sequential selection did not differentially select effect modifiers causative of survival C_LI Plain Language SummaryIn pharmacoepidemiology we oftentimes learn about the effectiveness of treatments within a smaller sample in attempt to understand how they would work in the broader population. When this sampling is non-random -- like when we use Real-World Data in the form of Electronic Medical Records or insurance claims -- what we learn in this sample may not translate to the broader population. In this study we show how we can publicly available data in the form of cancer registrations to better understand how these treatments work in the population. We show that what we learn about treatments within our highly selected sample that required patients to undergo expensive genetic testing does in fact translate to the population. We also provide a framework showing why this was the case: patients in our selected sample were very similar in their characteristics to the broader population.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting the need for escalation of care or death from repeated daily clinical observations and laboratory results in patients with SARS-CoV-2 during 2020: a retrospective population-based cohort study from the United Kingdom 91%
- Protection of previous SARS-CoV-2 infection is similar to that of BNT162b2 vaccine protection: A three-month nationwide experience from Israel 90%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 90%
Similar papers in this journal
- Simple Linear Cancer Risk Prediction Models with Novel Features Outperform Complex Approaches 94%
- Machine learning and mechanistic modeling for prediction of metastatic relapse in breast cancer 93%
- Histology-based Prediction of Therapy Response to Neoadjuvant Chemotherapy for Esophageal and Esophagogastric Junction Adenocarcinomas Using Deep Learning 92%
Similar papers in this journal
- Uncovering interpretable potential confounders in electronic medical records 93%
- Angiogenic and Immune Predictors of Neoadjuvant Axitinib Response in Renal Cell Carcinoma with Venous Tumour Thrombus 93%
- Pan-cancer analysis demonstrates that integrating polygenic risk scores with modifiable risk factors improves risk prediction 93%
Similar papers in this journal
- Income inequality and access to advanced immunotherapy for lung cancer: the case of Durvalumab in the Netherlands 93%
- Strength of Statistical Evidence for the Efficacy of Cancer Drugs: A Bayesian Re-Analysis of Trials Supporting FDA Approval 91%
- Information bias of vaccine effectiveness estimation due to informed consent for national registration of COVID-19 vaccination: estimation and correction using a data augmentation model 91%
Similar papers in this journal
- Quantifying absolute treatment effect heterogeneity for time-to-event outcomes across different risk strata: divergence of conclusions with risk difference and restricted mean survival difference 94%
- Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering 91%
- Association between household composition and severe COVID-19 outcomes in older people by ethnicity: an observational cohort study using the OpenSAFELY platform 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.