Comparison of Population Characteristics in Real-World Clinical Oncology Databases in the US: Flatiron Health-Foundation Medicine Clinico-Genomic Databases, Flatiron Health Research Databases, and the National Cancer Institute SEER Population-Based Cancer Registry
Snow, T.; Snider, J.; Comment, L.; Stergiopoulos, S.; Fisher, V.; McCusker, M.; Cho-Phan, C.
Show abstract
BackgroundThe Flatiron Health-Foundation Medicine Clinico-Genomic Databases (CGDBs) are de-identified, real-world data sources that link comprehensive genomic profiling (CGP) data with clinical data derived from electronic health records (EHRs) for patients with cancer. Comparing the CGDBs to the US population of patients with cancer allows researchers to understand the representativeness of a cohort when designing, conducting, and interpreting their analyses. The objective of this study was to compare the demographic and clinical characteristics of patients in the CGDBs with the Flatiron Health Research Databases (FHRDs) and The National Cancer Institutes Surveillance, Epidemiology, and End Results (SEER) population-based cancer registry. MethodsWe compared disease-specific CGDBs that had corresponding disease-specific FHRDs with relevant SEER patients using demographic and clinical characteristics of patients with cancer who had documented care from January 1, 2011 to March 31, 2021. For CGDBs where a corresponding disease-specific FHRD does not exist, comparisons were only done against SEER. The SEER Incidence Data 1975-2018 Research Database was used for this analysis, of which patients with a relevant cancer diagnosis from January 1, 2011 to December 31, 2018 were included. Subgroup analyses were performed to address potential biases related to temporal drifts and allow for a more direct comparison of the datasets as well as to examine biases that may be due to data missingness. The impact of the determination to reimburse for next generation sequencing (NGS) testing was not feasible to analyze given the most recent SEER data was available only through the end of 2018 at the time this study was conducted. ResultsThe overall distribution of cancer types was similar between the 22 CGDB databases and SEER. The overall distributions of gender and diagnosis year were similar across all databases. The CGDB has a lower proportion of patients who were aged 80 years or older at initial diagnosis compared to FHRD and SEER cohorts. However, narrower differences were observed in diseases where targeted therapies are approved and comprehensive genomic profiling is indicated (e.g., Melanoma, NSCLC). The proportion of incomplete records for race in the CGDB and FHRD was greater than in SEER. Completeness of stage varied by disease across all 3 cohorts, but was generally lower in CGDB and FHRD for clinical and data model design reasons. Overall the stage distributions for solid tumor cohorts were similar across CGDB and FHRD with SEER tending to have more earlier stage patients, which is expected given differences in data collection methods for the sources. ConclusionThis comparative analysis of real-world, US-based oncology databases provides crucial insights into the similarities and differences in patient characteristics across these three types of data sources. Observed variances could be due to several factors, including differences in CGP testing dynamics and data collection approaches used to create each of the databases. Ongoing monitoring and evaluation of the representativeness of these databases will be critical to help researchers and regulators contextualize evidence from the CGDBs, particularly as the CGDBs are expected to change over time due to increased adoption of CGP as part of routine clinical practice for a growing number of cancers.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COVID-19 Outcomes in Patients with Cancer: Findings from the University of California Health System Database 95%
- A Video Intervention to Improve Patient Understanding of Tumor Genomic Testing in Patients with Cancer 92%
- Added-value of whole exome and RNA Sequencing in advanced and refractory cancer patients with no molecular-based treatment recommendation based on a 90-gene panel 92%
Similar papers in this journal
- Comprehensive cancer-oriented biobanking resource of human samples for studies of post-zygotic genetic variation involved in cancer predisposition 92%
- Survival benefits of cytoreductive nephrectomy in patients with metastatic renal cell carcinoma: evidence from a SEER-based retrospective cohort study 92%
- Classification performance bias between training and test sets in a limited mammography dataset 92%
Similar papers in this journal
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 94%
- Use of natural language understanding to facilitate surgical de-escalation of axillary staging in patients with breast cancer 93%
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 92%
Similar papers in this journal
- Development and validation of multivariable machine learning algorithms to predict risk of cancer in symptomatic patients referred urgently from primary care 92%
- Reproducibility and transparency characteristics of oncology research evidence 92%
- Cohort-based association study of germline genetic variants with acute and chronic health complications of childhood cancer and its treatment: Genetic risks for childhood cancer complications Switzerland (GECCOS) study protocol 91%
Similar papers in this journal
- Development of a Single Molecule Counting Assay to Differentiate Chromophobe Renal Cancer and Oncocytoma in Clinics 92%
- Hormone Receptor-status Prediction in Breast Cancer Using Gene Expression Profiles and Their Macroscopic Landscape 91%
- NDRG1 expression is an independent prognostic factor in inflammatory breast cancer 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.