Comparison of Population Characteristics in Real-World Clinical Oncology Databases in the US: Flatiron Health, SEER, and NPCR
Ma, X.; Long, L.; Moon, S.; Adamson, B. J. S.; Baxi, S. S.
Show abstract
Background and ObjectiveThe Surveillance, Epidemiology, and End Results Program (SEER) program and the National Program of Cancer Registries (NPCR), are authoritative sources for population cancer surveillance and research in the US. An increasing number of recent oncology studies are based on the electronic health record (EHR)-derived de-identified databases created and maintained by Flatiron Health. This report describes the differences in the originating sources and data development processes, and compares baseline demographic characteristics in the cancer-specific databases from Flatiron Health, SEER, and NPCR, to facilitate interpretation of research findings based on these sources. MethodsPatients with documented care from January 1, 2011 through May 31, 2019 in a series of EHR-derived Flatiron Health de-identified databases covering multiple tumor types were included. SEER incidence data (obtained from the SEER 18 database) and NPCR incidence data (obtained from the US Cancer Statistics public use database) for malignant cases diagnosed from January 1, 2011 to December 31, 2016 were included. Comparisons of demographic variables were performed across all disease-specific databases, for all patients and for the subset diagnosed with advanced-stage disease. ResultsAs of May 2019, a total of 201,570 patients with 19 different cancer types were included in Flatiron Health datasets. In an overall comparison to national cancer registries, patients in the Flatiron Health databases had similar sex, age at initial diagnosis, and geographic distributions but appeared to be diagnosed with later stages of disease compared with patients in other datasets. For variables such as stage and race, Flatiron Health databases had a greater degree of incompleteness. There are variations in these trends by cancer types. ConclusionsThese three databases present general similarities in demographic and geographic distribution, but there are overarching differences across the populations they cover. Differences in data sourcing (medical oncology EHRs vs cancer registries), and disparities in sampling approaches and rules of data acquisition may explain some of these divergences. Furthermore, unlike the steady information flow entered into registries, the availability of medical oncology EHR-derived information reflects the extent of involvement of medical oncology clinics at different points in the specialty management of individual diseases, resulting in inter-disease variability. These differences should be considered when interpreting study results obtained with these databases.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 92%
- Southern European Prospective Investigation Into Childhood Cancer and Nutrition (EPIC kids ): Study Design and Protocol 92%
- Comprehensive cancer-oriented biobanking resource of human samples for studies of post-zygotic genetic variation involved in cancer predisposition 92%
Similar papers in this journal
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 94%
- Use of natural language understanding to facilitate surgical de-escalation of axillary staging in patients with breast cancer 93%
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 93%
Similar papers in this journal
- COVID-19 Outcomes in Patients with Cancer: Findings from the University of California Health System Database 94%
- Can Large Language Models Aid Caregivers of Pediatric Cancer Patients in Information Seeking? A Cross-Sectional Investigation 92%
- A Video Intervention to Improve Patient Understanding of Tumor Genomic Testing in Patients with Cancer 91%
Similar papers in this journal
- Development and validation of multivariable machine learning algorithms to predict risk of cancer in symptomatic patients referred urgently from primary care 93%
- Reproducibility and transparency characteristics of oncology research evidence 92%
- The impact of the COVID-19 pandemic and related control measures on cancer diagnosis in Catalonia:A time-series analysis of primary care electronic health records covering about 5 million people. 91%
Similar papers in this journal
- Missing data in the medical record for oncology patients: prevalence and association with outcomes 96%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 90%
- COVID-19 outcomes, risk factors and associations by race: a comprehensive analysis using electronic health records data in Michigan Medicine 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.