Biases in Race and Ethnicity Introduced by Filtering Electronic Health Records for 'Complete Data'
Acitores Cortina, J. M.; Fatapour, Y.; Zietz, M.; Brown, K. L.; Gisladottir, U.; Peter, D.; Bear Don't Walk, O. J.; Kuchi, A.; Srinivasan, A.; Liu, H.; Berkowitz, J. S.; Tsang, K.; Kivelson, S.; Friedrich, N.; Tatonetti, N. P.
Show abstract
ObjectiveIntegrated clinical databases from national biobanks have advanced the capacity for disease research. Data quality and completeness filters are used when building clinical cohorts to address limitations of data missingness. However, these filters may unintentionally introduce systemic biases when they are correlated with race and ethnicity. In this study, we examined the race/ethnicity biases introduced by applying common filters to four clinical records databases. Materials and MethodsWe used 19 filters commonly used in electronic health records research on the availability of demographics, medication records, visit details, observation periods, and other data types. We evaluated the effect of applying these filters on self-reported race and ethnicity. This assessment was performed across four databases comprising approximately 12 million patients. ResultsApplying the observation period filter led to a substantial reduction in data availability across all races and ethnicities in all four datasets. However, among those examined, the availability of data in the white group remained consistently higher compared to other racial groups after applying each filter. Conversely, the Black/African American group was the most impacted by each filter on these three datasets, Cedars-Sinai dataset, UK-Biobank, and Columbia University Dataset. Discussion and ConclusionOur findings underscore the importance of using only necessary filters as they might disproportionally affect data availability of minoritized racial and ethnic populations. Researchers must consider these unintentional biases when performing data-driven research and explore techniques to minimize the impact of these filters, such as probabilistic methods or the use of machine learning and artificial intelligence.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 93%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 93%
- Automatic phenotyping of electronical health record: PheVis algorithm 93%
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 92%
Similar papers in this journal
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 95%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 94%
- Development of a COVID-19 Application Ontology for the ACT Network 93%
Similar papers in this journal
- Diversity and inclusion: A hidden additional benefit of Open Data 94%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 94%
- From months to minutes: creating Hyperion, a novel data management system expediting data insights for oncology research and patient care 93%
Similar papers in this journal
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 94%
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 94%
- LinkR: an open source, low-code and collaborative data science platform for healthcare data analysis and visualization 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.