Using set visualization techniques to investigate and explain patterns of missing values in electronic health records
Ruddle, R. A.; Adnan, M.; Hall, M.
Show abstract
ObjectivesMissing data is the most common data quality issue in electronic health records (EHRs). Checks are typically limited to counting the number of missing values in individual fields, but researchers and organisations need to understand multi-field missing data patterns, and counts or numerical summaries are poorly suited to that. This study shows how set-based visualization enables multi-field missing data patterns to be discovered and investigated. DesignDevelopment and evaluation of interactive set visualization techniques to find patterns of missing data and generate actionable insights. Setting and participantsAnonymised Admitted Patient Care health records for NHS hospitals and independent sector providers in England. The visualization and data mining software was run over 16 million records and 86 fields in the dataset. ResultsThe dataset contained 960 million missing values. Set visualization bar charts showed how those values were distributed across the fields, including several fields that, unexpectedly, were not complete. Set intersection heatmaps revealed unexpected gaps in diagnosis, operation and date fields. Information gain ratio and entropy calculations allowed us to identify the origin of each unexpected pattern, in terms of the values of other fields. ConclusionsOur findings show how set visualization reveals important insights about multi-field missing data patterns in large EHR datasets. The study revealed both rare and widespread data quality issues that were previously unknown to an epidemiologist, and allowed a particular part of a specific hospital to be pinpointed as the origin of rare issues that NHS Digital did not know exist. ARTICLE SUMMARY Strengths and limitations of this studyO_LIThis study demonstrates the utility of interactive set visualization techniques for finding and explaining patterns of missing values in electronic health records, irrespective of whether those patterns are common or rare. C_LIO_LIThe techniques were evaluated in a case study with a large (16-million record; 86 field) Admitted Patient Care dataset from NHS hospitals. C_LIO_LIThere was only one data table in the dataset. However, ways to adapt the techniques for longitudinal data and relational databases are described. C_LIO_LIThe evaluation only involved one dataset, but that was from a national organisation that provides many similar datasets each year to researchers and organisations. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From months to minutes: creating Hyperion, a novel data management system expediting data insights for oncology research and patient care 94%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 94%
- A data management system for precision medicine 94%
Similar papers in this journal
- LinkR: an open source, low-code and collaborative data science platform for healthcare data analysis and visualization 94%
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 93%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 93%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 94%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
Similar papers in this journal
- A Simple Electronic Medical Record System Designed for Research 94%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 93%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.