A VAC4EU Systematic Review to Summarize and Critically Appraise Existing Phenotype Libraries Using Electronic Health Records
Mohammadi, S.; Campbell, C.; Sturkenboom, M. C.; Vaz, T. A.
Show abstract
BackgroundPharmacoepidemiology and population health studies using secondary analysis of electronic health care records (EHR) must define study variables through available electronic data. Defining a study variable starts with the identification of a phenotype, which is a defined set of criteria used to identify specific traits or medical conditions. In the real-world data perspective, a phenotype library is a collection of code lists or algorithms that standardize these sets of criteria. We conducted a systematic review of existing phenotype libraries to appraise their attributes, accessibility, interoperability, and portability. MethodsWe systematically searched three databases (Scopus, PubMed, and Web of Science) until June 2024, to identify studies on key characteristics of phenotype libraries. The search combined MeSH terms related to "electronic health records," "phenotype algorithm," and "phenotype library". Study parameters extracted included: library size, vocabularies, phenotype construction tools, validation and library management process, and portability in different sites. FindingsOf 134 articles, 26 met eligibility criteria, leaving nine articles related to eight unique phenotype libraries including CALIBER (Health Data Research UK (HDR UK) Phenotype Library or CALIBER), Centralized Interactive Phenomics Resource (CIPHER), ClinicalCodes Library, Manitoba Centre for Health Policy (MCHP) Concept Dictionary, Observational Health Data Sciences and Informatics (OHDSI) ATLAS, Open CodeLists, Phenotype Execution and Modeling Architecture (PhEMA) Workbench, Phenotype KnowledgeBase (PheKB). These libraries varied largely in size and vocabularies. Each library created rule-based phenotypes, though OHDSI and CIPHER also utilized machine learning. All libraries are both human and machine-readable. Validation processes varied and were only applied to some libraries. All libraries utilized a web-based platform and met at least the minimum requirements for library management, including phenotype definitions, metadata (if applicable), and version control. InterpretationsWe observed large variations in library features including phenotype construction. Transparency about phenotypes and creating computable phenotypes enhance portability and streamline the effective reuse of phenotypes for different systems. FundingThis investigation was supported by a Fellowship awarded by VAC4EU (Vaccine Collaboration for Europe) Phenotype Representation Model: An International and Streamlined Approach to Enhance RWE Studies (grant nr 2023/0001). Research in contextO_ST_ABSEvidence before this studyC_ST_ABSElectronic health data have been used extensively in epidemiology and health data science research for decades, as they offer a wealth of detailed real-world data which may be used to address important evidence gaps. Importance of such data sources has been strongly highlighted following the COVID-19 pandemic, which saw a massive increase in the number of scientific investigations utilizing electronic data to rapidly produce evidence to guide health policy. Recent development of multiple phenotype libraries has presented an important advancement in the field. Libraries serve as repositories for the construction and re-use of phenotypes built with electronic health data including diagnostic codes, laboratory values and demographic information. To our knowledge, a systematic review to identify and describe all existing phenotype libraries has not been undertaken following the COVID-19 pandemic undertaken following the COVID-19 pandemic. Added value of this studyThis study provides a comprehensive systematic review which identifies and describes all currently existing phenotype libraries. We summarize and compare phenotype construction processes, data sources, user interfaces, portability and algorithm validation practices across 8 individual phenotype libraries. We highlight how these libraries facilitate robust and transparency, open scientific practices in digital health research, and identify potential opportunities for innovation. This systematic review serves as an important benchmark study, providing a central documentation and description of phenotype libraries built to date. Implications of all the available evidenceThe use of phenotype libraries and collaboratively constructed and validated phenotypes in health research may greatly improve the robustness and impact of health research. We describe how libraries can be currently be used to improve research practice, as well as how existing libraries
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 91%
Similar papers in this journal
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 94%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 94%
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 93%
Similar papers in this journal
- Protocol for Development of a Reporting Guideline for Causal and Counterfactual Prediction Models 93%
- Data-driven discovery of changes in clinical code usage over time: a case-study on changes in cardiovascular disease recording in two English electronic health records databases (2001-2015) 92%
- Development and validation of multivariable prediction models for adverse COVID-19 outcomes in IBD patients 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.