Back

A VAC4EU Systematic Review to Summarize and Critically Appraise Existing Phenotype Libraries Using Electronic Health Records

Mohammadi, S.; Campbell, C.; Sturkenboom, M. C.; Vaz, T. A.

2024-12-16 health informatics
10.1101/2024.12.16.24319076 medRxiv
Show abstract

BackgroundPharmacoepidemiology and population health studies using secondary analysis of electronic health care records (EHR) must define study variables through available electronic data. Defining a study variable starts with the identification of a phenotype, which is a defined set of criteria used to identify specific traits or medical conditions. In the real-world data perspective, a phenotype library is a collection of code lists or algorithms that standardize these sets of criteria. We conducted a systematic review of existing phenotype libraries to appraise their attributes, accessibility, interoperability, and portability. MethodsWe systematically searched three databases (Scopus, PubMed, and Web of Science) until June 2024, to identify studies on key characteristics of phenotype libraries. The search combined MeSH terms related to "electronic health records," "phenotype algorithm," and "phenotype library". Study parameters extracted included: library size, vocabularies, phenotype construction tools, validation and library management process, and portability in different sites. FindingsOf 134 articles, 26 met eligibility criteria, leaving nine articles related to eight unique phenotype libraries including CALIBER (Health Data Research UK (HDR UK) Phenotype Library or CALIBER), Centralized Interactive Phenomics Resource (CIPHER), ClinicalCodes Library, Manitoba Centre for Health Policy (MCHP) Concept Dictionary, Observational Health Data Sciences and Informatics (OHDSI) ATLAS, Open CodeLists, Phenotype Execution and Modeling Architecture (PhEMA) Workbench, Phenotype KnowledgeBase (PheKB). These libraries varied largely in size and vocabularies. Each library created rule-based phenotypes, though OHDSI and CIPHER also utilized machine learning. All libraries are both human and machine-readable. Validation processes varied and were only applied to some libraries. All libraries utilized a web-based platform and met at least the minimum requirements for library management, including phenotype definitions, metadata (if applicable), and version control. InterpretationsWe observed large variations in library features including phenotype construction. Transparency about phenotypes and creating computable phenotypes enhance portability and streamline the effective reuse of phenotypes for different systems. FundingThis investigation was supported by a Fellowship awarded by VAC4EU (Vaccine Collaboration for Europe) Phenotype Representation Model: An International and Streamlined Approach to Enhance RWE Studies (grant nr 2023/0001). Research in contextO_ST_ABSEvidence before this studyC_ST_ABSElectronic health data have been used extensively in epidemiology and health data science research for decades, as they offer a wealth of detailed real-world data which may be used to address important evidence gaps. Importance of such data sources has been strongly highlighted following the COVID-19 pandemic, which saw a massive increase in the number of scientific investigations utilizing electronic data to rapidly produce evidence to guide health policy. Recent development of multiple phenotype libraries has presented an important advancement in the field. Libraries serve as repositories for the construction and re-use of phenotypes built with electronic health data including diagnostic codes, laboratory values and demographic information. To our knowledge, a systematic review to identify and describe all existing phenotype libraries has not been undertaken following the COVID-19 pandemic undertaken following the COVID-19 pandemic. Added value of this studyThis study provides a comprehensive systematic review which identifies and describes all currently existing phenotype libraries. We summarize and compare phenotype construction processes, data sources, user interfaces, portability and algorithm validation practices across 8 individual phenotype libraries. We highlight how these libraries facilitate robust and transparency, open scientific practices in digital health research, and identify potential opportunities for innovation. This systematic review serves as an important benchmark study, providing a central documentation and description of phenotype libraries built to date. Implications of all the available evidenceThe use of phenotype libraries and collaboratively constructed and validated phenotypes in health research may greatly improve the robustness and impact of health research. We describe how libraries can be currently be used to improve research practice, as well as how existing libraries

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.