Characterizing Variability of EHR-Driven Phenotype Definitions
Brandt, P. S.; Kho, A. N.; Luo, Y.; Pacheco, J. A.; Walunas, T. L.; Hakonarson, H.; Hripcsak, G.; Liu, C.; Shang, N.; Weng, C.; Walton, N.; Carrell, D. S.; Crane, P. K.; Larson, E.; Chute, C. G.; Kullo, I.; Carroll, R.; Denny, J. C.; Ramirez, A.; Wei, W.-Q.; Pathak, J.; Wiley, L. K.; Richesson, R.; Starren, J. B.; Rasmussen, L. V.
Show abstract
ObjectiveAnalyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the types of logical constructs used. Materials & MethodsA sample of 33 phenotype definitions used in research and published to the Phenotype KnowledgeBase (PheKB), that are represented using Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. ResultsMost of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DiscussionDespite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions is low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. ConclusionThe phenotypes analyzed show significant variation in specific logical, arithmetic and other operators, but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 96%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 96%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 95%
- De-novo FAIRification via an Electronic Data Capture system by automated transformation of filled electronic Case Report Forms into machine-readable data 95%
- Genomic Considerations for FHIR; eMERGE Implementation Lessons 94%
Similar papers in this journal
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 95%
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 94%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.