A Query Taxonomy Describes Performance of Patient-Level Retrieval from Electronic Health Record Data
Chamberlin, S. R.; Bedrick, S. D.; Cohen, A. M.; Wang, Y.; Wen, A.; Liu, S.; Liu, H.; Hersh, W.
Show abstract
Performance of systems used for patient cohort identification with electronic health record (EHR) data is not well-characterized. The objective of this research was to evaluate factors that might affect information retrieval (IR) methods and to investigate the interplay between commonly used IR approaches and the characteristics of the cohort definition structure. We used an IR test collection containing 56 test patient cohort definitions, 100,000 patient records originating from an academic medical institution EHR data warehouse, and automated word-base query tasks, varying four parameters. Performance was measured using B-Pref. We then designed 59 taxonomy characteristics to classify the structure of the 56 topics. In addition, six topic complexity measures were derived from these characteristics for further evaluation using a beta regression simulation. We did not find a strong association between the 59 taxonomy characteristics and patient retrieval performance, but we did find strong performance associations with the six topic complexity measures created from these characteristics, and interactions between these measures and the automated query parameter settings. Some of the characteristics derived from a query taxonomy could lead to improved selection of approaches based on the structure of the topic of interest. Insights gained here will help guide future work to develop new methods for patient-level cohort discovery with EHR data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review 94%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 92%
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 92%
Similar papers in this journal
- A Comparative Analysis of System Features Used in the TREC-COVID Information Retrieval Challenge 93%
- Detecting Goals of Care Conversations in Clinical Notes with Active Learning 93%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 93%
Similar papers in this journal
- Evaluation of Patient-Level Retrieval from Electronic Health Record Data for a Cohort Discovery Task 95%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 93%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 92%
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 93%
- Temporal Relationship of Computed and Structured Diagnoses in Electronic Health Record Data 92%
- An Ontology-based Approach to Guide and Document Variable and Data Source Selection and Data Integration Process to Support Integrative Data Analysis in Cancer Outcomes Research 91%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 94%
- Extracting social determinants of health from electronic health records: development and comparison of rule-based and large language models-based methods 93%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.