A computable phenotype for patients with SARS-CoV2 testing that occurred outside the hospital
Wang, L.; Zipursky, A.; Geva, A.; McMurry, A.; Mandl, K. D.; Miller, T. A.
Show abstract
ObjectiveTo identify a cohort of COVID-19 cases, including when evidence of virus positivity was only mentioned in the clinical text, not in structured laboratory data in the electronic health record (EHR). Materials and MethodsStatistical classifiers were trained on feature representations derived from unstructured text in patient electronic health records (EHRs). We used a proxy dataset of patients with COVID-19 polymerase chain reaction (PCR) tests for training. We selected a model based on performance on our proxy dataset and applied it to instances without COVID-19 PCR tests. A physician reviewed a sample of these instances to validate the classifier. ResultsOn the test split of the proxy dataset, our best classifier obtained 0.56 F1, 0.6 precision, and 0.52 recall scores for SARS-CoV2 positive cases. In an expert validation, the classifier correctly identified 90.8% (79/87) as COVID-19 positive and 97.8% (91/93) as not SARS-CoV2 positive. The classifier identified an additional 960 positive cases that did not have SARS-CoV2 lab tests in hospital, and only 177 of those cases had the ICD-10 code for COVID-19. DiscussionProxy dataset performance may be worse because these instances sometimes include discussion of pending lab tests. The most predictive features are meaningful and interpretable. The type of external test that was performed is rarely mentioned. ConclusionCOVID-19 cases that had testing done outside of the hospital can be reliably detected from the text in EHRs. Training on a proxy dataset was a suitable method for developing a highly performant classifier without labor intensive labeling efforts.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 94%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
Similar papers in this journal
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 95%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 94%
- Distinguishing Admissions Specifically for COVID-19 from Incidental SARS-CoV-2 Admissions: A National Retrospective EHR Study 94%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 94%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 93%
- Identification of an ANCA-Associated Vasculitis Cohort Using Deep Learning and Electronic Health Records 92%
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- Modeling physician variability to prioritize relevant medical record information 94%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 93%
Similar papers in this journal
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 95%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 95%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.