Using Primary Care Text Data And Natural Language Processing To Monitor COVID-19 In Toronto, Canada
Meaney, C.; Moineddin, R.; Kalia, S.; Aliarzadeh, B.; Greiver, M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSObjectiveC_ST_ABSTo investigate whether a rule-based natural language processing (NLP) system, applied to primary care clinical text data, can be used to monitor COVID-19 viral activity in Toronto, Canada. DesignWe employ a retrospective cohort design. We include primary care patients with a clinical encounter between January 1, 2020 and December 31, 2020 at one of 44 participating clinical sites. Setting and ContextThe study setting is Toronto, Canada. During the study timeframe the city experienced a first wave of COVID-19 in spring 2020; followed by a second viral resurgence beginning in the fall of 2020. Methods and DataStudy objectives are descriptive. We use an expert derived dictionary, pattern matching tools and a contextual analyzer to classify documents as 1) COVID-19 positive, 2) COVID-19 negative, or 3) unknown COVID-19 status. We apply the COVID-19 biosurveillance system across three primary care electronic medical record text streams: 1) lab text, 2) health condition diagnosis text and 3) clinical notes. We enumerate COVID-19 entities in the clinical text and estimate the proportion of patients with a positive COVID-19 record. We construct a primary care COVID-19 NLP-derived time series and investigate its correlation with other external public health series: 1) lab confirmed COVID-19 cases, 2) COVID-19 hospitalizations, 3) COVID-19 ICU admissions, and 4) COVID-19 intubations. ResultsOver the study timeframe 1,976 COVID-19 positive documents, and 277 unique COVID-19 entities were identified in the lab text. 539 COVID-19 positive documents and 121 unique COVID-19 entities were identified in the health condition diagnosis text. And 4,018 COVID-19 positive documents, and 644 unique COVID-19 entities were identified in the clinical notes. A total of 196,440 unique patients were observed over the study timeframe, of which 4,580 (2.3%) had at least one positive COVID-19 document in their primary care electronic medical record. We constructed an NLP-derived COVID-19 time series describing the temporal dynamics of COVID-19 positivity status over the study timeframe. The NLP derived series correlates strongly with external public health series under investigation. ConclusionsUsing a rule-based NLP system we identified hundreds of unique COVID-19 entities, and thousands of COVID-19 positive documents, across millions of clinical text documents. Future work should continue to investigate how high quality, low-cost, passively collected primary care electronic medical record clinical text data can be used for COVID-19 monitoring and surveillance.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 94%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 93%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
Similar papers in this journal
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 96%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 95%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 93%
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 94%
- Development of a COVID-19 Application Ontology for the ACT Network 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.