Moving Biosurveillance Beyond Coded Data: AI for Symptom Detection from Physician Notes
McMurry, A.; Zipursky, A. R.; Geva, A.; Olson, K. L.; Jones, J.; Ignatov, V.; Miller, T.; Mandl, K. D.
Show abstract
BackgroundReal-time surveillance of emerging infectious diseases necessitates a dynamically evolving, computable case definition, which frequently incorporates symptom-related criteria. For symptom detection, both population health monitoring platforms and research initiatives primarily depend on structured data extracted from electronic health records. ObjectiveTo validate and test an artificial intelligence (AI) based Natural Language Processing (NLP) pipeline for detecting COVID-19 symptoms from physician notes. MethodsSubjects in this retrospective cohort study are patients 21 years old and younger, who presented to a pediatric emergency department (ED) at a large academic childrens hospital between March 1, 2020 and May 31, 2022. ED notes for all patients were processed with an NLP pipeline tuned to detect the mention of 11 COVID-19 symptoms based on CDC criteria. For a gold standard, 3 subject matter experts labeled 226 ED notes and had strong agreement (F1=98.6; PPV=97.2; Recall=100.0). F1, PPV, and recall were used to compare the performance of both NLP and ICD-10 to the gold standard chart review. As a formative use case, variations in symptom patterns were measured across SARS-Cov2 variant eras. ResultsThere were 85,678 ED encounters during the study period, 4.0% with patients with COVID-19. NLP was more accurate at identifying encounters with patients that had any of the COVID-19 symptoms (F1=79.6) than ICD-10 codes (F1=45.1%). NLP accuracy was higher for positive symptoms (recall=93%) than ICD-10 (recall=30%). However, ICD-10 accuracy was higher for negative symptoms (specificity=99.4%) than NLP (specificity=91.7%). Congestion or runny nose showed the highest accuracy difference: NLP F1=82.8%, ICD-10 F1=4.2%. Prevalence of NLP symptoms among patients with COVID-19 differed across variant eras. And patients with COVID-19 were more likely to have each symptom than patients without this disease. Effect sizes (odds ratios) varied across pandemic eras. ConclusionsThis study establishes the value of AI based NLP as a highly effective tool for real-time COVID-19 symptom detection in pediatric patients, outperforming traditional ICD-10 methods. It also reveals the evolving nature of symptom prevalence across different virus variants, underscoring the need for dynamic, technology-driven approaches in infectious disease surveillance.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical prediction rule for SARS-CoV-2 infection from 116 U.S. emergency departments 94%
- Derivation and validation of a triage tool for acutely ill adults with suspected COVID-19: The PRIEST observational cohort study 94%
- Predicting patients with false negative SARS-CoV-2 testing at hospital admission: A retrospective multi-center study 94%
Similar papers in this journal
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 92%
- The Validity of the Parsley Symptom Index: an e-PROM designed for Telehealth 89%
- Is virtual care the new normal? Evidence supporting Covid-19’s durable transformation on healthcare delivery 89%
Similar papers in this journal
- Distinguishing Admissions Specifically for COVID-19 from Incidental SARS-CoV-2 Admissions: A National Retrospective EHR Study 94%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 91%
Similar papers in this journal
- High proportion of post-acute sequelae of SARS-CoV-2 infection in individuals 1-6 months after illness and association with disease severity in an outpatient telemedicine population 94%
- Clinical symptoms among ambulatory patients tested for SARS-CoV-2 94%
- Demographic disparities in clinical outcomes of COVID-19: data from a statewide cohort in South Carolina 93%
Similar papers in this journal
- Score for Emergency Risk Prediction (SERP): An Interpretable Machine Learning AutoScore–Derived Triage Tool for Predicting Mortality after Emergency Admissions 92%
- Consistency of performance of adverse outcome prediction models for hospitalized COVID-19 patients 91%
- Self-reported Memory Problems Eight Months after Non-Hospitalized COVID-19 in a Large Cohort 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.