Back

Long Covid symptoms and diagnosis in primary care: a cohort study using structured and unstructured data in The Health Improvement Network primary care database

Shah, A. D.; Subramanian, A. D.; Lewis, J.; Dhalla, S.; Ford, E.; Haroon, S.; Kuan, V.; Nirantharakumar, K.

2023-01-09 epidemiology
10.1101/2023.01.06.23284202 medRxiv
Show abstract

BACKGROUNDLong Covid is a widely recognised consequence of COVID-19 infection, but little is known about the burden of symptoms that patients present with in primary care, as these are typically recorded only in free text clinical notes. Our objectives were to compare symptoms in patients with and without a history of COVID-19, and investigate symptoms associated with a Long Covid diagnosis. METHODSWe used primary care electronic health record data from The Health Improvement Network (THIN), a Cegedim database. We included adults registered with participating practices in England, Scotland or Wales. We extracted information about 89 symptoms and Long Covid diagnoses from free text using natural language processing. We calculated hazard ratios (adjusted for age, sex, baseline medical conditions and prior symptoms) for each symptom from 12 weeks after the COVID-19 diagnosis. FINDINGSWe compared 11,015 patients with confirmed COVID-19 and 18,098 unexposed controls. Only 20% of symptom records were coded, with 80% in free text. A wide range of symptoms were associated with COVID-19 at least 12 weeks post-infection, with strongest associations for fatigue (adjusted hazard ratio (aHR) 3.99, 95% confidence interval (CI) 3.59, 4.44), shortness of breath (aHR 3.14, 95% CI 2.88, 3.42), palpitations (aHR 2.75, 95% CI 2.28, 3.32), and phlegm (aHR 2.88, 95% CI 2.30, 3.61). However, a limited subset of symptoms were recorded within 7 days prior to a Long Covid diagnosis in more than 20% of cases: shortness of breath, chest pain, pain, fatigue, cough, and anxiety / depression. CONCLUSIONNumerous symptoms are reported to primary care at least 12 weeks after COVID-19 infection, but only a subset are commonly associated with a GP diagnosis of Long Covid.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.