Characterization of long-term patient-reported symptoms of COVID-19: an analysis of social media data
Banda, J. M.; Adderley, N.; Ahmed, W.-U.-R.; AlGhoul, H.; Alser, O.; Alser, M.; Areia, C.; Cogenur, M.; Fister, K.; Gombar, S.; Huser, V.; Jonnagaddala, J.; Lai, L.; Leis, A.; Mateu, L.; Mayer, M. A.; Minty, E.; Morales, D. R.; Natarajan, K.; Paredes, R.; Periyakoil, V. S.; Prats-Uribe, A.; Ross, E. G.; Singh, G. V.; Subbian, V.; Vivekanantham, A.; Prieto-Alhambra, D.
Show abstract
As the SARS-CoV-2 virus (COVID-19) continues to affect people across the globe, there is limited understanding of the long term implications for infected patients1-3. While some of these patients have documented follow-ups on clinical records, or participate in longitudinal surveys, these datasets are usually designed by clinicians, and not granular enough to understand the natural history or patient experiences of long COVID. In order to get a complete picture, there is a need to use patient generated data to track the long-term impact of COVID-19 on recovered patients in real time. There is a growing need to meticulously characterize these patients experiences, from infection to months post-infection, and with highly granular patient generated data rather than clinician narratives. In this work, we present a longitudinal characterization of post-COVID-19 symptoms using social media data from Twitter. Using a combination of machine learning, natural language processing techniques, and clinician reviews, we mined 296,154 tweets to characterize the post-acute infection course of the disease, creating detailed timelines of symptoms and conditions, and analyzing their symptomatology during a period of over 150 days.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 94%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 94%
- Real-time clinician text feeds from electronic health records 94%
Similar papers in this journal
- Machine learning approach to dynamic risk modeling of mortality in COVID-19: a UK Biobank study 94%
- Using explainable machine learning to identify patients at risk of reattendance at discharge from emergency departments 93%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
Similar papers in this journal
- Hospital-wide Natural Language Processing summarising the health data of 1 million patients 95%
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 93%
- Use of large language models as a scalable approach to understanding public health discourse 92%
Similar papers in this journal
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 93%
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 92%
- Extraction of Crohn's Disease Clinical Phenotypes from Clinical Text Using Natural Language Processing 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.