SEED: Symptom Extraction from English Social Media Posts using Deep Learning and Transfer Learning
Magge, A.; OConnor, K.; Scotch, M.; Graciela, G.-H.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThe increase of social media usage across the globe has fueled efforts in digital epidemiology for mining valuable information such as medication use, adverse drug effects and reports of viral infections that directly and indirectly affect population health. Such specific information can, however, be scarce, hard to find, and mostly expressed in very colloquial language. In this work, we focus on a fundamental problem that enables social media mining for disease monitoring. We present and make available SEED, a natural language processing approach to detect symptom and disease mentions from social media data obtained from platforms such as Twitter and DailyStrength and to normalize them into UMLS terminology. Using multi-corpus training and deep learning models, the tool achieves an overall F1 score of 0.86 and 0.72 on DailyStrength and balanced Twitter datasets, significantly improving over previous approaches on the same datasets. We apply the tool on Twitter posts that report COVID19 symptoms, particularly to quantify whether the SEED system can extract symptoms absent in the training data. The study results also draw attention to the potential of multi-corpus training for performance improvements and the need for continuous training on newly obtained data for consistent performance amidst the ever-changing nature of the social media vocabulary.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 94%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 93%
Similar papers in this journal
- Users’ Reactions on Announced Vaccines against COVID-19 Before Marketing in France: Analysis of Twitter posts 95%
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 94%
- Information retrieval in an infodemic: the case of COVID-19 publications 94%
Similar papers in this journal
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 94%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 94%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.