Data Veracity of Patients and Health Consumers Reported Adverse Drug Reactions on Twitter: Key Linguistic Features, Twitter Variables, and Association Rules
Lyu, T.; Eidson, A.; Jun, J.; Zhou, X.; Cui, X.; Liang, C.
Show abstract
Adverse drug reactions (ADRs) lead to high disease burden and health expenditure. Aside from traditional data sources used for pharmacovigilance, social media have emerged as an important supplemental data source for monitoring patients and consumers reported ADRs. Recently, there have been increasing concerns about the data veracity of ADRs extracted from social media. Our objective is to categorize different levels of data veracity and explore influential linguistic features and Twitter variables as they may be used for screening ADRs for high data veracity. We annotated a corpus of ADRs with linguistic features validated by clinical experts. Multinomial logistic regression was applied to investigate the associations between the linguistic features and levels of data veracity. We found that using first-person pronouns, expressing negative sentiment, ADR and drug name being in the same sentence were significantly associated with higher levels of data veracity (all p < 0.05), using medical terminology and less indications were associated with good data veracity (p < 0.05), less drug numbers were marginally associated with good data veracity (p = 0.053). These findings suggest an opportunity of developing machine learning models for automatic screening of ADRs from Twitter using identified key linguistic features, Twitter variables, and association rules.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automatic Gender Detection in Twitter Profiles for Health-related Cohort Studies 94%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 93%
- Development and Evaluation of Machine Learning Models for the Detection of Emergency Department Patients with Opioid Misuse from Clinical Notes 92%
Similar papers in this journal
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 94%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 91%
- Abusers indoors and coronavirus outside: an examination of public discourse about COVID-19 and family violence on Twitter using machine learning 91%
Similar papers in this journal
- Which social media platforms facilitate monitoring the opioid crisis? 93%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 92%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 91%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 92%
- Emergence and Evolution of Big Data Analytics in HIV Research: Bibliometric Analysis of Federally Sponsored Studies 2000-2019 91%
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 90%
Similar papers in this journal
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 93%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 90%
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.