Back

Data Veracity of Patients and Health Consumers Reported Adverse Drug Reactions on Twitter: Key Linguistic Features, Twitter Variables, and Association Rules

Lyu, T.; Eidson, A.; Jun, J.; Zhou, X.; Cui, X.; Liang, C.

2020-11-05 health informatics
10.1101/2020.11.03.20225532 medRxiv
Show abstract

Adverse drug reactions (ADRs) lead to high disease burden and health expenditure. Aside from traditional data sources used for pharmacovigilance, social media have emerged as an important supplemental data source for monitoring patients and consumers reported ADRs. Recently, there have been increasing concerns about the data veracity of ADRs extracted from social media. Our objective is to categorize different levels of data veracity and explore influential linguistic features and Twitter variables as they may be used for screening ADRs for high data veracity. We annotated a corpus of ADRs with linguistic features validated by clinical experts. Multinomial logistic regression was applied to investigate the associations between the linguistic features and levels of data veracity. We found that using first-person pronouns, expressing negative sentiment, ADR and drug name being in the same sentence were significantly associated with higher levels of data veracity (all p < 0.05), using medical terminology and less indications were associated with good data veracity (p < 0.05), less drug numbers were marginally associated with good data veracity (p = 0.053). These findings suggest an opportunity of developing machine learning models for automatic screening of ADRs from Twitter using identified key linguistic features, Twitter variables, and association rules.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.