OnSIDES (ON-label SIDE effectS resource) Database : Extracting Adverse Drug Events from Drug Labels using Natural Language Processing Models
Tanaka, Y.; Chen, H. Y.; Belloni, P.; Gisladottir, U.; Kefeli, J.; Patterson, J.; Srinivasan, A.; Zeitz, M.; Sirdeshmukh, G.; Berkowitz, J.; LaRow Brown, K.; Tatonetti, N. P.
Show abstract
Adverse drug events (ADEs) are the fourth leading cause of death in the US and cost billions of dollars annually in increased healthcare costs. However, few machine-readable databases of ADEs exist, limiting the opportunity to study drug safety on a broader, systematic scale. Recent advances in Natural Language Processing methods, such as BERT models, present an opportunity to accurately extract relevant information from unstructured biomedical text. As such, we fine-tuned a PubMedBERT model to extract ADE terms from descriptive text in FDA Structured Product Labels for prescription drugs. With this model, we achieve an F1 score of 0.90, AUROC of 0.92, and AUPR of 0.95 at extracting ADEs from the labels "Adverse Reactions". We further utilize this method to extract serious ADEs from labels "Boxed Warnings", and ADEs specifically noted for pediatric patients. Here, we present OnSIDES (ON-label SIDE effectS resource), a compiled, computable database of drug-ADE pairs generated with this method. OnSIDES contains more than 3.6 million drug-ADE pairs for 3,233 unique drug ingredient combinations extracted from 47,211 labels. Additionally, we expand this method to extract ADEs from drug labels of other major nations/regions - Japan, the UK, and the EU - to build a complementary OnSIDES-INTL database. To present potential applications, we used OnSIDES to predict novel drug targets and indications, analyze enrichment of ADEs across drug classes, and predict novel ADEs from chemical compound structures. We conclude that OnSIDES can be utilized as a comprehensive resource to study and enhance drug safety. One Sentence SummaryOnSIDES is a large, comprehensive database of adverse drug events extracted from drug labels using natural language processing methods.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Algorithmic identification of treatment-emergent adverse events from clinical notes using large language models: a pilot study in inflammatory bowel disease 94%
- DrugWAS: Leveraging drug-wide association studies to facilitate drug repurposing for COVID-19 93%
- Model-informed Deep Q-Networks to Guide Infliximab Dosing in Pediatric Crohn's Disease 91%
Similar papers in this journal
- Large-scale empirical identification of candidate comparators for pharmacoepidemiological studies 95%
- Standardization of drug names in the FDA Adverse Event Reporting System: The DiAna dictionary 94%
- Preventable deaths involving medicines in England and Wales, 2013-22: a systematic case series of coroners’ reports 92%
Similar papers in this journal
Similar papers in this journal
- Evaluating risk detection methods to uncover ontogenic-mediated adverse drug effect mechanisms in children 95%
- Changing word meanings in biomedical literature reveal pandemics and new technologies 93%
- Expanding a Database-derived Biomedical Knowledge Graph via Multi-relation Extraction from Biomedical Abstracts 92%
Similar papers in this journal
- A Framework to Quantify Disparities in Pharmacogenomic Treatment Concordance and Drug Response Outcomes 93%
- Assessing The Net Financial Benefits Of Employing Digital Endpoints In Clinical Trials 91%
- In silico PK predictions in Drug Discovery: Benchmarking of Strategies to Integrate Machine Learning with Empiric and Mechanistic PK modelling 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.