Wikipedia Drug Safety Advisory Committee: Distilling a Drug Adverse Effect Reference Set Using Wisdom of the Crowd
Bilu, Y.; Yanover, C.
Show abstract
BackgroundLarge datasets of relational medical data, such as the adverse effects of drugs or vaccines, typically attain their large size, by relying on automatic, or semi-automatic, methods for generation. This often comes with a compromise on the precision of the generated data, which can be at least partially alleviated by having experts curate the data. AimSince having experts review a large dataset can be costly and time consuming, here we suggest using Wikipedia for this task - that is, augment the automatic generation step by an automatic curation step based on the expert knowledge accumulated in Wikipedia. MethodsTo curate a dataset of adverse drug effects (ADEs), we suggest retrieving the Wikipedia page associated with the drug, and checking whether the ADE appears in the sections describing adverse effects. Drug indications, typically described in the opening paragraph of the page, are similarly filtered out. ResultsWe use the method to curate two large adverse drug effect datasets and show that the obtained datasets have a much higher precision relative to their originating ones. ConclusionsAlgorithms which aim to infer drug-ADE relations, should at the very least be able to identify the "clear cut" cases. The high-precision benchmark constructed herein may therefore be a valuable resource for the evaluation of such algorithms.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- KG-Bench: Benchmarking Graph Neural Network Algorithms for Drug Repurposing 95%
- Design and application of a knowledge network for automatic prioritization of drug mechanisms 95%
- DTI-Voodoo: machine learning over interaction networks and ontology-based background knowledge predicts drug-target interactions 95%
Similar papers in this journal
- A Scoping Review of Artificial Intelligence Applications in Clinical Trial Risk Assessment 94%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 94%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.