Back

Wikipedia Drug Safety Advisory Committee: Distilling a Drug Adverse Effect Reference Set Using Wisdom of the Crowd

Bilu, Y.; Yanover, C.

2021-11-01 health informatics
10.1101/2021.11.01.21265741 medRxiv
Show abstract

BackgroundLarge datasets of relational medical data, such as the adverse effects of drugs or vaccines, typically attain their large size, by relying on automatic, or semi-automatic, methods for generation. This often comes with a compromise on the precision of the generated data, which can be at least partially alleviated by having experts curate the data. AimSince having experts review a large dataset can be costly and time consuming, here we suggest using Wikipedia for this task - that is, augment the automatic generation step by an automatic curation step based on the expert knowledge accumulated in Wikipedia. MethodsTo curate a dataset of adverse drug effects (ADEs), we suggest retrieving the Wikipedia page associated with the drug, and checking whether the ADE appears in the sections describing adverse effects. Drug indications, typically described in the opening paragraph of the page, are similarly filtered out. ResultsWe use the method to curate two large adverse drug effect datasets and show that the obtained datasets have a much higher precision relative to their originating ones. ConclusionsAlgorithms which aim to infer drug-ADE relations, should at the very least be able to identify the "clear cut" cases. The high-precision benchmark constructed herein may therefore be a valuable resource for the evaluation of such algorithms.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.