Back

CONORM: Context-Aware Entity Normalization for Adverse Drug Event Detection

Yazdani, A.; Rouhizadeh, H.; Bornet, A.; Teodoro, D.

2023-09-26 health informatics
10.1101/2023.09.26.23296150 medRxiv
Show abstract

Adverse drug events (ADEs) are a critical aspect of patient safety and pharmacovigilance, with significant implications for patient outcomes and public health monitoring. The increasing availability of electronic health records, social media, and online patient forums provides valuable yet challenging unstructured data sources for ADE surveillance. To address these challenges, we introduce CONORM, a novel framework integrating named entity recognition (NER) and entity normalization (EN) for ADE resolution across diverse textual domains. CONORM comprises CONORM-NER and CONORM-EN, featuring a dual-encoder architecture with dynamic context refining (DCR). The DCR mechanism adaptively combines isolated entity embeddings with contextual representations. Our analyses demonstrate this approach effectively adjusts model behavior according to text formality, enhances precision on out-of-distribution concepts, and substantially reduces normalization errors compared to context-agnostic baselines. CONORM was evaluated on tweets, forum posts, and structured product labels, achieving end-to-end F1-scores of 63.86%, 72.45%, and 84.99%, respectively, surpassing existing solutions by an average margin of 35%. These results highlight CONORMs robust adaptability across domains, enabled by DCRs effective context utilization. CONORM offers a scalable, reproducible solution for pharmacovigilance, with pre-computed target embeddings enhancing inference efficiency. Its generalization establishes it as a robust tool for ADE surveillance. Source code is publicly available at https://github.com/ds4dh/CONORM.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.