Beyond associations: A benchmark Causal Relation Extraction Dataset (CRED) of disease-causing genes, its comparative evaluation, interpretation and application
Bansal, N.; R C, S. D.; Pathak, A.; Narayanan, M.
Show abstract
Information on causal relationships is essential to many sciences, including biomedical science, and beneficial (e.g., causative rather than merely associative gene-disease relations can lead to better treatments). Despite much work on Relation Extraction (RE), automatically extracting causal relations from large text corpora remains less explored. Few existing studies on CRE (Causal RE) are limited to extracting causality within a sentence or for a particular disease, mainly due to the lack of a diverse benchmark dataset. Here, we carefully curate a new CRE Dataset (CRED) of 3639 (causal and non-causal) gene-disease pairs, spanning 204 diseases and 500 genes, within or across sentences of 267 published abstracts. CRED is assembled in two phases to reduce class imbalance, and its inter-annotator agreement is 89%. To assess CREDs utility in classifying causal vs. non-causal pairs, we compared multiple classifiers and found SVM (Support Vector Machine) trained on embeddings from a deep learning transformer model called BioBERT to perform the best (F1 score 0.70). CRED outperformed a state-of-the-art RE dataset in terms of classifier performance and model interpretability, i.e., whether the model focuses importance/attention on words with causal connotations in abstracts. Moving from benchmark to real-world settings, application of our CRED-trained BioBERT+SVM model on all PubMed abstracts on Parkinsons disease (PD) revealed both well- and less-studied PD-causing genes. For instance, genes predicted to be causal for PD in at least 50 abstracts by our model were already linked to PD in books; and lends confidence to further explore the other genes predicted to be causal in fewer abstracts. Our systematically curated and evaluated CRED, and its associated classification model and gene-disease causality scores, thus offer concrete resources for advancing future research in CRE from biomedical literature.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A Sequence Labeling Framework for Extracting Drug-Protein Relations from Biomedical Literature 97%
- DISEASES 2.0: a weekly updated database of disease-gene associations from text mining and data integration 96%
- SynLethDB 2.0: A web-based knowledge graph database on synthetic lethality for novel anticancer drug discovery 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Interpretable and Generalizable Attention-Based Model for Predicting Drug-Target Interaction Using 3D Structure of Protein Binding Sites: SARS-CoV-2 Case Study and in-Lab Validation 94%
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 94%
- GexMolGen: Cross-modal Generation of Hit-like Molecules via Large Language Model Encoding of Gene Expression Signatures 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.