GeDex: A consensus Gene-disease Event Extraction System based on frequency patterns and supervised learning
Olayo-Alarcon, R.; Morales-Soto, L.; Velazquez-Ramirez, D. A.; Munguia-Reyes, A.; Balderas-Martinez, Y. I.; Mendez-Cruz, C. F.; Collado-Vides, J.
Show abstract
MotivationThe genetic mechanisms involved in human diseases are fundamental in biomedical research. Several databases with curated associations between genes and diseases have emerged in the last decades. Although, due to the demanding and time consuming nature of manual curation of literature, they still lack large amounts of information. Current automatic approaches extract associations by considering each abstract or sentence independently. This approach could potentially lead to contradictions between individual cases. Therefore, there is a current need for automatic strategies that can provide a literature consensus of gene-disease associations, and are not prone to making contradictory predictions. ResultsHere, we present GeDex, an effective and freely available automatic approach to extract consensus gene-disease associations from biomedical literature based on a predictive model trained with four simple features. As far as we know, it is the only system that reports a single consensus prediction from multiple sentences supporting the same association. We tested our approach on the curated fraction of DisGeNet (f-score 0.77) and validated it on a manually curated dataset, obtaining a competitive performance when compared to pre-existing methods (f-score 0.74). In addition, we effectively recovered associations from an article collection of chronic pulmonary diseases, and discovered that a large proportion is not reported in current databases. Our results demonstrate that GeDex, despite its simplicity, is a competitive tool that can successfully assist the curation of existing databases. AvailabilityGeDex is available at https://bitbucket.org/laigen/gedex/src/master/ and can be used as a docker image https://hub.docker.com/r/laigen/gedex Contactcmendezc@ccg.unam.mx Supplementary informationSupplementary material are available at bioRxiv online.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DISEASES 2.0: a weekly updated database of disease-gene associations from text mining and data integration 96%
- LSD600: the first corpus of biomedical abstracts annotated with lifestyle–disease relations 96%
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 95%
Similar papers in this journal
- FunCoup 6: advancing functional association networks across species with directed links and improved user experience 94%
- CROssBAR: Comprehensive Resource of Biomedical Relations with Deep Learning Applications and Knowledge Graph Representations 94%
- Assessing the impact of transcriptomics data analysis pipelines on downstream functional enrichment results 92%
Similar papers in this journal
- SKiM-GPT: Combining Biomedical Literature-Based Discovery with Large Language Model Hypothesis Evaluation 96%
- Using BERT to identify drug-target interactions from whole PubMed 93%
- Towards a standard benchmark for phenotype-driven variant and gene prioritisation algorithms: PhEval - Phenotypic inference Evaluation framework 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.