Building a literature knowledge base towards transparent biomedical AI
Huang, Y.; Han, Z.; Luo, X.; Luo, X.; Gao, Y.; Zhao, M.; Tang, F.; Wang, Y.; Chen, J.; Li, C.; Lu, X.; Qiu, J.; Deng, F.; Jiao, T.; Xue, D.; Feng, F.; Vu, T. H. H.; Guan, L.; Cartailler, J.-P.; Stitzel, M.; Chen, S.; Brissova, M.; Parker, S.; Liu, J.
Show abstract
As artificial intelligence (AI) continues to advance and scale up in biomedical research, concerns about AIs trustworthiness and transparency have grown. There is a critical need to systematically bring accurate and relevant biomedical knowledge into AI applications for transparency and provenance. Knowledge graphs have emerged as a powerful tool that integrates heterogeneous knowledge by explicitly describing biomedical knowledge as entities and relationships between entities. However, PubMed, the largest and most comprehensive repository of biomedical knowledge, exists primarily as unstructured text and is under utilized for advanced machine learning tasks. To address the challenge, we developed LiteralGraph, a computational framework to extract biomedical terms and relationships from PubMed literature into a unified knowledge graph. Using this framework, we established the Genomic Literature Knowledge Base (GLKB), which consolidates 14,634,427 biomedical relationships between 3,276,336 biomedical terms from over 33 million PubMed abstracts and nine well-established biomedical repositories. The database is coupled with RESTful APIs and a user-friendly web interface that makes it accessible to researchers for various usages. We demonstrated the broad utility of GLKB towards transparent AI in three distinct application scenarios. In the LLM grounding scenario, we developed a Retrieval Augmented Generation (RAG) agent to reduce LLM hallucination in biomedical question answering. In the hypothesis generation scenario, we elucidated the potential functions of RFX6 in type 2 diabetes (T2D) using the vast evidence from PubMed articles. In the machine learning scenario, we utilized GLKB to provide semantic knowledge in predictive tasks and scientific fact-checking.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DISEASES 2.0: a weekly updated database of disease-gene associations from text mining and data integration 96%
- SynLethDB 2.0: A web-based knowledge graph database on synthetic lethality for novel anticancer drug discovery 96%
- A Sequence Labeling Framework for Extracting Drug-Protein Relations from Biomedical Literature 96%
Similar papers in this journal
- ARAX: a graph-based modular reasoning tool for translational biomedicine 96%
- FORUM: Building a Knowledge Graph from public databases and scientific literature to extract associations between chemicals and diseases 95%
- BioMedGraphica: An All-in-One Platform for Biomedical Prior Knowledge and Omic Signaling Graph Generation 95%
Similar papers in this journal
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 96%
- Gilda: biomedical entity text normalization with machine-learned disambiguation as a service 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.