VaxKG: Integrating The Vaccine Ontology And VIOLIN For Advanced Vaccine Queries And LLM-Powered Chat Systems
Yeh, F.-Y.; Asato, M.; Zheng, J.; He, Y.
Show abstract
Vaccine research faces challenges in integrating diverse biomedical datasets. While the Vaccine Investigation and Online Information Network (VIOLIN) provides comprehensive vaccine data, implemented in traditional relational models limit complex analysis. Similarly, the Vaccine Ontology (VO) offers standardized semantic frameworks but lacks comprehensive empirical data. This study addresses these limitations by developing the Vaccine Knowledge Graph (VaxKG) that integrates VIOLINs dataset with VOs standardized terminology. Using Neo4j, we transformed 12 core VIOLIN tables into a graph structure enriched with VO concepts. The resulting knowledge graph comprises 28,123 VIOLIN data nodes and 101,282 VO resource nodes, connected by 412,865 relationships. Our comparative analysis of Brucella and Influenza vaccines demonstrates VaxKGs ability to enable complex semantic queries and reveal insights unavailable from either resource alone. We further demonstrate VaxKGs utility through VaxChat, a large language model application that leverages the VaxKG as Retrieval-Augmented Generation (RAG) for intuitive vaccine information access.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Knowledge Graph-based Thought: a knowledge graph enhanced LLMs framework for pan-cancer question answering 94%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 93%
- Health Data Nexus: An Open Data Platform for AI Research and Education in Medicine 93%
Similar papers in this journal
- Extract, Transform, Load Framework for the Conversion of Health Databases to OMOP 94%
- Datavzrd: Rapid programming- and maintenance-free interactive visualization and communication of tabular data 94%
- Understanding signaling and metabolic paths using semantified and harmonized information about biological interactions 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.