KRAKEN: A provenance-tracked knowledge graph for multiomic and wellness research
Glen, A. K.; Witherington, D.; Leslie, T.; Baumgartner, A.; Fernando, A.; Vemuri, B.; Nahman, O.; Glusman, G.; Hood, L.; Pflieger, L.; Rappaport, N.
Show abstract
Existing general-purpose biomedical knowledge graphs tend to focus on disease mechanisms and drug repurposing, leaving multiomic and wellness-relevant content underrepresented. KRAKEN (Knowledge Research & Analysis Kit for Evidence Networks) addresses this gap by integrating existing graphs (including Translator KG Open, RTX-KG2, and ROBOKOP) with specialized sources such as RefMet, LIPID MAPS, NIH Common Data Elements, Polygenic Score Catalog, and derived wellness measures including biological age and biological BMI. The resulting graph spans ~15M nodes and ~113M edges across 62 entity types. KRAKEN adopts the Biolink Model as its semantic layer, ensuring compatibility with standardized resources emerging from the NIH NCATS Biomedical Data Translator program. A lightweight, modular build system rebuilds the full graph (including entity resolution), with peak memory consumption under 48 GB, and supports flexible inclusion or exclusion of sources, allowing the user to scope the graph to a domain of interest. Built-in analytical tools include multi-hop reasoning, subgraph extraction, text, vector and hybrid entity search, and enrichment analyses, all accessible through an interactive web interface, a REST API, and a Model Context Protocol server, the last enabling direct consumption by agentic and LLM-based systems. KRAKEN is freely available at https://app.krakenkg.com.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CROssBAR: Comprehensive Resource of Biomedical Relations with Deep Learning Applications and Knowledge Graph Representations 92%
- FunCoup 6: advancing functional association networks across species with directed links and improved user experience 92%
- Integration of the Drug-Gene Interaction Database (DGIdb) with open crowdsource efforts 91%
Similar papers in this journal
Similar papers in this journal
- DISEASES 2.0: a weekly updated database of disease-gene associations from text mining and data integration 92%
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 92%
- LSD600: the first corpus of biomedical abstracts annotated with lifestyle–disease relations 92%
Similar papers in this journal
- Orchestrating and sharing large multimodal data for transparent and reproducible research 92%
- A Platform for Oncogenomic Reporting and Interpretation 92%
- A user's guide to the online resources for data exploration, visualization, and discovery for the Pan-Cancer Analysis of Whole Genomes project (PCAWG) 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.