Scaling Biomedical Knowledge Graph Retrieval for Interpretable Reasoning: Applications to Clinical Diagnosis Prediction
Cheng, H.; Wu, Y.; Khatwani, S.; Kruse, M.; Dligach, D.; Miller, T.; Afshar, M.; Gao, Y.
Show abstract
Biomedical knowledge graphs (KGs) organize molecular mechanisms, biological pathways, and clinical concepts into structured representations that support diagnostic reasoning. As these graphs grow in scale and connectivity, scalable and interpretable multi-hop retrieval over deep biomedical graph structures has become a major computational bottleneck. We present LogosKG, a hardware-optimized retrieval system that enables efficient k-hop traversal over very large biomedical KGs using symbolic graph formulations and hardwareefficient execution. By integrating degreeaware partitioning, cross-partition routing, and on-demand caching, LogosKG scales to billionedge graphs while preserving retrieval fidelity. Experiments demonstrate substantial efficiency improvements over CPU- and GPU-based baselines. Using diagnosis-oriented retrieval workloads as a downstream case study, we show that scalable access to deep, high-hop biomedical graph structures enables interpretable diagnostic evidence propagation. A clinician-aligned LLM-as-judge evaluation further indicates that high-hop KG retrieval improves reasoning quality in terms of accuracy, comprehensibility, and succinctness, underscoring the value of deep graph retrieval for diagnostic reasoning.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Causal Analysis for Multivariate Integrated Clinical and Environmental Exposures Data 92%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 92%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.