The Data Distillery: A Graph Framework for Semantic Integration and Querying of Biomedical Data
Mohseni Ahooyi, T.; Stear, B.; Simmons, J. A.; Metzger, V. T.; Kumar, P.; Evangelista, J. E.; Clarke, D. J. B.; Xie, Z.; Kim, H.; Jenkins, S. L.; Maurya, M. R.; Ramachandran, S.; Fahy, E.; Imam, F. T.; Kokash, N.; Roth, M. E.; Fullem, R.; Jevtic, D.; Mihajlovic, A.; Tiemeyer, M.; Gillespie, T. H.; Bakker, C.; Schroeder, A. J.; Markowski, J.; Nedzel, J.; Hill, D. D.; Terry, J.; Nemarich, C.; Park, P.; Ardlie, K. G.; Vora, J.; Mazumder, R.; Ranzinger, R.; de Bono, B.; Subramaniam, S.; Grethe, J. S.; Yang, J. J.; Lambert, C. G.; Resnick, A.; Milosavljevic, A.; Ma'ayan, A.; Silverstein, J. C.; Tay
Show abstract
The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CROssBAR: Comprehensive Resource of Biomedical Relations with Deep Learning Applications and Knowledge Graph Representations 95%
- FunCoup 6: advancing functional association networks across species with directed links and improved user experience 94%
- Echtvar: Compressed variant representation for rapid annotation and filtering of SNPs and indels 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.