BiomarkerKB: FAIR and Integrated Biomarker Knowledge Connecting Biomolecular and Clinical Data Types
Masood, D.; Kim, M.; Vora, J.; Kahsay, R.; McNeeley, P.; Kim, S.; Kulkarni, S. V.; Natale, D. A.; Maurya, M.; Ramachandran, S.; Gupta, S.; Bologa, C. G.; DeNapoli, T. S.; Metzger, V. T.; Kumar, P.; Ahmed, N.; Evangelista, J. E.; Kelly, S. C.; Sepulveda, J.; Ma'ayan, A.; Silverstein, J.; Taylor, D. M.; Crichton, D. J.; Mahabal, A.; Yang, J. J.; Lambert, C. G.; Subramaniam, S.; Tiemeyer, M.; Ranzinger, R.; Mazumder, R.
Show abstract
Biomarkers are essential tools for disease detection, risk assessment, therapeutic monitoring, and precision medicine. However, biomarker data are dispersed across heterogeneous resources, inconsistently reported in the literature, and rarely standardized for computational use. This fragmentation limits reproducibility, cross-study integration, and the discovery of novel biomarker and disease relationships. We developed BiomarkerKB, a knowledgebase designed to harmonize and integrate biomarker information under a standardized data model. The model follows the FDA-NIH BEST biomarker definition and captures both core fields (biomarker entity, disease/condition, exposure agent) and contextual metadata (specimen, biomarker role, evidence, provenance). Biomarker data and related annotations were either curated from publications or collected from public resources (e.g., OpenTargets, GWAS Catalog, ClinVar, CIViC, OncoMX) and were also contributed by the Common Fund Data Coordinating Centers and the Early Detection Research Network (EDRN). Standardization was achieved using ontologies and reference resources such as Disease Ontology, UBERON, UniProtKB, and HUGO Gene Nomenclature Committee (HGNC) gene symbols. BiomarkerKB data were ingested into a Neo4j-based knowledge graph and integrated with the Common Fund Data Ecosystem (CFDE) Knowledge Graph. The initial release of BiomarkerKB contains over 200,000 biomarker-disease associations spanning genes, proteins, metabolites, glycans, and chemical elements. The knowledge graph comprises more than 300,000 nodes and 1.2 million edges, enabling structured exploration of biomarker relationships within CFDE data as demonstrated through the knowledge graph query-based use cases presented in this study. A publicly accessible web portal (https://biomarkerkb.org) provides keyword search, filtering, data downloads, and access to graph visualization to support both researchers and computational analyses. BiomarkerKB addresses a critical gap in biomarker informatics by providing an integrated, FAIR (Findable, Accessible, Interoperable, and Reusable), and unified framework for biomarker knowledge exploration and discovery.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- OmicsFootPrint: a framework to integrate and interpret multi-omics data using circular images and deep neural networks 94%
- CROssBAR: Comprehensive Resource of Biomedical Relations with Deep Learning Applications and Knowledge Graph Representations 93%
- MetaOmGraph: a workbench for interactive exploratory data analysis of large expression datasets 92%
Similar papers in this journal
- Using semantic search to find publicly available gene-expression datasets 93%
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 93%
- FORUM: Building a Knowledge Graph from public databases and scientific literature to extract associations between chemicals and diseases 93%
Similar papers in this journal
- Strategies and Techniques for Quality Control and Semantic Enrichment with Multimodal Data: A Case Study in Colorectal Cancer with eHDPrep 94%
- New implementation of data standards for AI research in precision oncology. Experience from EuCanImage 93%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.