Back

Beyond Identifier Matching: An Empirical Characterization of Failure Modes in Biomedical Knowledge Graph Integration

Hu, S.; Cheng, H.; Gillenwater, L.; Manpearl, K.; Mandava, A.; Wang, Y.; Pividori, M.; Stranger, B.; Krishnan, A.; Greene, C.; Gao, Y.

2026-05-28 health informatics
10.64898/2026.05.26.26354182 medRxiv
Show abstract

Objective. Biomedical knowledge graphs (KGs) such as PrimeKG, Hetionet, UMLS, and PharmGKB are increasingly used as the substrate for downstream machine-learning, retrieval-augmented generation, drug-repurposing, and electronic health record (EHR) augmentation pipelines. The dominant assumption in published work is that integrating two or more such KGs is a tractable engineering step solved by identifier (ID) matching. This paper interrogates that assumption empirically. We quantify how much concept overlap survives realistic alignment, and we characterize the new failure modes introduced by the methods that practitioners reach for when ID matching is insufficient. Materials and Methods. We compared four widely used biomedical KGs (PrimeKG, Hetionet v1.0, the full UMLS Metathesaurus, and PharmGKB) across eleven node types using a tiered alignment pipeline: (1) direct ID matching for nodes sharing a primary vocabulary; (2) cross-ontology bridging using standard mappings (e.g., MONDO-DOID, HPO-UMLS, HPO-UMLS-MeSH for side effects, NCBI Gene-HGNC-UMLS, UBERON-FMA/SNOMEDCT_US/NCI/MeSH for anatomy); (3) ClinicalBERT cosine-similarity grouping at threshold >= 0.98 for over-segmented disease nodes, with a deterministic suffix-stripping canonicalizer; (4) exact name matching for ontology-poor types (anatomy, REACTOME pathways); and (5) embedding-based fuzzy matching with UMLS lookup (SapBERT and ClinicalBERT) for free-text microbiome concepts. We applied the pipeline to a 698-concept gut-microbiome benchmark spanning taxa, pathways, and disease labels, validated grouping decisions against the curated SSSOM mappings released by the MONDO project, and audited the ClinicalBERT consolidation against five clinical-genetics case studies drawn from the literature. Results. Per-type pairwise coverage was strikingly asymmetric. Genes/proteins and the three Gene Ontology categories aligned cleanly across PrimeKG and Hetionet (mutual coverage 94-99%), but disease overlap was sparse: only 0.7% of PrimeKG individual disease nodes mapped to Hetionet, rising to 2.0% after MONDO grouping (versus 78.7% and 18.4% from the Hetionet side). PrimeKG-to-UMLS coverage spanned 100% (effect/phenotype via HPO) down to 20.8% (REACTOME pathways), with drugs at 73.7% and anatomy at 58.8%. PrimeKG-to-PharmGKB drug coverage required up to two bridging hops (DrugBank -> UMLS -> RxNorm/ATC/MeSH). Bigger was not uniformly more complete: on a 698-concept microbiome drug benchmark, Hetionet missed 0 concepts while PrimeKG missed 16. ClinicalBERT-based grouping consolidated 22,205 raw MONDO disease nodes into 17,080 groups but introduced three reproducible failure modes documented in case studies: (i) peer over-merging: for example, all 22 osteogenesis imperfecta subtypes collapsed into a single node despite distinct severity classes; (ii) parent-child collapse: e.g. acute myeloid leukemia merged with myeloid leukemia, erasing the acute/chronic distinction that drives clinical management; and (iii) lexical false positives: neurofibromatosis and schwannomatosis grouped together despite cellular-pathology differences. Discussion. Identifier matching alone is a weak baseline for biomedical KG integration. Cross-ontology bridges and embedding-based consolidation expand coverage but do so at the cost of clinically meaningful resolution, and the resulting failures are systematic rather than random. Reporting only aggregate coverage statistics obscures these losses, which propagate silently into downstream tasks. Conclusion. We provide reusable per-type coverage tables, a taxonomy of three integration failure modes, and concrete recommendations for downstream studies that depend on a unified biomedical KG. We argue that future KG integration work should report per-type coverage and per-cluster confidence rather than aggregate match rates.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 3%
7.7%
2
Nature Medicine
125 papers in training set
Top 0.2%
7.1%
3
GENETICS
483 papers in training set
Top 1%
5.4%
4
Genome Biology
637 papers in training set
Top 2%
5.1%
5
Frontiers in Bioinformatics
49 papers in training set
Top 0.1%
4.8%
6
Communications Medicine
113 papers in training set
Top 0.6%
4.2%
7
Nature Communications
5641 papers in training set
Top 31%
4.2%
8
eBioMedicine
183 papers in training set
Top 0.7%
4.0%
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.7%
10
PLOS ONE
5266 papers in training set
Top 41%
2.6%
11
Cell Genomics
172 papers in training set
Top 2%
2.4%
50% of probability mass above
12
npj Digital Medicine
118 papers in training set
Top 2%
2.3%
13
Scientific Reports
3612 papers in training set
Top 45%
2.3%
14
Journal of the American Medical Informatics Association
71 papers in training set
Top 1%
2.3%
15
iScience
1154 papers in training set
Top 14%
1.9%
16
JAMIA Open
42 papers in training set
Top 0.8%
1.9%
17
Med
39 papers in training set
Top 0.2%
1.9%
18
JMIR Medical Informatics
18 papers in training set
Top 0.5%
1.7%
19
BMC Medical Informatics and Decision Making
43 papers in training set
Top 1%
1.7%
20
GigaScience
212 papers in training set
Top 2%
1.7%
21
PLOS Computational Biology
1863 papers in training set
Top 15%
1.5%
22
Journal of Biomedical Informatics
47 papers in training set
Top 0.9%
1.3%
23
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
24
Genetics in Medicine
78 papers in training set
Top 0.8%
1.1%
25
BMC Medical Genomics
50 papers in training set
Top 0.9%
1.1%
26
Database
61 papers in training set
Top 0.7%
1.1%
27
Scientific Data
209 papers in training set
Top 2%
1.1%
28
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.1%
29
Patterns
78 papers in training set
Top 2%
1.0%
30
npj Antimicrobials and Resistance
11 papers in training set
Top 0.2%
1.0%