Back

MR-KG: A knowledge graph of Mendelian randomization evidence powered by large language models

Liu, Y.; Burton, J.; Gatua, W.; Hemani, G.; Gaunt, T. R.

2025-12-15 health informatics
10.64898/2025.12.14.25342218 medRxiv
Show abstract

BackgroundThe exponential growth of Mendelian randomization (MR) literature has created challenges for systematically organising and synthesising evidence, with key information fragmented across heterogeneous publications. We present MR-KG, a knowledge graph resource using large language models (LLMs) to systematically extract and structure published MR evidence at scale. MethodsWe evaluated eight OpenAI and local LLMs for extracting structured information from MR study abstracts. Two reviewers independently assessed extraction quality across 100 randomly selected studies. We applied the pipeline to over 15,000 MR studies published between 2003 and 2025, implementing trait profile similarity matching to identify related research questions and evidence profile similarity matching to assess concordance across studies. ResultsThe LLM extractors achieved high scores across all assessment dimensions, and the resulting MR-KG resource provides comprehensive coverage of the MR literature. Semantic similarity matching revealed that whilst researchers explore conceptually related domains, they typically examine distinct exposure-outcome pairs. Temporal analysis documented substantial increases in trait diversity per study and significant improvements in reporting completeness following STROBE-MR guidelines. Reproducibility analysis showed that whilst a majority of replicated trait pairs achieve high concordance, a non-trivial minority showed discordant results, with substantial domain-specific variation. In addition, semantic matching quality varied substantially across disease categories. ConclusionsLLM-based extraction can address information overload in MR research by extracting structured, queryable knowledge from fragmented literature. MR-KG enables systematic evidence synthesis, reveals domain-specific reproducibility patterns, and provides a continuously updated resource for the research community. This approach is generalisable to other research domains facing similar challenges.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.