AI semantics for biomedical data integration
McLaughlin, J.; Puig-Barbe, A.; Ibrahim, A.; Pava, D.; Pendlington, Z. M.; Matentzoglu, N.; Sollis, E.; Foreman, A.; Wilson, R.; Lopez Gomez, F.; Harris, L.; Adeleye, Y.; Kaur, S.; Meldal, B.; Smedley, D.; Parkinson, H.
Show abstract
Researchers increasingly need to explore hypotheses that span multimodal data across different scales, organisms, and domains. In practice, this requires connecting knowledge across fragmented databases with incompatible APIs and heterogeneous annotation practices. Large language model (LLM) agents can automate this data integration process, but grounding LLM agent outputs in scientifically correct sources of truth remains a significant challenge. Here we describe our deployment of a novel AI semantics workflow using LLM agents to enable scalable data integration, grounded in biological knowledge in the form of ontologies. Our workflow comprises (1) a multi-agent system curating scientific knowledge across ontologies using the Ontology Lookup Service (OLS) as grounding; (2) an LLM embedding service to enable interoperability between scientific databases by mapping ontology terms; and (3) GrEBI, a knowledge graph and Model Context Protocol (MCP) server enabling LLM agents to conduct cross-cutting, multi-omic biomedical queries. O_FIG O_LINKSMALLFIG WIDTH=187 HEIGHT=200 SRC="FIGDIR/small/742514v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@1fc525aorg.highwire.dtl.DTLVardef@82c4f7org.highwire.dtl.DTLVardef@15173fdorg.highwire.dtl.DTLVardef@961f1c_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Echtvar: Compressed variant representation for rapid annotation and filtering of SNPs and indels 90%
- Single-Cell Trajectory Inference for Detecting Transient Events in Biological Processes 90%
- FunCoup 6: advancing functional association networks across species with directed links and improved user experience 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.