Back

Deepgeosearch: Llm-Powered Schemaless Retrievalfor Biomedical Data Discovery

Singh, D.; Jatav, S.; Lathwal, S.; Kazmi, A.; Luthra, S.

2025-12-27 bioinformatics
10.64898/2025.12.27.696662 bioRxiv
Show abstract

Public biomedical repositories contain extensive, valuable datasets, yet identifying datasets that precisely match specific study requirements remains inefficient. Conventional keyword- and schema-based systems frequently fall short when queries encompass multiple biological and experimental facets. To address this, we developed DeepGEOSearch, an LLM-powered, schema-less retrieval system that interprets dataset text directly rather than relying solely on predefined metadata fields. The system extracts key study attributes, harmonizes terminology across sources, and ranks datasets by contextual alignment with natural-language queries while providing verifiable evidence for each match. In this work, we applied DeepGEOSearch for the Gene Expression Omnibus (GEO). The framework is repository-agnostic and can integrate new metadata sources without manual relabeling. This design enables complex compositional queries that existing tools cannot support. In evaluations against strong baselines using a curated evaluation benchmark comprising a mix of query complexities, DeepGEOSearch achieved more than 90% precision and best recall, with the largest performance gains observed on complex, real-world queries. DeepGEOSearch consistently identifies relevant datasets overlooked by conventional search tools, accelerating dataset discovery and improving reuse of public biomedical data.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.