Back

Dug: A Semantic Search Engine Leveraging Peer-Reviewed Literature to Span Biomedical Data Repositories

Waldrop, A. M.; Cheadle, J. B.; Bradford, K.; Braswell, N. T.; Watson, M.; Crerar, A.; Ball, C.; Kebede, Y.; Schreep, C.; Linebaugh, P. J.; Hiles, H.; Boyles, R.; Bizon, C.; Krishnamurthy, A.; Cox, S.

2021-07-09 bioinformatics
10.1101/2021.07.07.451461 bioRxiv
Show abstract

MotivationAs the number of public data resources continues to proliferate, identifying relevant datasets across heterogenous repositories is becoming critical to answering scientific questions. To help researchers navigate this data landscape, we developed Dug: a semantic search tool for biomedical datasets utilizing evidence-based relationships from curated knowledge graphs to find relevant datasets and explain why those results are returned. ResultsDeveloped through the National Heart, Lung, and Blood Institutes (NHLBI) BioData Catalyst ecosystem, Dug has indexed more than 15,911 study variables from public datasets. On a manually curated search dataset, Dugs total recall (total relevant results/total results) of 0.79 outperformed default Elasticsearchs total recall of 0.76. When using synonyms or related concepts as search queries, Dug (0.36) far outperformed Elasticsearch (0.14) in terms of total recall with no significant loss in the precision of its top results. Availability and ImplementationDug is freely available at https://github.com/helxplatform/dug. An example Dug deployment is also available for use at https://search.biodatacatalyst.renci.org/. Contactawaldrop@rti.org or scox@renci.org

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.