Back

Methodological Issues with Search in MEDLINE: A Longitudinal Query Analysis

Burns, C. S.; Nix, T.; Shapiro, R.; Huber, J.

2020-05-22 bioinformatics
10.1101/2020.05.22.110403 bioRxiv
Show abstract

This study compares the results of data collected from a longitudinal query analysis of the MEDLINE database hosted on multiple platforms that include PubMed, EBSCOHost, Ovid, ProQuest, and Web of Science in order to identify variations among the search results on the platforms after controlling for search query syntax. We devised twenty-nine sets of search queries comprised of five queries per set to search against the five MEDLINE database platforms. We ran our queries monthly for a year and collected search result count data to observe changes. We found that search results vary considerably depending on MEDLINE platform, both within sets and across time. The variation is due to trends in scholarly publication that include publishing online first versus publishing in journal issues, which leads to metadata differences in the bibliographic record; to differences in the level of specificity among search fields provided by the platforms; to database integrity issues that lead to large fluctuations in monthly search results based on the same query; and to database currency issues that arise due to when each platform updates its MEDLINE file. Specific bibliographic databases, like PubMed and MEDLINE, are used to inform clinical decision-making, create systematic reviews, and construct knowledge bases for clinical decision support systems. Since they serve as essential information retrieval and discovery tools that help identify and collect research data and are used in a broad range of fields and as the basis of multiple research designs, this study should help clinicians, researcher, librarians, informationalists, and others understand how these platforms differ and inform future work in their standardization.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.