Back

Finding Case Report Nuggets: A Web-Based Tool for Mining and Enhancing the Value of Clinical Case Reports

Holt, A. W.; Smalheiser, N. R.

2025-11-15 health informatics
10.1101/2025.11.13.25340162 medRxiv
Show abstract

ObjectivesCase reports are eyewitness reports of medical phenomena, such as adverse effects of treatments, outcomes of new surgical techniques, descriptions of rare diseases, unusual presentations of common diseases, or emerging infectious outbreaks. Although any single case report may be confounded, biased or erroneous, observations that are separately reported in multiple independent publications are more likely to be reliable, and so the accumulated evidence should have more value than any single report on its own. This notion led us to analyze the case reports literature in search of nuggets: collections of multiple case reports that describe similar main findings. Materials and MethodsTo identify nuggets in collections of case reports retrieved in PubMed queries, semantic similarities among the case reports were computed based on titles and main finding sentences extracted from the abstracts, and then grouped into communities with a graph database. The initial communities were then merged with a secondary hierarchical clustering process. ResultsComputed nuggets of size 4-100 articles are displayed along with large language model (LLM)-computed summaries, the title of the nuggets central article, and hyperlinks for viewing as well as export to our companion tool Anne OTate for further analysis. A variety of advanced options are also offered; users can optionally submit feedback on the quality of computed nuggets. DiscussionOur free, public tool https://arrowsmith.psych.uic.edu/casereports facilitates the identification of nuggets and their summarization and mining. This should enhance the value of case report evidence and assist clinicians as well as those performing evidence syntheses of the published literature.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.