Finding Case Report Nuggets: A Web-Based Tool for Mining and Enhancing the Value of Clinical Case Reports
Holt, A. W.; Smalheiser, N. R.
Show abstract
ObjectivesCase reports are eyewitness reports of medical phenomena, such as adverse effects of treatments, outcomes of new surgical techniques, descriptions of rare diseases, unusual presentations of common diseases, or emerging infectious outbreaks. Although any single case report may be confounded, biased or erroneous, observations that are separately reported in multiple independent publications are more likely to be reliable, and so the accumulated evidence should have more value than any single report on its own. This notion led us to analyze the case reports literature in search of nuggets: collections of multiple case reports that describe similar main findings. Materials and MethodsTo identify nuggets in collections of case reports retrieved in PubMed queries, semantic similarities among the case reports were computed based on titles and main finding sentences extracted from the abstracts, and then grouped into communities with a graph database. The initial communities were then merged with a secondary hierarchical clustering process. ResultsComputed nuggets of size 4-100 articles are displayed along with large language model (LLM)-computed summaries, the title of the nuggets central article, and hyperlinks for viewing as well as export to our companion tool Anne OTate for further analysis. A variety of advanced options are also offered; users can optionally submit feedback on the quality of computed nuggets. DiscussionOur free, public tool https://arrowsmith.psych.uic.edu/casereports facilitates the identification of nuggets and their summarization and mining. This should enhance the value of case report evidence and assist clinicians as well as those performing evidence syntheses of the published literature.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 93%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 93%
- Predicting Emerging Themes in Rapidly Expanding COVID-19 Literature with Dynamic Word Embedding Networks and Machine Learning 91%
Similar papers in this journal
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 94%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 94%
- An interactive retrieval system for clinical trial studies with context-dependent protocol elements 94%
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 93%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 93%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.