Back

Creating an Indexing Scheme for Case Series Articles

Shahidehpour, A.; Holt, A. W.; Troy, A. M.; Menke, J. D.; Smalheiser, N. R.

2025-12-29 health informatics
10.64898/2025.12.19.25342712 medRxiv
Show abstract

ObjectivesCase reports and case series comprise a significant portion of the biomedical literature, yet unlike case reports, the National Library of Medicine does not index case series as a Publication Type. This hurts clinicians and researchers ability to retrieve, identify and analyze evidence from this type of study. Materials and MethodsPubMed articles mentioning "case series" in title or abstract were characterized to learn what are considered to be case series by the authors themselves. We then set aside articles better indexed as other standard publication types - case reports, cohort studies, reviews and clinical trials -- as well as those that discuss (rather than report the results of) case series studies, to create a corpus of typical case series articles. A random sample of these articles was evaluated by two annotators who confirmed that the great majority satisfy a formal definition of "case series". ResultsThe corpus was utilized in an automated transformer-based machine learning indexing model. Case series performance of this model on hold-out data was excellent (precision = 0.887, recall = 0.952, F1 = 0.918, PR-AUC = 0.941) and manual evaluation of 100 articles tagged as "case series" revealed that 88% satisfied a formal definition of case series. Discussion and ConclusionThis study demonstrates the feasibility of automatically indexing case series articles. Indexing should enhance their discoverability, and hence their medical value, for evidence synthesis groups as well as general users of the biomedical literature.

Published in BMC Medical Research Methodology (predicted rank #4) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.