Back

Clinical study type classification, validation, and PubMed filter comparison with natural language processing and active learning

van IJzendoorn, D. G. P.; Habets, P. C.; Vinkers, C. H.; Otte, W. M.

2022-11-03 epidemiology
10.1101/2022.11.01.22281685 medRxiv
Show abstract

Each day, many thousands of new studies are published. Identifying specific study types with high sensitivity and specificity may improve searchability and accelerate updating systematic reviews and meta-analyses. Machine learning transformer models could facilitate this identification process if sufficient training data is available. We used an active learning strategy to construct a large training set (n=50,000) and fine-tuned the PubMedBERT language model to classify PubMed abstracts as randomized controlled trials, human studies, systematic reviews with and without meta-analyses, protocols, and rodent studies. In an external dataset (n=5,000), the average sensitivity and specificity across study types were 0.94 and 0.96, respectively. PubMeds internal filters had a low sensitivity for both systematic reviews with meta-analysis (0.175, CI: 0.057-0.293) and randomized controlled trials (0.256, CI: 0.119-0.393). We applied this labeling to all 34 million PubMed abstracts currently available and provide the results within an online meta-information platform (EvidenceHunt). In conclusion, we show that study type classification in PubMed is opportune, given the available language models. The high accuracy in this study invites extending these models to more elaborate and hierarchical identification schemes.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.