A machine learning model to support the screening for methods guidance articles in MEDLINE: A performance evaluation of ASReview simulation mode
Abdelkader, W.; Xie, D.; Lokker, C.; Chu, L.; Schandelmaier, S.; Saha, A.; Afzal, M.; Iorio, A.
Show abstract
BackgroundAdvances in clinical research methods are frequently published in biomedical journals, but identifying these articles remains challenging due to their rapid growth and insufficient indexing in biomedical databases. These challenges hinder the curation of methodologically focused resources like the Library of Guidance for Health Scientists (LIGHTS). Traditional screening approaches, such as Boolean search strategies and manual abstract screening, are inefficient and resource-intensive, limiting the feasibility of regularly updating LIGHTS. Machine learning (ML), particularly active learning models, presents a promising solution to improve the efficiency of article screening. ObjectivesThis study evaluates the performance of ASReviews active learning feature in identifying relevant methods guidance articles using pre-labeled data in simulation mode. MethodsUsing pre-labeled dataset composed of 1500 methods guidance articles and 20000 clinical studies, categorized as relevant or irrelevant, we trained and compared multiple simulation models in ASReview using various classifiers and feature extraction models. These included combinations of Support Vector Machine (SVM), Naive Bayes (NB), Neural Network with Sentence BERT (sBERT), Doc2Vec, and TF-IDF. Model performance was evaluated based on screening burden, recall, Work Saved over Sampling (WSS), and precision. All model combinations used maximum query and dynamic double sampling settings. ResultsAt 95-99.5% recall, SVM with TF-IDF required the fewest screened records (6.87-7.66% burden), while SVM with Doc2Vec achieved the best overall performance at 100% recall with only 11.47% screening burden (WSS@100 = 88.5%) in 42 minutes. Models using sBERT for feature extraction performed comparably through 99.5% recall but exhibited severe performance degradation at 100% recall, requiring screening of over 65% of the corpus. ConclusionClassical feature extraction methods, TF-IDF and Doc2Vec, paired with SVM outperform deep learning embeddings methods. ASReview in this controlled setting is a feasible tool for screening methodological literature. Future work should include prospective, human-in-the-loop experiments that embed the Doc2Vec-based SVM pipeline in comparison to human screening.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.