Investigating the impact of weakly supervised data on text mining models of publication transparency: a case study on randomized controlled trials
Jiang, L.; Hoang, L.; Kilicoglu, H.
Show abstract
Lack of large quantities of annotated data is a major barrier in developing effective text mining models of biomedical literature. In this study, we explored weak supervision strategies to improve the accuracy of text classification models developed for assessing methodological transparency of randomized controlled trial (RCT) publications. Specifically, we used Snorkel, a framework to programmatically build training sets, and UMLS-EDA, a data augmentation method that leverages a small number of existing examples to generate new training instances, for weak supervision and assessed their effect on a BioBERT-based text classification model proposed for the task in previous work. Performance improvements due to weak supervision were limited and were surpassed by gains from hyperparameter tuning. Our analysis suggests that refinements to the weak supervision strategies to better deal with multi-label case could be beneficial.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Scalable information extraction from free text electronic health records using large language models 94%
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 93%
- Quantitative bias analysis in practice: Review of software for regression with unmeasured confounding 91%
Similar papers in this journal
- ZIBGLMM: Zero-Inflated Bivariate Generalized Linear Mixed Model for Meta-Analysis with Double-Zero-Event Studies 93%
- Evaluation of statistical methods used to meta-analyse results from interrupted time series studies: a simulation study 93%
- Fast and frugal decision tree for the rapid critical appraisal of systematic reviews 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.