Active learning pipeline to automatically identify candidate terms for a CDSS ontology: measures, experiments, and performance
Alluri, S.; Komatineni, K.; Goli, R.; Hubig, N.; Min, H.; Gong, Y.; Sittig, D. F.; Robinson, D.; Biondich, P.; Wright, A.; Nohr, C.; Law, T. D.; Faxvaag, A.; Boyce, R. D.; Gimbel, R. W.; Rennert, L.; Jing, X.
Show abstract
ObjectiveTo explore new strategies to make the document selection process more transparent, reproducible, and effective for the active learning process. The ultimate goal is to leverage active learning in identifying keyphrases to facilitate ontology development and construction, to streamline the process, and help with the long-term maintenance. MethodsThe active learning pipeline used a BILSTM-CRF model and over 2900 abstracts retrieved from PubMed relevant to clinical decision support systems. We started the model training with synthetic labeled abstracts, then used different strategies to select domain experts annotated abstracts (gold standards). Random sampling was used as the baseline. Recall, F1 (beta = 1, 5, and 10) scores are used as measures to compare the performance of the active learning pipeline by different strategies. ResultsWe tested four novel document-level uncertainty aggregation strategies--KPSum, KPAvg, DOCSum, and DOCAvg--that operate over standard token-level uncertainty scores such as Maximum Token Probability (MTP), Token Entropy (TE), and Margin. All strategies show significant improvement in early active learning cycles ({theta} to {theta}2) for recall and F1. The systematic evaluations show that KPSum (actual order) shows consistent improvement in both recall and F1 and KPSum (actual order) shows better results than the random sampling results. The document order (actual versus reverse) does not seem to play a critical role across strategies in model learning and performance in our datasets, although in some strategies, actual order shows slightly more effective results. The weighted F1 (beta = 5 and 10) provided complementary results to raw recall and F1 (beta = 1). ConclusionWhile prior work on uncertainty sampling typically focuses on token-level uncertainty metrics within generic NER tasks, our work advances this line of research by introducing a higher-level abstraction: document-level uncertainty aggregation. With a human-in-the-loop Active Learning pipeline, it can effectively prioritize high-impact documents, improve early-cycle recall, and reduce annotation effort. Our results show promise in automating part of ontology construction and maintenance work, i.e., monitoring and screening new publications to identify candidate keyphrases. However, future work needs to improve the model performance to make it usable in real-world operations.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A fast, accurate, and generalisable heuristic-based negation detection algorithm for clinical text 93%
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 93%
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 93%
Similar papers in this journal
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 94%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 92%
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
- Uncertainty in Deep Learning for EEG under Dataset Shifts 92%
- Graph Neural Network Modelling as a potentially effective Method for predicting and analyzing Procedures based on Patient Diagnoses 92%
Similar papers in this journal
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 93%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 92%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.