Evaluating machine learning approaches for multi-label classification of unstructured electronic health records with a generative large language model
Vithanage, D.; Deng, C.; Wang, L.; Yin, M.; Alkhalaf, M.; Zhang, Z.; Zhu, Y.; Soewargo, A. C.; Yu, P.
Show abstract
Multi-label classification of unstructured electronic health records (EHR) is challenging due to the semantic complexity of textual data. Identifying the most effective machine learning method for EHR classification is useful in real-world clinical settings. Advances in natural language processing (NLP) using large language models (LLMs) offer promising solutions. Therefore, this experimental research aims to test the effects of zero-shot and few-shot learning prompting, with and without parameter-efficient fine-tuning (PEFT) and retrieval-augmented generation (RAG) of LLMs, on the multi-label classification of unstructured EHR data from residential aged care facilities (RACFs) in Australia. The four clinical tasks examined are agitation in dementia, depression in dementia, frailty index, and malnutrition risk factors, using the Llama 3.1-8B. Performance evaluation includes accuracy, macro-averaged precision, recall, and F1 score, supported by non-parametric statistical analyses. Results indicate that both zero-shot and few-shot learning, regardless of the use of PEFT and RAG, demonstrate equivalent performance across the clinical tasks when using the same prompting template. Few-shot learning consistently outperforms zero-shot learning when neither PEFT nor RAG is applied. Notably, PEFT significantly enhances model performance in both zero-shot and few-shot learning; however, RAG improves performance only in few-shot learning. After PEFT, the performance of zero-shot learning is equal to that of few-shot learning across clinical tasks. Additionally, few-shot learning with RAG surpasses zero-shot learning with RAG, while no significant difference exists between few-shot learning with RAG and zero-shot learning with PEFT. These findings offer crucial insights into LLMs for researchers, practitioners, and stakeholders utilizing LLMs in clinical document analysis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 95%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 95%
- Evaluating and mitigating unfairness in multimodal remote mental health assessments 95%
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 95%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 95%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 94%
Similar papers in this journal
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 96%
- Using Artificial Intelligence to Learn Optimal Regimen Plan for Alzheimer’s Disease 95%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.