Customize Deep Learning-based De-Identification Systems Using Local Clinical Notes - A Study of Sample Size
Yang, X.; Bian, J.; Wu, Y.
Show abstract
Electronic Health Records (EHRs) are a valuable resource for both clinical and translational research. However, much detailed patient information is embedded in clinical narratives, including a large number of patients identifiable information. De-identification of clinical notes is a critical technology to protect the privacy and confidentiality of patients. Previous studies presented many automated de-identification systems to capture and remove protected health information from clinical text. However, most of them were tested only in one institute setting where training and test data were from the same institution. Directly adapting these systems without customization could lead to a dramatic performance drop. Recent studies have shown that fine-tuning is a promising method to customize deep learning-based NLP systems across different institutes. However, its still not clear how much local data is required. In this study, we examined the customizing of a deep learning-based de-identification system using different sizes of local notes from UF Health. Our results showed that the fine-tuning could significantly improve the model performance even on a small local dataset. Yet, when the local data exceeded a threshold (e.g., 700 notes in this study), the performance improvement became marginal.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 95%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 95%
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 94%
Similar papers in this journal
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 96%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 93%
Similar papers in this journal
- pathCLIP: Detection of Genes and Gene Relations from Biological Pathway Figures through Image-Text Contrastive Learning 95%
- Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach 95%
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 94%
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 94%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 93%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.