A Comparison of Representation Learning Methods for Medical Concepts in MIMIC-IV
Wu, X.; Liu, Z.; Zhao, Y.; Yang, Y.; Clifton, D. A.
Show abstract
ObjectiveTo compare and release the diagnosis (ICD-10-CM), procedure (ICD-10-PCS), and medication (NDC) concept (code) embeddings trained by Latent Dirichlet Allocation (LDA), Word2Vec, GloVe, and BERT, for more efficient electronic health record (EHR) data analysis. Materials and MethodsThe embeddings were pre-trained by the four aforementioned models separately using the diagnosis, procedure, and medication information in MIMIC-IV. We interpreted the embeddings by visualizing them in 2D space and used the silhouette coefficient to assess the clustering ability of these embeddings. Furthermore, we evaluated the embeddings in three downstream tasks without fine-tuning: next visit diagnoses prediction, ICU patients mortality prediction, and medication recommendation. ResultsWe found that embeddings pre-trained by GloVe have the best performance in the downstream tasks and the best interpretability for all diagnosis, procedure, and medication codes. In the next-visit diagnosis prediction, the accuracy of using GloVe embeddings was 12.2% higher than the baseline, which is the random generator. In the other two prediction tasks, GloVe improved the accuracy by 2%-3% over the baseline. LDA, Word2Vec, and BERT marginally improved the results over the baseline in most cases. Discussion and ConclusionGloVe shows superiority in mining diagnoses, procedures, and medications information of MIMIC-IV compared with LDA, Word2Vec, and BERT. Besides, we found that the granularity of training samples can affect the performance of models according to the downstream task and pre-train data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Creating a computer assisted ICD coding system: performance metric choice and use of the ICD hierarchy 96%
- An Open-Set Semi-Supervised Multi-Task Learning Framework for Context Classification in Biomedical Texts 96%
- ARCH: Large-scale Knowledge Graph via Aggregated Narrative Codified Health Records Analysis 95%
Similar papers in this journal
- Comparing neural language models for medical concept representation and patient trajectory prediction 96%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 95%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 94%
Similar papers in this journal
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 95%
- A Transformer-Based Model Trained on Large Scale Claims Data for Prediction of Severe COVID-19 Disease Progression 95%
- pathCLIP: Detection of Genes and Gene Relations from Biological Pathway Figures through Image-Text Contrastive Learning 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.