GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records
Yang, X.; Pour Nejatian, N.; Shin, H. C.; Smith, K.; Parisien, C.; Compas, C.; Martin, c.; Flores, M.; Zhang, Y.; Magoc, T.; Harle, C.; Lipori, G.; Mitchell, D.; Hogan, W.; Shenkman, E.; Bian, J.; Wu, Y.
Show abstract
ObjectiveTo develop a large pretrained clinical language model from scratch using transformer architecture; systematically examine how transformer models of different sizes could help 5 clinical natural language processing (NLP) tasks at different linguistic levels. MethodsWe created a large corpus with >90 billion words from clinical narratives (>82 billion words), scientific literature (6 billion words), and general English text (2.5 billion words). We developed GatorTron models from scratch using the BERT architecture of different sizes including 345 million, 3.9 billion, and 8.9 billion parameters, compared GatorTron with three existing transformer models in the clinical and biomedical domain on 5 different clinical NLP tasks including clinical concept extraction, relation extraction, semantic textual similarity, natural language inference, and medical question answering, to examine how large transformer models could help clinical NLP at different linguistic levels. Results and ConclusionGatorTron scaled up transformer-based clinical language models to a size of 8.9 billion parameters and achieved state-of-the-art performance on 5 clinical NLP tasks of different linguistic levels targeting various healthcare information documented in unstructured electronic health records (EHRs). The proposed GatorTron models performed remarkably better in much complex clinical NLP tasks such as natural language inference (9.6% and 7.5% improvements) and question answering (9.5% and 7.77% improvements) compared with existing smaller clinical transformer models (i.e., BioBERT and ClinicalBERT), demonstrating the potential of large transformer-based clinical models for advanced medical artificial intelligent (AI) applications such as question answering.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 95%
- Zero Shot Health Trajectory Prediction Using Transformer 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
Similar papers in this journal
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 95%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 93%
- Repurposing Non-pharmacological Interventions for Alzheimer’s Diseases through Link Prediction on Biomedical Literature 93%
Similar papers in this journal
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 96%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 93%
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 93%
Similar papers in this journal
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 94%
- KG-COVID-19: a framework to produce customized knowledge graphs for COVID-19 response 92%
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.