Back

Task-Oriented Predictive (Top)-BERT: Novel Approach for Predicting Diabetic Complications Using a Single-Center EHR Data

Islam, H.; Bartlett, G.; Pierce, R. P.; Rao, P.; Waitman, L. R.; Song, X.

2024-04-16 health informatics
10.1101/2024.04.15.24305843 medRxiv
Show abstract

In this study, we assess the capacity of the BERT (Bidirectional Encoder Representations from Transformers) framework to predict a 12-month risk for major diabetic complications--retinopathy, nephropathy, neuropathy, and major adverse cardiovascular events (MACE) using a single-center EHR dataset. We introduce a task-oriented predictive (Top)-BERT architecture, which is a unique end-to-end training and evaluation framework utilizing sequential input structure, embedding layer, and encoder stacks inherent to BERT. This enhanced architecture trains and evaluates the model across multiple learning tasks simultaneously, enhancing the models ability to learn from a limited amount of data. Our findings demonstrate that this approach can outperform both traditional pretraining-finetuning BERT models and conventional machine learning methods, offering a promising tool for early identification of patients at risk of diabetes-related complications. We also investigate how different temporal embedding strategies affect the models predictive capabilities, with simpler designs yielding better performance. The use of Integrated Gradients (IG) augments the explainability of our predictive models, yielding feature attributions that substantiate the clinical significance of this study. Finally, this study also highlights the essential role of proactive symptom assessment and the management of comorbid conditions in preventing the advancement of complications in patients with diabetes.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.