Back

Identifying Health Conditions in Older Adults in Textual Health Records Using Deep Learning-Based Natural Language Processing

Lin, J.; Korpi, T.; Kuukka, A.; Tirkkonen, A.-K.; Kariluoto, A.; Kaijansinkko, J.; Satamo, M.; Pajulammi, H.; Haapanen, M. J.; Hayrynen, S.; Pursiainen, E.; Ciovica, D.; von Bonsdorff, M. B.; Jylhava, J.

2024-10-10 health informatics
10.1101/2024.10.08.24315141 medRxiv
Show abstract

ImportanceMany clinically significant health conditions are frequently underreported, underdiagnosed or recorded only in unstructured textual health records, yet they contain critical information for patient assessment, care and prognosis. ObjectiveTo determine whether deep learning-based natural language processing employed for named entity recognition can effectively identify health conditions, such as incontinence, falls, mobility limitations and loneliness in unstructured textual electronic health records. The identified conditions were further used to predict all-cause mortality. DesignThis cohort study utilized electronic health records from public primary, secondary, tertiary, long-term and home care from 2010 to 2022, providing up to 12 years of follow-up. The named entity recognition task to identify incontinence, falls, mobility limitations, and loneliness was implemented using Googles Bidirectional Encoder Representations from Transformers deep learning model pre-trained for the Finnish language. Diagnostic codes for incontinence and falls were collected for comparisons. SettingRetrospective electronic health records across the Central Finland wellbeing services county. ParticipantsStructured summary data and 10.6 million free-text entries from 102,525 patients aged 50 to 80 years at baseline. ExposureIncontinence, falls, mobility limitations and loneliness were considered as exposures. Main Outcomes and MeasuresThe performance of the named entity recognition models was evaluated by precision, recall and F1 scores benchmarked against human ratings. Cox regression models were used to assess and compare NER- and diagnostic code-identified falls and incontinence onsets in predicting all-cause mortality. ResultsThe deep learning model demonstrated excellent performance with recall, precision and F1 scores of 0.86, 0.88, and 0.87 for falls; 0.84, 0.78, and 0.81 for incontinence; 0.86, 0.84, and 0.85 for mobility limitations and 0.91, 0.84, and 0.87 for loneliness, respectively. Compared to diagnostic codes, named entity recognition identified greater numbers of falls (31987 vs 4090) and incontinence (7059 vs 3873) onsets and yielded greater hazard ratios: 1.31 vs 1.04 for falls and 1.99 vs 0.65 for incontinence. Conclusions and RelevanceDeep learning-based named entity recognition models reliably identified incontinence, falls, loneliness and mobility limitations in free-text medical records, presenting new opportunities to use unstructured clinical data to identify vulnerable patients and apply the method in research applications. KEY POINTSO_ST_ABSQuestionC_ST_ABSCan deep learning-based natural language processing (NLP) identify health conditions, such as incontinence, falls, mobility limitations and loneliness in unstructured electronic health records (EHRs)? FindingsThe results of this cohort study demonstrate that a deep-learning NLP model can effectively identify incontinence, falls, mobility limitations and loneliness in textual EHR data. This approach also results in improved mortality prediction compared to available diagnostic codes. MeaningNLP approaches could be used to identify underreported and underdiagnosed health conditions in textural EHR data, enabling identification of vulnerable and at-risk patients.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.