Identifying Health Conditions in Older Adults in Textual Health Records Using Deep Learning-Based Natural Language Processing
Lin, J.; Korpi, T.; Kuukka, A.; Tirkkonen, A.-K.; Kariluoto, A.; Kaijansinkko, J.; Satamo, M.; Pajulammi, H.; Haapanen, M. J.; Hayrynen, S.; Pursiainen, E.; Ciovica, D.; von Bonsdorff, M. B.; Jylhava, J.
Show abstract
ImportanceMany clinically significant health conditions are frequently underreported, underdiagnosed or recorded only in unstructured textual health records, yet they contain critical information for patient assessment, care and prognosis. ObjectiveTo determine whether deep learning-based natural language processing employed for named entity recognition can effectively identify health conditions, such as incontinence, falls, mobility limitations and loneliness in unstructured textual electronic health records. The identified conditions were further used to predict all-cause mortality. DesignThis cohort study utilized electronic health records from public primary, secondary, tertiary, long-term and home care from 2010 to 2022, providing up to 12 years of follow-up. The named entity recognition task to identify incontinence, falls, mobility limitations, and loneliness was implemented using Googles Bidirectional Encoder Representations from Transformers deep learning model pre-trained for the Finnish language. Diagnostic codes for incontinence and falls were collected for comparisons. SettingRetrospective electronic health records across the Central Finland wellbeing services county. ParticipantsStructured summary data and 10.6 million free-text entries from 102,525 patients aged 50 to 80 years at baseline. ExposureIncontinence, falls, mobility limitations and loneliness were considered as exposures. Main Outcomes and MeasuresThe performance of the named entity recognition models was evaluated by precision, recall and F1 scores benchmarked against human ratings. Cox regression models were used to assess and compare NER- and diagnostic code-identified falls and incontinence onsets in predicting all-cause mortality. ResultsThe deep learning model demonstrated excellent performance with recall, precision and F1 scores of 0.86, 0.88, and 0.87 for falls; 0.84, 0.78, and 0.81 for incontinence; 0.86, 0.84, and 0.85 for mobility limitations and 0.91, 0.84, and 0.87 for loneliness, respectively. Compared to diagnostic codes, named entity recognition identified greater numbers of falls (31987 vs 4090) and incontinence (7059 vs 3873) onsets and yielded greater hazard ratios: 1.31 vs 1.04 for falls and 1.99 vs 0.65 for incontinence. Conclusions and RelevanceDeep learning-based named entity recognition models reliably identified incontinence, falls, loneliness and mobility limitations in free-text medical records, presenting new opportunities to use unstructured clinical data to identify vulnerable patients and apply the method in research applications. KEY POINTSO_ST_ABSQuestionC_ST_ABSCan deep learning-based natural language processing (NLP) identify health conditions, such as incontinence, falls, mobility limitations and loneliness in unstructured electronic health records (EHRs)? FindingsThe results of this cohort study demonstrate that a deep-learning NLP model can effectively identify incontinence, falls, mobility limitations and loneliness in textual EHR data. This approach also results in improved mortality prediction compared to available diagnostic codes. MeaningNLP approaches could be used to identify underreported and underdiagnosed health conditions in textural EHR data, enabling identification of vulnerable and at-risk patients.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
- Beyond Metrics to Methods: A Scoping Review of Large Language Models for Detection of Social Drivers of Health in Clinical Notes 93%
Similar papers in this journal
- Identifying clusters of people with Multiple Long-Term Conditions using Large Language Models: a population-based study 94%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 92%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 92%
Similar papers in this journal
- Machine learning approach to dynamic risk modeling of mortality in COVID-19: a UK Biobank study 94%
- Systematic identification of rare disease patients in electronic health records enables evaluation of clinical outcomes 94%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
Similar papers in this journal
- An external validation of the QCovid risk prediction algorithm for risk of mortality from COVID-19 in adults: national validation cohort study in England 93%
- Understanding COVID-19 trajectories from a nationwide linked electronic health record cohort of 56 million people: phenotypes, severity, waves & vaccination 92%
- Remote Covid Assessment in Primary Care (RECAP) risk prediction tool: derivation and real-world validation studies 92%
Similar papers in this journal
- Estimating excess mortality in people with cancer and multimorbidity in the COVID-19 emergency 92%
- Identifying markers of health-seeking behaviour and healthcare access in UK electronic health records 92%
- Development and validation of automated computer aided-risk score for predicting the risk of in-hospital mortality using first electronically recorded blood test results and vital signs for COVID-19 hospital admissions: a retrospective development and validation study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.