Automatic identification of risk factors for SARS-CoV-2 positivity and severe clinical outcomes of COVID-19 using Data Mining and Natural Language Processing
Schoening, V.; Liakoni, E.; Drewe, J.; Hammann, F.
Show abstract
ObjectivesSeveral risk factors have been identified for severe clinical outcomes of COVID-19 caused by SARS-CoV-2. Some can be found in structured data of patients Electronic Health Records. Others are included as unstructured free-text, and thus cannot be easily detected automatically. We propose an automated real-time detection of risk factors using a combination of data mining and Natural Language Processing (NLP). Material and methodsPatients were categorized as negative or positive for SARS-CoV-2, and according to disease severity (severe or non-severe COVID-19). Comorbidities were identified in the unstructured free-text using NLP. Further risk factors were taken from the structured data. Results6250 patients were analysed (5664 negative and 586 positive; 461 non-severe and 125 severe). Using NLP, comorbidities, i.e. cardiovascular and pulmonary conditions, diabetes, dementia and cancer, were automatically detected (error rate [≤]2%). Old age, male sex, higher BMI, arterial hypertension, chronic heart failure, coronary heart disease, COPD, diabetes, insulin only treatment of diabetic patients, reduced kidney and liver function were risk factors for severe COVID-19. Interestingly, the proportion of diabetic patients using metformin but not insulin was significantly higher in the non-severe COVID-19 cohort (p<0.05). Discussion and conclusionOur findings were in line with previously reported risk factors for severe COVID-19. NLP in combination with other data mining approaches appears to be a suitable tool for the automated real-time detection of risk factors, which can be a time saving support for risk assessment and triage, especially in patients with long medical histories and multiple comorbidities.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and validation of an algorithm to estimate the risk of severe complications of COVID-19 to prioritise vaccination 93%
- Replicating a COVID-19 study in a national England database to assess the generalisability of research with regional electronic health record data 93%
- Development and validation of multivariable prediction models for adverse COVID-19 outcomes in IBD patients 93%
Similar papers in this journal
- Development and validation of a clinical risk score to predict the risk of SARS-CoV-2 infection from administrative data: a population-based cohort study from Italy 94%
- ChatGPT- versus human-generated answers to frequently asked questions about diabetes: a Turing test-inspired survey among employees of a Danish diabetes center 94%
- Regional performance variation in external validation of four prediction models for severity of COVID-19 at hospital admission: An observational multi-centre cohort study 93%
Similar papers in this journal
Similar papers in this journal
- Practical barriers and facilitators experienced by patients, pharmacists and physicians to the implementation of pharmacogenomic screening in Dutch outpatient hospital care – an explorative pilot study 91%
- Urine-based detection of biomarkers indicative of chronic kidney disease in a patient cohort from Ghana 91%
- Development and validation of decision rules models to stratify coronary artery disease, diabetes, and hypertension risk in preventive care: cohort study of returning UK Biobank participants 91%
Similar papers in this journal
- Developing and Evaluating Mappings of ICD-10 and ICD-10-CM Codes to PheCodes 95%
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 94%
- Using convolutional neural network to predict remission of diabetes after gastric bypass surgery: a machine learning study from the Scandinavian Obesity Surgery Register 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.