Back

Comparative assessment of machine learning algorithms to predict severity of disease in COVID-19 patients based on eight cofactors

Patel, T. S.; Patel, D. P.; El-Sayed, H.; Sanyal, M.; Shrivastav, P. S.

2024-10-14 bioinformatics
10.1101/2024.10.11.617884 bioRxiv
Show abstract

Machine learning is one of the important tools to diagnose and predict the diseased state accurately and effectively. The COVID-19 pandemic caused due to severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has become one of the most researched healthcare topics worldwide. Machine learning algorithms can find efficient and reliable ways to predict the COVID-19 from vast amounts of existing health care data, allowing faster, effective, and more accurate diagnosis with lower risk based on the symptoms. Based on the countrywide data published by the Israeli Ministry of Health, we propose a system that detects COVID-19 instances using simple variables. The COVID-19 dataset used in the study consisted of 278848 patients samples with five different symptoms, namely cough, fever, sore throat, shortness of breath, and headache, apart from other basic information like age, gender, and test indication excluding confirmed COVID-19 result. The data was analyzed using traditional supervised machine learning algorithms namely, Decision tree, Support vector machine, Random Forest, Logistic regression, k-nearest neighbor, and Naive Bayes based on eight cofactors with high accuracy rate ([≥] 0.9450). Apart from Support vector machine, all other algorithms displayed better performance based on the AUC score calculated using the receiver operator characteristic (ROC) curve. This study also highlights the significant differences between precision, recall and accuracy for each model. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=97 SRC="FIGDIR/small/617884v1_ufig1.gif" ALT="Figure 1"> View larger version (19K): org.highwire.dtl.DTLVardef@33bb11org.highwire.dtl.DTLVardef@3e93aaorg.highwire.dtl.DTLVardef@508f76org.highwire.dtl.DTLVardef@faa9e2_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphical AbstractC_FLOATNO C_FIG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.