Back

Data-Driven Inference of COVID-19 Clinical Prognosis

Salas, J.; Pulido, D.; Montoya, O.; Ruiz, I.

2020-09-01 public and global health
10.1101/2020.08.27.20183202 medRxiv
Show abstract

Knowing the most likely clinical prognosis for a patient infected with SARS-Cov-2 could offer guidelines for tracking their medical evolution, improving attention, and assigning resources. Aiming to assess a patients status quantitatively, we explore the analysis of existing clinical information using data-driven methods. Our goal is to extract the characteristics distinguishing between those COVID-19 patients that improve and those who die. In our approach, we select the relevant features using the algorithm of Boruta, a wrapper framework that takes input from classifiers generating relevance assessment of the predictors. Using the extracted features, we train machine learning classifiers, including Random Forests, Support Vector Machine, Extreme Gradient Boosting, and Neural Networks. We assess the performance of the classifiers using Precision-Recall and ROC analysis, establishing the ranges at which risk assessment permits effective decision-making. Our research highlights that local regions present unique sets of essential features, that it is possible to construct effective classifiers based on clinical data, and that an ensemble of classifiers results in the best performing discriminant.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.