Comparative Analysis of Decision Trees on Two COVID-19 Symptom Datasets
Saengamnatdej, S.; Molee, P. W.; Warnnissorn, P.
Show abstract
ObjectiveThis study compares decision trees on two COVID-19 symptom datasets to assess their performance and feature importance in predicting and understanding infection patterns. MethodsWe created decision trees on Israeli and Swedish COVID-19 infection datasets. Performance metrics were used to assess their predictive capabilities, and feature importance analysis identified significant variables in the decision-making process. ResultsThe study observed different performance levels of decision trees on the COVID-19 datasets. The Swedish dataset achieved high accuracy and F1-score without hyperparameter tuning, while the Israeli dataset improved significantly with Extreme Gradient Boosting. Dataset characteristics impact the selection of an optimal decision tree algorithm. The key variable in both datasets was sore throat. ConclusionThis study compares decision trees on COVID-19 infection datasets, emphasizing the importance of dataset characteristics in selecting an optimal algorithm. Identifying significant features enhances understanding of infection patterns, benefiting decision-making and prediction accuracy in infectious disease analysis.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 94%
- A multipurpose machine learning approach to predict COVID-19 negative prognosis in Sao Paulo, Brazil 94%
- Machine learning to predict retention and viral suppression in South African HIV treatment cohorts 94%
Similar papers in this journal
- Enhanced Formulation of Precision Probiotics through Active Machine Learning 93%
- Harnessing multi-output machine learning approach and dynamical observables from network structure to optimize COVID-19 intervention strategies 92%
- Prediction of high-risk liver cancer patients from their mutation profile: Benchmarking of mutation calling techniques 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.