Back

Comparing Classifier Performance to Predict Infectious Diseases

Gonzalez, R. G.

2023-05-10 health informatics
10.1101/2023.05.06.23289606 medRxiv
Show abstract

We compared the accuracy of the machine learning classifier algorithms: Random Forest, Naive Bayes, Decision Tree, and Artificial Neural Network to predict zoonoses using the Random Forest extracted features and the serology data for seven different zoonotic diseases as the targets. We identified Random Forest and Naive Bayes as having the best performance overall. The Random Forest models above did well using Positive Predictive Value (PPV), Area Under the Curve (AOC) and Receiver Operating Characteristic (ROC) performance measures in identifying the positive cases for each of the diseases which is imperative when it comes to being able to identify the disease and then use this information to implement prevention and medical aid to specific areas and people where it is most needed. It also does well in predicting the negative values which is important to ensure the negatives are not false negatives. Naive Bayes was found to be the best choice for accuracy and performance. NB works well because it treats each feature as independent and thus, any change in one feature will not affect the other in the NB model. Decision Tree could not capture the data and thus, underfit during the first initial modeling and after hyper tuning. Artificial Neural Network overfit the model by capturing all the data including noise in the initial model, but underfit after hyper tuning. Both Decision Tree and Artificial Neural Network classifier algorithms are not recommended as classifiers for this dataset. StatementsThere are no conflicts of interest in this work. All methods were carried out in accordance with relevant guidelines and regulations. All experimental protocols were approved by the Forestry Administration of Cambodia. Informed consent was obtained from all subjects and/or their legal guardian(s) at the beginning of the survey.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.