Back

Alternative Approaches for Modelling COVID-19:High-Accuracy Low-Data Predictions

Agarwal, D. K.; De, S.; Shukla, O.; Checker, A.; Mittal, A.; Borah, A.; Gupta, D.

2020-07-25 epidemiology
10.1101/2020.07.22.20159731 medRxiv
Show abstract

BackgroundNumerous models have tried to predict the spread of COVID-19. Many involve myriad assumptions and parameters which cannot be reliably calculated under current conditions. We describe machine-learning and curve-fitting based models using fewer assumptions and readily available data. MethodsInstead of relying on highly parameterized models, we design and train multiple neural networks with data on a national and state level, from 9 COVID-19 affected countries, including Indian and US states and territories. Further, we use an array of curve-fitting techniques on government-reported numbers of COVID-19 infections and deaths, separately projecting and collating curves from multiple regions across the globe, at multiple levels of granularity, combining heavily-localized extrapolations to create accurate national predictions. FindingsWe achieve an R2 of 0{middle dot}999 on average through the use of curve-fits and fine-tuned statistical learning methods on historical, global data. Using neural network implementations, we consistently predict the number of reported cases in 9 geographically- and demographically-varied countries and states with an accuracy of 99{middle dot}53% for 14 days of forecast and 99{middle dot}1% for 24 days of forecast. InterpretationWe have shown that curve-fitting and machine-learning methods applied on reported COVID-19 data almost perfectly reproduce the results of far more complex and data-intensive epidemiological models. Using our methods, several other parameters may be established, such as the average detection rate of COVID-19. As an example, we find that the detection rate of cases in India (even with our most lenient estimates) is 2.38% - almost a fourth of the world average of 9% [1].

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.