Back

Predictive models for hospitalization and mortality in dengue using SINAN data: study protocol for development, temporal validation, and performance evaluation

Delpino, F. M. M.; Magalhaes, D.; Peres, I. T.; Gusberti, T.; de Lima, C. J.; Bozza, F. A.; Ranzani, O.; Bastos, L.

2026-08-04 infectious diseases
10.64898/2026.08.03.26359583 medRxiv
Show abstract

Background: Dengue continues to place a heavy clinical and organizational burden on health care systems, particularly during epidemics, when the high volume of cases adds pressure on triage, decisions regarding hospitalization, and the monitoring of patients at higher risk of severe outcomes. Despite the growing body of literature on dengue prediction, many studies still exhibit heterogeneity in outcomes, insufficiently detailed analytical designs, and a lack of validation. Objective: To describe the protocol for a study on the development and validation of predictive models for the outcomes of hospitalization among reported cases and mortality among hospitalized patients. Methods: A retrospective study will be conducted using secondary data from the Brazilian Notifiable Diseases Information System (SINAN), covering the period from 2017 to 2025. We will build two independent models: one to predict hospitalization among reported dengue cases and another to predict mortality among patients hospitalized for dengue. The protocol will follow TRIPOD+AI guidelines and include prior definition of eligible predictors, restriction to variables available at the clinically appropriate time of decision-making, handling of missing data, comparison between regression and machine learning algorithms, internal and temporal validation, and assessment of discrimination, calibration, and clinical utility. Conclusion: The study aims to establish a transparent and reproducible analytical protocol to support the development of risk models that could be applied to clinical screening and surveillance for dengue.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.