Back

Mind the Baseline: The Hidden Impact of Reference Model Selection on Forecast Assessment

Stapper, M.; Funk, S.

2025-08-05 epidemiology
10.1101/2025.08.01.25332807 medRxiv
Show abstract

Baseline models are essential reference points for evaluating forecasting methods, yet their selection often receives insufficient attention. We present a systematic framework for baseline model selection in epidemiological forecasting, establishing criteria for suitable baselines and demonstrating the consequences of different choices. Analysing data from COVID-19 and influenza forecast hubs, we evaluated ten baseline model frameworks. Our results reveal that baseline selection profoundly impacts forecast evaluation: for influenza, the proportion of models outperforming the baseline ranged from 11% to 100% depending on the baseline chosen. No single baseline satisfied all evaluation criteria. The choice of baseline also affected model rankings, with some baselines producing substantially different orderings of forecast model performance. We found that well-calibrated baselines do not necessarily align with good forecast performance, highlighting a fundamental tension in baseline selection. These findings highlight the need for careful baseline selection in forecast evaluation, particularly in collaborative efforts where fair comparison across multiple models is essential. We provide practical recommendations for baseline selection and suggest strategies for improving evaluation fairness when ideal baselines cannot be identified.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.