To cite this paper use one of the standards below:
In multi-series forecasting, a model trained on sliding windows is deployed to forecast known or unseen realizations. Prior work on cross-validation for time series focuses on the single-series regime and defaults to temporal splits, leaving the multi-series setting underexplored. We compare four strategies across seven datasets, evaluating calibration bias, model-selection ability, and autoregressive robustness under both temporal and group generalization. Results reveal a bias--coverage trade-off: shuffled $k$-fold attains the lowest mean absolute bias in both scenarios, yet its confidence intervals rarely cover the test error under group holdout. Group-aware and temporal splits trade slightly higher bias for markedly better coverage, and are statistically indistinguishable from each other on both axes. Well-calibrated strategies also select near-optimal window sizes, and findings extend to multi-step prediction. Any strategy respecting the deployment scenario's structure is a reasonable default, once expanding-window folds too small to represent the deployed model are discarded before averaging.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper