Confidence Interval Estimation for Machine Learning Models in Forecasting Infectious Diseases

Taewan Goo, Kyulhee Han, Hanbyul Song, Jiwon Park, Zhe Liu, Jooha Oh, Sayooj Aby Jose, Taesung Park · 2024

Forecasting models have been instrumental in managing the COVID-19 pandemic by facilitating efficient resource distribution and proper interventions. Although various forecasting models—such as mathematical models, statistical models, and machine learning models—are available, a few models account for the uncertainty in their predictions. The lack of uncertainty information constrains the reliability of forecasting results, making them less applicable for use in decision-making. Confidence Intervals (CIs) are widely used in statistical inference to provide uncertainty information. In this study, we introduce a framework that provides bootstrap-based CIs readily applicable to various forecasting models. The key concept of this framework is repeated training on bootstrap data based on resampled residuals from the forecasting model. After fitting the forecasting model, bootstrap datasets are generated considering time series characteristics to retrain the forecasting models. Then, CIs are calculated based on the distribution of prediction values. This framework was applied to forecasting COVID-19 outcomes—such as daily confirmed COVID-19 cases, deaths, and ICU patients—in South Korea. Our results demonstrate that bootstrap-based CIs can be successfully applied to provide additional information on the reliability of prediction. In conclusion, our framework can provide uncertainty information without requiring complex assumptions, regardless of the forecasting model or response variable type. Therefore, this approach can support public health policymakers in gaining a deeper understanding of epidemic trends, thereby offering valuable insights for decision-making.

Read the paper · More papers on PaperTik