Missing Data Imputation for LSTM-Based Flood Early Warning System in Jakarta

Akmarina Khairunnisa, Bagus Sartono, Muhammad Nur Aidi · International Research Journal of Innovations in Engineering and Technology · 2025

The technological advancements of data storage capacity and computational capabilities have implications for the recording of time series data with increasingly narrow intervals, called high-frequency time series data. Sensor data, as a prominent example of high-frequency time series generated through the utilization of the Internet of Things (IoT), is susceptible to issues related to missing data due to the likelihood of device failures. Furthermore, both the quantity and quality of data significantly impact the performance of forecasting models. This study examines the effects of imputing missing data within a forecasting workflow for sensor data that records water levels at four observation sites. The analysis will be conducted by evaluating 6 imputation methods in a simulation study using 10 datasets with 18 missing scenarios each. The forecasting outcomes of the IMV-LSTM (Interpretable Multi Variable Long ShortTerm Memory) model, trained using empirical data reconstructed through the best imputation methods from the simulation study, will also be evaluated. The results indicate that the imputed data using the KalmanStructural method enhances forecast accuracy, evidenced by a 32% reduction in RMSE compared to the model trained on data without imputation treatment as the benchmark. Additionally, imputed data employing Kalman-ARIMA improves the performance of the IMVLSTM model, yielding a 29% lower RMSE compared to the benchmark. The best-performing model demonstrates that the forecasts of water levels deviate by only approximately 0.1% from the actual data.

Read the paper · More papers on PaperTik