Interpretable multi-criteria selection of imputation models to enhance fault tolerance in real-time IoT data streams
Dimitris Gkoulis, Anargyros Tsadimas, George Kousiouris, Cleopatra Bardaki, George J. Dimitrakopoulos, Dimosthenis Anagnostopoulos, Μάρα Νικολαϊδου · Internet of Things · 2026
Edge-based IoT deployments must preserve data continuity when sensor observations are delayed, corrupted, or missing. Missing values should be repaired near the source, within stream timing budgets, and under the CPU, memory, and energy limits of single-board edge devices. Since no single imputation model is optimal across all stream states, short gaps, long tail gaps, sparse buffers, and diurnal regimes may favor different lightweight forecasting methods, requiring accurate, transparent, and simple online model selection. This work presents a real-time imputation framework combining a rolling-buffer streaming engine with a periodically re-inducible, online-executed decision-table selector. The engine regularizes incoming measurements, maintains per-stream quality indicators, and exposes a common tail-extrapolation interface for lightweight forecasting models. In scheduled calibration, it replays reference streams under controlled missingness, evaluates all candidate models under identical conditions, and logs accuracy and runtime evidence. Feature engineering, quantile-based discretization, aggregation, and a parametric cost function produce a compact decision table. Online, the deployed table acts as interpretable if-then logic to select a model for each imputation event; re-induction may be periodic or drift-triggered, with cadence set by domain timing and stability requirements. Evaluation uses a Harokopio University weather-station case study with ten-minute temperature, humidity, wind speed, barometric pressure, and solar radiation streams. Results show the reference selector improves sensor-tolerance-compliant imputations over any individual configured model while keeping mean end-to-end latency below the sampling period on Raspberry Pi-class hardware. The study validates a reference instantiation and feasibility, not universal optimality; architecture and selection remain application-agnostic and deployment-specific.