Synthetic Data Generation for Time Series Imputation: Comparing the Foundation Model Chronos with Established Methods
Sebastião Santos Lessa, Alexandre Lucas · 2025
Accurately imputing missing data is critical in time series analysis. The present work compares Foundation Model Chronos against Linear Interpolation, K-Nearest Neighbor Imputer, and Gaussian Mixture Model Imputer with three types of missing data patterns: random, short sequential chunks, and a long sequential chunk. These results confirm that for random missing values, KNN and interpolation yield the highest performance, while Chronos outperforms these on sequences. Indeed, however, for longer sequences of missing values, Chronos starts suffering from cascading errors which eventually allow the simpler imputation methods to outrank it. Another test with limited quantities of training data showed different tradeoffs for the different methods. Unlike KNN and interpolation, which smooth out the gaps, Chronos generates variable synthetic data. This can be beneficial in tasks which require control or simulation. The results highlight the strengths and weaknesses of the imputers and, therefore, offer practical insights into trade-offs between computational complexities, accuracy, and suitability for time series imputation scenarios.