Enhancing data quality in wastewater processes: Missing data imputation with deep Variational Autoencoders and genetic algorithms
Christian Kazadi Mbamba, Philip C. Keymer, Maira Alvi, Sebastian Olivier Nymann Topalian, Fareed Ud Din, Damien John Batstone · Computers & Chemical Engineering · 2025
Missing data is a persistent challenge in wastewater analysis, often leading to biased results and reduced accuracy. This study introduces an innovative Automated Machine Learning (AutoML) framework that combines deep learning-based variational autoencoders (VAEs) and genetic algorithms (GAs) to address this issue. VAEs are employed to impute missing values by learning latent data representations, while GAs optimize the VAE architecture and hyperparameters, including the size of the latent space. The framework is specifically designed to handle the complex and nonlinear relationships in wastewater datasets. The framework was trained and validated using data from a full-scale water resource recovery facility. The imputed data from the optimized VAE, developed using the GA-based AutoML framework, is then used to train predictive models. Experimental evaluations demonstrate the effectiveness of the proposed approach over traditional imputation methods. The results reveal that the models can accurately predict key variables such as ammonia nitrogen (NH 4 -N), nitrate nitrogen (NO 3 -N), pH, and biogas flow rate, using imputed data. The scalability and adaptability of this framework make it valuable for real-time wastewater monitoring and predictive analytics.