Real-Time Audio Spectrogram Inpainting

Akshesh Shah · 2025

In this study, we explore the use of Vector Quantized Variational Autoencoders (VQ-VAE) for real-time audio spectrogram inpainting, with a focus on minimizing environmental impact. We leverage the capabilities of VQ-VAE, a generative model that compresses and quantizes spectrograms into compact code maps, to regenerate missing or corrupted parts of audio signals. Our experiments, conducted on both the MNIST image dataset and the NSynth audio dataset, demonstrate the model’s effectiveness in reconstructing spectrograms. We also introduce a real-time CO2 tracker to monitor the energy consumption and carbon footprint of our training processes. Our results highlight the potential of VQ-VAE in the field of computer music, providing innovative tools for musicians while addressing the environmental costs of machine learning.

Read the paper · More papers on PaperTik