Decompression of Bluetooth-transmitted Audio using Super Resolution for Low-Latency Applications
Allysa Joy Atienza, Andrei Calabano, Marie Lourdes Manalo, Sophia Beatrice Salandanan, Crisron Rudolf Lucas, Franz de Leon, Charleston Dale Ambatali, Carl Timothy Tolentino · 2021 International Conference on Information and Communication Technology Convergence (ICTC) · 2021
Bluetooth devices experience a common trade-off between quality and latency. This affects applications that require fast and accurate transmission of audio signals, especially in the medical and musical industry. In this paper, the use of super resolution techniques such as the Convolutional Neural Networks (CNN) and Generative Adversarial Networks (GAN) were utilized to improve the quality of Bluetooth-transmitted audio signals. We have shown that these two models were able to improve a number of speech and non speech audio signals based on our performance metrics: (1) Signal-to-Noise Ratio (SNR); (2) Log Spectral Distance (LSD); (3) Perceptual Evaluation of Audio Quality (PEAQ); and (4) Multi Stimulus test with Hidden Reference and Anchor (MUSHRA). Both models were able to improve the Bluetooth-transmitted audio signals, although the GAN model produced better results on both the objective and subjective evaluation tests compared to the CNN model. For the SNR, LSD, PEAQ, and MUSHRA, the GAN model averaged 20.4170 dB, 2.1408 dB, −2.7190, and 42.6500 respectively, while the CNN model averaged −1.7190 dB, 1.9410 dB, −2.7190, and 17.1600 respectively. For this project, the subjective test MUSHRA bear more weight in the results as the objective tests shows inconsistencies and cannot be heavily relied on.