Deep-learning-based speech enhancement under rough-focus conditions with optical laser microphone

Yuki NAKANO, Kazumichi MIYAZATO, Yuting Geng, Kenta Iwai, Takanobu Nishiura · NOISE-CON proceedings · 2024

Optical laser microphones have attracted attention for acoustic systems capable of recording the target speech from a distance. An optical laser microphone measures the speech-induced vibration by focusing the laser beam on the surface of the vibrating object. A recording method called rough-focus recording using an unfocused laser beam, enables wide-range recording and robust recording against changes in focal length. However, with rough-focus recording, the broad laser beam coverage leads to insufficient intensity of the reflected laser for accurate acoustical signal measurement, resulting in speech-quality degradation, such as the inclusion of noise, and attenuation of high-frequency components in the acquired speech. To solve this problem, deep-learning-based speech-enhancement methods for optical laser microphones have been proposed. Such methods require separate models for different focus settings, exhibiting a lack of adaptability to changing focus settings. We propose a speech-enhancement method for training a single model for various focus settings. This model is trained with speech signals recorded across different focus settings to enhance speech recorded in various focus settings. Experimental results indicate that the this model trained with the proposed method performs equivalent to or better than a model trained with the conventional models.

Read the paper · More papers on PaperTik