Noise Estimation of Self-Coupled Laser Microphone Using Transformer-Based Source Separation Models
Takemasa Okita, Norio Tsuda, Daisuke Mizushima · 2024
Transformer-based models excel in a wide range of deep learning areas, such as natural language processing and image recognition, and have recorded the highest accuracy in source separation, which separates the speech of each source from a mixture of multiple speech sounds. The laser microphone used in this study has more noise superimposed on it than general microphones and has not yet been used in practical applications. Therefore, in this study, we performed source separation using Transformer-based models to reduce noise in laser microphones and compared the improvement in scale-independent signal-to-distortion ratio (SI-SDRi) with a conventional model (convolution-based model). Transformer-based models are structurally capable of capturing temporal dependencies with high accuracy. Therefore, the separated speech is very close to the target signal, and sufficient noise reduction is achieved. The audio after separation was also analyzed for noise specific to the laser microphone. It was found that the noise specific to the laser microphone was present in the low frequency band below 100 Hz and that a constant noise such as white noise was superimposed in the other frequency bands.