Bayesian Microphone Array Processing

Takuma Otsuka · Kyoto University Research Information Repository (Kyoto University) · 2014

This dissertation presents Bayesian models of microphone array processing for computational auditory scene analysis in multisource environments.In such environments where multiple sounds are inevitably observed at a time, the decomposition function that extracts constituent sound source signals from the observed mixture of audio signals is essential for robust auditory processing because most audio decoding algorithms, such as speech recognition and sound source classification, assume clean and isolated audio signals as an input.We develop microphone array techniques that provide three fundamental functions: sound source separation, localization, and removal of reverberation (also known as dereverberation) to cope with multiple sound sources in practical environments.Microphone array processing must overcome the following three auditory uncertainties for achieving a robust decomposition function in real environments: (1) uncertainty in the number of sound sources, (2) reverberation in indoor environments, and (3) dynamic environments such as moving sound sources.These uncertainty issues have been partly addressed:(1) Sound source separation methods assuming unknown number of sources ignore the reverberation, which results in degraded separation performance.(2) Methods coping with sound source separation and dereverberation simultaneously are limited to the case where the microphones outnumber the sound sources.(3) As for microphone array processing in dynamic environments, the main topic has been focused on only the localization function.We overcome these uncertainties by using Bayesian nonparametrics so that our models can have an infinitely extensible flexibility to express the data and deal with observed mixture signals containing any number of sound sources.When the number of sound sources is uncertain, the selection of model complexity to handle the mixture signal is problematic because the model should be flexible enough to explain the observed mixture signal.Our Bayesian i

Read the paper · More papers on PaperTik