Application of a microphone array to the automatic structuring of meeting recordings

Futoshi Asano, Jun Ogata, Yosuke Matsusaka · The Journal of the Acoustical Society of America · 2006

A microphone array system applied to the automatic structuring of meeting recordings is introduced. First, speakers in a meeting are identified at every time block (0.25 s) by sound localization using a microphone array, and then the speech events are extracted (structuring). Next, overlaps and insertions of speech events such as ‘‘Uh-huh,’’ which greatly reduce the performance of the automatic speech recognition, are separated out. The difficulty in eliminating overlaps and insertions is that the overlapping sections are sometimes very short, and data sufficiently long for conventional separation techniques, such as independent component analysis, cannot be obtained. In this system, a maximum-likelihood beamformer is modified so that all the information necessary for constructing the separation filter, such as speaker location and spatial characteristics of noise, are estimated from the recorded data based he structuring information. Evaluation of the performance of the speech event separation by automatic speech recognition is also presented.

Read the paper · More papers on PaperTik