A model of event detection for perceptual extraction of the temporal structures of speech

Satomi Tanaka, Minoru Tsuzaki, Hiroaki Kato · 聴覚研究会資料 = Proceedings of the auditory research meeting · 2006

Abstract We introduce a new computational model, the “event-plausibility” model, as an extension of the loudness-jump model, which has been proposed with the aim of extracting temporal structures in speech based on simulated auditory processing. The main characteristic of the new model is to overcome a drawback of the loudness jump model, i.e., insensitivity to potential boundaries where the jump in loudness estimates is not sufficiently large. In the new model, the increases in auditory activation level in each tonotopic sub-band are computed using the Auditory Image Model, and they are used as an index for a potential new event. We compare the performance of the proposed model to that of the loudness-jump model by estimating speaking rates in a Japanese speech database. The results of the event-plausibility model demonstrate its advantage over the loudness-jump model. Keyword temporal structures,event detection,Auditory Image Model 1. Introduction A speech signal is a sequence of several sub-events and we perceive temporal features on the relation of these sub-events, such as rhythms, timing, and tempi (speaking rates). In natural speech, acoustical boundaries between consecutive sub-events are not necessarily clear, but one approach to overcoming this problem is to use linguistic knowledge. Although the importance of linguistic knowledge cannot be defied, the challenge of event detection based simply on auditory signal processing is also indispensable. Kato et al. have challenged this approach by investigating several characteristics in perceptual judgments on the temporal modification of speech signals, and consequently have proposed the “loudness-jump” model [1-4]. The loudness-jump model provides differences in loudness between successive syllables as a function of time. They found that perceptual performances for detecting temporal modification could be better explained by the loudness-jump concept than by a linguistic constraint such as the Japanese mora structure. However, the loudness-jump model has a problem. Detection of new events depends simply on the loudness of speech signals. Accordingly, the loudness-jump model cannot extract changes in signals without any change in loudness. Actually, there is a case that energy distribution along the signal’s frequency axis suddenly changes whereas loudness shows no change, and we can perceive the difference of signals and detect the arrival of a new event even in that case. To overcome this drawback of the loudness-jump model, we constructed a new simulation model, extending the loudness-jump concept by using the auditory image model (AIM) [5-7]. The AIM can provide a reasonably realistic simulation of the signal processing in the human auditory periphery. It produces multi-channel outputs that simulate the neural activity level of each frequency (tonotopic) sub-band. In the proposed model, information from each frequency band is sent through a nonlinear process, achieving cross-channel integration. The result of this integration can be assumed to indicate the level of plausibility of new events, and the model detects new events on a event-plausibility contour.

Read the paper · More papers on PaperTik