Connectionist Temporal Classification-based Sound Event Encoder for Converting Sound Events into Onomatopoeic Representations

Koichi Miyazaki, Tomoki Hayashi, Tomoki Toda, Kazuya Takeda · 2018

In this paper, we propose a sound event encoder for converting sound events into their onomatopoeic representations. The proposed method uses connectionist temporal classification (CTC) as an end-to-end approach to directly convert a sequence of feature vectors of each sound event into a corresponding onomatopoeic word representation which accurately represents each sound and can be intuitively understood. Moreover, to address the issue of the ambiguity of onomatopoeic representations among different individuals, we develop a database of sound events and their corresponding typical onomatopoeic representations as accepted by multiple listeners. To evaluate the performance of our proposed method, we conduct objective and subjective evaluations. Experimental results demonstrate that the proposed sound event encoder is capable of converting sound events into their onomatopoeic representations with a 74.5% subjective acceptability rating, and that use of typical onomatopoeic representations, as approved by multiple subjects, yields significant improvement, resulting in an acceptability rate of 81.8%.

Read the paper · More papers on PaperTik