PronounSE: SFX Synthesizer from Language-Independent Vocal Mimic Representation
Riki Takizawa, Shigeyuki Hirai · 2024
Sound creators make various sound effects (SFX) depending on auditory events utilizing knowledge, techniques, and experience. These are challenging tasks for inexperienced creators. In this research, we focus on the fact that it is relatively easy for anyone to mimic SFX with utterances, and we propose a novel interactive technique called PronounSE. It can synthesize SFX to reflect subtle sound nuances by the vocal representation with language-independent sound mimicry. PronounSE consists of a Transformer that converts a mel-spectrogram of utterance mimic sound into a mel-spectrogram of SFX, and iSTFTNet, as a neural-vocoder, reconstructs a waveform for a synthesized SFX. We built a dataset for PronounSE that especially picked explosion sounds which we could easily represent with many nuances. This paper describes the model of PronounSE, the dataset of explosion sounds and their various vocal mimic representations of plural people, and the results of synthesized SFXs from untrained representations interactively.