Generation of Similar Sounds Using GPTs in a Sound Effects Synthesis System

Kakeru Iwamoto, Hironori Uchida, Yujie Li, Yoshihisa Nakatoh · 2025

Nowadays, the market size for services such as gaming and video streaming has been expanding, and along with this, the demand for essential sound effects in content production is also increasing. Additionally, audio generation AI, such as Diff-Sound, has emerged in response to this growing demand. However, these AI systems face challenges when tasked with generating content based on prompts that contain multiple events. In this study, we propose a method for improving the generation of sound effects by utilizing GPTs to preprocess prompts containing multiple events, based on a system used in previous research. This preprocessing technique involves using GPTs to decompose prompts with multiple events into simpler, single-event components. As a result, the generated sound effects are improved through this process.

Read the paper · More papers on PaperTik