Understanding the Use of AI-Based Audio Generation Models by End-Users
Jiyeon Han, Eunseo Yang, Uran Oh · 2024
With the growing popularity of video platforms, the demand for copyright-free audio sources for adding background music to videos is also expected to increase. While text-to-audio-generation models can be useful for this purpose, little is known about how people perceive and use these models. To understand how generation models for audio are used and to identify their strengths and weaknesses compared to typical audio search engines, we conducted a user study with 16 participants, where they were asked to choose matching background music after watching muted videos. Findings show that participants appreciate the search engine for recommending search keywords and displaying multiple results although the outcome does not fully reflect the intent. In contrast, the generation model posed challenges in choosing proper prompts but excelled in finding the desired music. Based on these results, we suggest design considerations to improve the usability of the audio generation model for end-users.