Toward a Third-Kind Voice for Conversational Agents in an Era of Blurring Boundaries Between Machine and Human Sounds
Jeesun Oh, Hyeonjeong Im, Sangsu Lee · 2024
The voice of widely used conversational agents (CAs) is standardized to be highly intelligible, yet it still sounds machine-generated due to its artificial qualities. With advancements in deep neural networks, voice synthesis technology has become nearly indistinguishable from a real person. The voice enables users to discern the speakers’ identities and significantly impacts user perception, particularly in voice-only interactions. While more natural, human-sounding voices are generally preferred, their use in CAs raises potential ethical dilemmas, such as eliciting unwanted social responses or confusing the nature of the speaker. In this evolving landscape, it is necessary to understand the voice characteristics from multiple facets of voice design for CAs. Therefore, our study examines the voice characteristics of both artificial-sounding and human-sounding voices. Then, we propose a ‘third-kind’ of voice that considers the characteristics of each voice type. This discussion contributes to the debate on the future direction of voice design in the field of Conversational User Interface research.