Context, Perception, Production: A Model of Vocal Persona
Camille Noufi, Lloyd May, Jonathan C Berger · 2023
We present a contextualized production-perception model of vocal persona based on a deductive thematic analysis of interviews with voice and performance experts. The model formalizes how the vocal persona frames and bounds expressive vocal interactions (both biological and synthesized), and centers a person's agency over their communicative role. This article provides insights into opportunities for improvement in Voice User Interfaces (VUI) and Augmentative \& Assistive Communications (AAC) technologies based on the study's results. The proposed contextualized production-perception model fills an important gap in the literature on expressive and interactive speech technologies by identifying a key missing mechanism in contextualized vocal communication models and providing a resource for more nuanced approaches to expressive speech synthesis methods. Incorporating vocal persona into expressive vocal synthesis has the potential to significantly enhance the level of agency and embodiment experienced by VUI and AAC users during communication, resulting in a heightened sense of authenticity and an improved relationship with their environment.