Beyond Functional Speech Synthesis
Rupal M. Patel, Geoffrey S. Meltzner, Markus Toman · Cambridge University Press eBooks · 2021
As synthetic voices enter the mass market, there is an increasing need for voice personalisation, that is, a voice for the text-to-speech system that not only conveys information but also exudes a persona much like the human voice. We begin this chapter with a historical overview of the field starting with model-based approaches, to concatenative systems and finally to contemporary implementations of parametric synthesis. We then examine how the confluence of increased computational speed at reduced costs, the availability of large data sets, and advances in machine learning and artificial intelligence enable a whole new approach to speech synthesis, including the ability to create high-quality personalised voices. We then examine the role of crowdsourcing in developing a scalable method for voice customisation and adaptation. We discuss the benefits and challenges of acquiring recordings of novice voice talent of all ages from around the world and using these recordings for voice building. We conclude by discussing the impact of personalised voices and implications for future work.