Real-time voice adaptation with abstract normalization and sound-indexed based search
Mads Midtlyng, Yuji Sato · 2016
This paper proposes a two-step based real-time voice adaptation system in the field of speech processing. Step one combines recording and pre-processing to construct a voice profile. Secondly is the real-time raw input of the voice's adaptation to a target voice. The fact that individual voices' structure are habitually varying, we suggest a method for converting into a comparable format. The new method is called abstract normalization which cuts the voice data into smaller sounds and generate an abstracted, simplified version of the data using a level of abstraction along with parameter fitting. The normalized data is used to generate a sound-index which is a sequence hash that represents the data in a simpler fashion. The indices are used to compare different sounds/voices for adaptation. This effectively transforms the speech-related challenges into a search problem rather than a biometric one. To assess the approach, voice profile data are compared against each other as a method to verify the sound-index. Ultimately, a real-time voice input using alternating levels of abstraction is run against a Norwegian voice profile. The degree of adaptation success is measured in percentage, and experimental results show that while accuracy is not yet excellent, the concept was validated.