Whisper Restoration Combining Real- and Source-Model Filtered Speech for Clinical and Forensic Applications

Francesco Roberto Dani, Sonia Cenceschi, Alice Albanesi, Elisa Colletti, Alessandro Trivilini · 2023

This chapter focuses on the structure and development of the algorithmic component of VocalHUM, a smart system aiming to enhance the intelligibility of patients&s; whispered speech in real time, based on audio only. VocalHUM is primary designed for patients in a temporary or prolonged state of physical and vocal frailty (e.g., respiratory infection, geriatric weakness, or partial/total paralysis) with the ultimate purpose to facilitate patient–caregiver communication and improve the use of Speech-to-Text tool and voice commands. To date, its Whisper-to-Speech algorithm focuses on the Italian language and combines real whispered speech, synthetized vowels (generated by additive synthesis) and consonant enhancement techniques. The chapter discusses choices that make us move from neural networks to the generative approach, paying attention to the addressed issues for an application in real contexts. The first results are presented both in laboratory and real context, and their possible application in different fields, such as digital audio forensics, are discussed.

Read the paper · More papers on PaperTik