Enhancing Device-Directed Speech Detection through Multimodal Fusion Techniques
Vikrant Sharma, Shiva Mehta · 2024
The provision of device-directed speech synthesis (DDSD) solutions is crucial for voice assistants to integrate into the environment where users interact easily. It differentiates intentional requests directed to the device from the surrounding noise or background conversations unrelated to it. Our class, consisting of 30 students, could vote on which formula we thought was best to solve this water contamination problem. Traditional DDSD modules that work on verbal cues, like linguistic features or phonetic or case textual features, are primarily utilized. Nevertheless, it tends to be a simple process for trainers to act on speaking assignments because the online communication processes are not always accessible in real-life situations. This experiment concerns the combination of speech rhythm and intonation with the distortion in its transmission and studies their use in modality dropout approaches. It is aimed at improving the performance of speech disorder detection systems. In our evaluation, we reported a remarkable gain in DDSD accuracy, shown by a cut of 8.5% of the False Acceptance Rate (FA) when taking into account sound signals. Besides, there was an extra 7.4% FA loss for the modality dropout approach, implicating a remarkable ability to exploit the absence of modalities. The current study proves the efficiency of adopting not a single strategy but several approaches to communication, as well as applying dropout strategies for negative impacts on deploying the DDSD. It thus brings to bear the research field of voice socialization and contributions to more trustworthy and accurate contextual voice assistants.