A multi-frame blocking for signal segmentation in voice command recognition

Achmad Fanany Onnilita Gaffar, Rheo Malani, Supriadi Supriadi, Agusma Wajiansyah, Arief Bramanto Wicaksono Putra · 2020

Frame blocking segment the voice signals into overlapping frames to ensure information is maintained. Various sound features in the time domain can be extracted from the signal representation in the form of this frame. The frame blocking stage allows spectral distortion to appear at the beginning and the end of each frame, which gives rise to discontinuous signal pieces. It will appear as part of the voice feature if extracted directly in the time domain. In the frequency domain, these discontinuous signal pieces are removed using windowing techniques. However, the use of windowing techniques sometimes has an impact on changing signal information at the beginning and the end of the frame. This study proposes a multi-frame blocking method to optimize the frame blocking stage. The aim of the proposed method is to maintain the integrity of the signal information while minimizing the number of frames without losing information. In this study, the sample of the voice signal is a command signal used to control the basic movements of the robot. The absolute portion of PSD (Power Spectral Density) have used as a voice signal feature where the average values represented as grayscale images. The results of studies have shown an increase in the performance of the voice signal recognition stage by 18.29% when compared to using only conventional frame blocking methods.

Read the paper · More papers on PaperTik