Comparison of the effectiveness of two methods of formalization of voice interaction
Ivan Naydonov · Bulletin of the National Technical University «KhPI» Series New solutions in modern technologies · 2018
The article is devoted to the study of the effectiveness of formalization of voice interaction without the transformation of voice information into text, based on the use of a reflex voice control system consisting of a phonemic transcript that converts a sound recording to a phonemic representation, and a classification core that determines the content and control of the received phonemic set. The purpose of the paper is to compare the effectiveness of machine learning methods for formalizing voice interaction on an example of a support system for vehicle dispatching using a reflex voice control system. In order to verify the effectiveness of the constructed models, an iterative process of data collection (in accordance with the model of voice interaction in the form of a tree of scenarios) and formalization modeling was carried out, which included analysis of the results and the expansion of metrics for the accuracy of the evaluation for unbalanced samples (precision, recall, F-score). At the initial stage, voice data of 23 speakers was collected, with an average of 45 samples per reaction. The simulation results on a minimum set of data by both methods showed an accuracy of no more than 50%, which is insufficient for practical application. On the second stage, an average of 310 voice samples were collected for each of the 3 simple-context reactions, a total of 925 reactions. The simulation by the method of intelligent reflex systems showed a accuracy of about 60%, which is also insufficient, and the accuracy of method of convolutional neural networks is slightly more than 90%, which is acceptable. In order to confirm the efficiency of the method of intelligent reflex systems, two stages was insufficient, the hypothesis about insufficient quality of sound recordings and high level of noise as obstacles to the effectiveness of the formalization model was advanced, prospects for conducting the next stage of the research were outlined. A conclusion is made about the effectiveness of the reflex voice control system and its ability to determine in practice the content and control of the received phonemic set without converting the voice information into a text form.