Advanced Speech Processing for Speaker Authentication in Communication Systems
Vinod Gujral, Jetendra Joshi, Praveen Medikonda, Nishchay Grover · 2018
In this era of artificial intelligence and machine learning, there are many small as well as large-scale companies and organizations, which are working on projects to produce goods and services, these are intelligent enough to perform task on their own. The major focus in building intelligent machines leads to the replication of senses of hearing and sight, which involves the use of audio and image processing. Through this paper, we are proposing a technique devised by us during our research and development project based on research we did on existing techniques in terms of their performance and how we can use them in a combination. The technique, which we have proposed in this paper, is for recognizing and authenticating the speaker based on both frequency and phase features of the speech and analyzing it on frame-by-frame basis. It has often to see the problem of corrupted/synthetic speech, which allows unrestricted access in different situations or in simple terms, we can say that anyone can spoof the system. This happens due to limited dataset or a dataset, which contains very less samples of synthetic speech. Due to this, it becomes difficult to train the model and we get negative results. Therefore, in this research we have devised a technique and tested it to eliminate such problems. Further, we have utilized this technique in the implementation of sending voice messages in a Vehicular Ad-Hoc Network (VANETs) and because of this technique its easier for a driver to trust on the message sent by an another driver in the network.