Speech Recognition with Gender Identification and Speaker Diarization
M. Naveen, Abraham Sudharson Ponraj · 2020 IEEE International Conference for Innovation in Technology (INOCON) · 2020
Speech recognition is the ability of a machine or program to identify words and phrases in spoken language and convert them to a machine-readable format. Capturing the voice using the mic from all direction. Direction of arrival is being recorded with an angle of 360 degree. The voice is being detected using Voice Activity Detection to avoid capturing the noise in the surrounding. Captured voice is being processed using Acoustic Echo Cancellation. The voice is being analysed using speech algorithm. Voice is being separated based on their decibels (db). In the proposed system the voice is being recorded and is being converted to transcripts by processing it with various speech models. The Automatic Speech Recognition (ASR) is the process of converting an unknown speech waveform into the corresponding orthographic transcription. The speech signal using the algorithms of analysis synchronized with the pitch frequency. The online recogniser is being used reads a list of wav files and outputs the result into a text format. Speaker Diarization identifies who speaks when during the conversation on the recorded wav. The model attains the accuracy of 92 % for speech recognition. This research can further used in transcription platform and other smart voice devices.