German Speech Recognition System using DeepSpeech
Jiahua Xu, Kaveen Matta, Shaiful Islam, Andreas Nürnberger · 2020
Speech recognition focus on the translation of speech from an audio format to a text. Popular models are available for the English language as open source in the domain of voice/speech recognition; however, German language open models and training schemes are rather rare. An end-to-end real-time German speech-to-text system based on multiple German language datasets is worthy of more attention and further investigation. In this paper, we combined multiple German datasets on the market and optimizes the Deep-speech for training a real-time German speech-to-text model. A GUI is also proposed for functionality demonstration. Our model performs considerably well compared to other state-of-the-art since we utilized noisy data to replicate real-life scenarios. We released our fully trained German model along with its parameter configurations to promote the diversification of the open-source model for the German language.