On context-dependent neural networks and speaker adaptation
Jan Zelinka, Jan Trmal, Luděk Müller · 2012
This paper describes evaluation of a neural network based hybrid LVCSR system. The novelty of the evaluated hybrid system lies in speaker adaptation techniques that are employed to increase performance of neural networks for context-dependent phonetic units modeling. The performance comparison is done as follows. First, performances of different hybrid systems employing either a context-independent neural network or a context-dependent neural network are compared. Second, the influence of the recently published speaker adaptation technique called MELT is evaluated. Furthermore, several possible approaches to conversion of posterior probabilities into observation likelihoods, which are necessary for a hybrid LVSCR systems, are described and discussed in this paper.