Deep segmental neural networks for speech recognition

Ossama Abdel‐Hamid, Li Deng, Dong Yu, Hui Jiang · 2013

Hybrid systems which integrate the deep neural network (DNN) and hidden Markov model (HMM) have recently achieved re-markable performance in many large vocabulary speech recog-nition tasks. These systems, however, remain to rely on the HMM and assume the acoustic scores for the (windowed) frames are independent given the state, suffering from the same difficulty as in the previous GMM-HMM systems. In this pa-per, we propose the deep segmental neural network (DSNN), a segmental model that uses DNNs to estimate the acoustic scores of phonemic or sub-phonemic segments with variable lengths. This allows the DSNN to represent each segment as a single unit, in which frames are made dependent on each other. We describe the architecture of the DSNN, as well as its learning and decoding algorithms. Our evaluation experiments demon-strate that the DSNN can outperform the DNN/HMM hybrid systems and two existing segmental models including the seg-mental conditional random field and the shallow segmental neu-ral network.

Read the paper · More papers on PaperTik