GMM and HMM training by aggregated EM algorithm with increased ensemble sizes for robust parameter estimation
Takahiro Shinozaki, Tatsuya Kawahara · IEEE International Conference on Acoustics Speech and Signal Processing · 2008
In order to compensate for the weaknesses of the expectation maximization (EM) algorithm to over-training and to improve model performance for new data, we have recently proposed aggregated EM (Ag-EM) algorithm that introduces bagging-like approach in the framework of the EM algorithm and have shown that it gives similar improvements as cross-validation EM (CV EM) over conventional EM. However, a limitation with the experiments was that the number of multiple models used in the aggregation operation or the ensemble size was fixed to a small value. Here, we investigate the relationship between the ensemble size and the performance as well as giving a theoretical discussion with the order of the computational cost. The algorithm is first analyzed using simulated data and then applied to large vocabulary speech recognition on oral presentations. Both of these experiments show that Ag-EM outperforms CV-EM by using larger ensemble sizes.