Estimation of instants of significant excitation using accumulated energy function of DCT

N. Sripriya, T. Nagarajan · 2013

This correspondence proposes an effective algorithm for estimation of instants of significant excitation that are exclusively present in the voiced speech. The proposed method is based on the idea that the fundamental frequency signal can be generated by the reconstruction of the signal using the first few DCT coefficients. The critical window including the required number of DCT coefficients is identified using the accumulated energy function (AEF) of DCT. This AEF function is evaluated by computing the energies of the signals constructed by expanding the window each time to include the next DCT coefficient. Many work reported in the literature yield better estimates of these instants of excitation present in the voiced speech. However, their estimations are not restricted to voiced regions alone causing spurious instants in non-voiced regions. The major advantage of this method is its inbuilt extended ability to differentiate voiced/non-voiced speech preventing spurious instants in non-voiced regions. The well-known method for instants estimation, DYPSA, is reviewed, evaluated and, the new algorithm is compared with it using the CMU Arctic database. The results clearly show that the generation of 85% of the spurious instants in non-voiced region are avoided without any serious compromise in the identification rate(IDR). Moreover, this technique is time-effective and has a tuning parameter to control the miss rate(MR) and the false alarm rate(FAR).

Read the paper · More papers on PaperTik