Extracting nasality from speech signals

Henning Reetz · The Journal of the Acoustical Society of America · 1990

“Nasality” is commonly defined as resulting from speaking with a more or less lowered velum. Although this phenomenon appears in every language, sometimes as a distinctive feature, and can be perceived clearly, there is no reliable way to automatically extract it from the speech signal. An algorithm is being developed that segments the speech waveform into “nasal” and “non-nasal” parts. The algorithm works without presegmenting and does not need information about whether nasal sounds are present or not. The algorithm searches voiced areas in the speech signal by means of a pitch extraction algorithm. In these areas, a resonance in the range of 200–400 Hz is searched. The frequency of this resonance may vary between speakers, but is assumed to be constant for individual speakers. If this resonance with constant frequency persists for more than 100 ms, the area is labeled nasal. The bandwidth of the low resonance can be used as a rough indicator for the amount of nasality existing in the speech signal.

Read the paper · More papers on PaperTik