Top-Down Speech Recognition

Robert H. Chen, Chelsea Chen · 2022

The amplitudes of sound as a function of frequency is called sound spectra that shows the discrete location of formants that reveal the speech characteristics of loudness, pitch, intonation, and accent. Human hearing is critically dependent on the perception of proportion, which is more distinctly manifested by a logarithmic rather than a linear scale. The Fourier coefficient formulas represent the components of a waveform by extracting a single period of the wave and finding the area of that period for a given frequency, one frequency at a time. The technology generally comprised a bank-of-filters front-end analyzer for first separating the very different voice pitches, such as men from women, and producing a set of signals representing the energy of a sound in a given frequency band, thereby creating the sound spectra of an utterance. One can easily form waveforms on a graphics calculator or personal computer by just adding the sine and cosine functions with different amplitudes and frequencies.

Read the paper · More papers on PaperTik