Systolic devices for speech processing
Bryan Beresford‐Smith, Jens Breckling, H. Schroder · 2003
A frequently used method of representing words in the area of speech recognition is as a sequence of LPC (linear prediction coefficient) vectors, which are real-valued 12- to 16-dimensional vectors each representing a time slice of the speech signal. It can be associated with particular phonems (or transitions between phonems). One method of transferring speech in a real-time environment is through the use of codebooks; i.e. a set of representatives of LPC vectors. The method proposed is based on the following idea: a sequence of N LPC vectors is produced out of a few minutes of speech. Then a distance matrix is calculated using any suitable distance function such as the Itakura distance. Then this distance matrix can be analyzed by repeatedly looking for rows with a maximal number of small-valued entries. These rows refer to clusters in the LPC vector space and thus can be assumed to be good representatives to be used as members of the codebook.>