GPU acceleration of automated speech recognition for mobile devices
Richard Veitch, Roger Woods, Louis-Marie Aubert · 2011
The implementation of a complex, large vocabulary, speech recognition application on a modern graphic processors (GPUs) is presented. The parallel single instruction, multiple data (SIMD) architecture is effectively exploited by performing various optimizations to expose the algorithmic parallelism. The work addresses particularly the realization of the Gaussian calculation, a key function. The result is an implementation that runs 3.75 faster than real-time and gives a tenfold speedup when compared to a highly optimized sequential CPU-based implementation. The work is also compared with some earlier work involved in building the same system on a Virtex 5-based, Alpha Data XRC-5T1 reconfigurable computer.