Signal processing algorithms and architectures
Hassan Masud Ahmed · 1982
The advent of the Very Large Scale Integration (VLSI) technology has provided the ability to construct large systems on a single silicon chip. This dissertation is concerned with exploiting this ability to design a powerful signal processing chip capable of efficiently implementing such popular algorithms as the discrete Fourier transform, ladder filters and associated matrix algebra operations. The latter include Givens rotations and Cholesky factorization. The goal of the present work is to efficiently map algorithms onto architectures by maintaining a close link with the theoretical basis of a particular signal processing method. It is shown that all of the algorithms considered can be cast into a mathematical framework involving generalized vector rotations. Such rotation operations provide a natural description of the algorithms and the computational complexity measured in terms of these elementary operations is much lower than in terms of the usual measure of total number of multiplications. Thus, unlike present day signal processing computers which emphasize rapid multiplication, the signal processing architectures in this thesis are based on the ability to perform vector rotations in generalized coordinate systems. It is shown that the CORDIC algorithm of Volder provides a convenient implementation of vector rotations with only simple components such as adders, registers and shifters. Unfortunately, throughput is severely compromised owing to the need for performing special operations to account for the limited region of convergence and spurious scale constants inherent to the method. New techniques to circumvent these problems with no additional hardware and only a marginal speed penalty are described. Further speed enhancements through the use of a newly developed method known as hybrid CORDIC are discussed. Additionally, floating point CORDIC (FLORDIC) algorithms that are conceptually simpler than their fixed point counterparts are developed and the connection of CORDIC to the convergence computation methods is shown. The architecture of a dual CORDIC block chip is described for a target application of real time speech analysis. The resulting chip is shown to have a higher throughput per area than conventional chips based on fast multiplications. This is attributed to the close match of the present chip to the algorithms. Large mesh connected processor architectures for matrix factorization are developed which are also closely matched to the algorithms. Individual processing elements in the mesh are based on CORDIC operations, in fact on the aforementioned signal processing chip. Finally, a new technique for signal detection in additive Gaussian noise is developed with a view towards ease of implementation. It is based on ladder filters and may be implemented using the signal processing chip mentioned above.