The arithmetic cube II and memory-based architectures for data structure manipulation

Thomas P. Kelliher · 1993

The Arithmetic Cube II (AC II) is a digital signal processing system, designed to compute the discrete Fourier transform and cyclic convolutions. The architecture is the first which implements the so-called small-n algorithms. It is capable of computing a 1008 point complex-in, complex-out DFT in 3.5 ms., a rate equivalent to commercial systems running at much higher clock frequencies. University-build systems such as the AC II, although currently rare, demonstrate that it is possible for universities to build complete systems, from custom chips to user-level software. This situation will result in numerous advances in computing as application specific systems are designed and built to solve old problems in new ways. Although the AC II's primary purpose was as a proof of concept design of a small-n algorithms architecture, the goal of demonstrating a complete system seems greater in comparison. One of the key design problems which faced the author was where and how to transpose the two-dimensional matrices which are the AC II's data. It turned out that the best place for performing the transposition operation was within the memory sub-system. The design work on the memory sub-system of the AC II led the author naturally to searching for transposition and rotation architectures for other architectures. Transposition can become significantly more difficult for more general dataflows. This thesis shows that the problem of matrix transposition can be solved AT$\sp2$ optimally by a VLSI architecture which is completely systolic and has very simple control requirements. This thesis also shows that the transposer architecture may be extended to handle the rotation of multi-dimensional cubes (matrices). This is a critical sub-operation in multi-dimensional signal processing problems. The resulting VLSI architecture is also AT$\sp2$ optimal for cube rotations. Both architectures are completely programmable as to the length of dimension of the input. In addition, the rotator is programmable as to the number of dimensions. The additional routing network required by the rotator may be mapped directly onto the routing network of the transposer; only additional control logic is required.

Read the paper · More papers on PaperTik