High speed computer arithmetic architectures
Hosahalli R. Srinivas · 1994
Redundant arithmetic number systems are gaining popularity in computationally intensive environments particularly due to their carry-free addition/subtraction property. In carry-free addition, the length of carry-propagation is independent of word-length. This property has enabled arithmetic operations to be performed much faster than with conventional binary number systems, where the delay of an operation depends on the propagation of carry signals. As a result, the floating point unit of most of the high-performance computing chips use redundant arithmetic. Also, the need for high-sample rate operations in digital signal processing (DSP) has resulted in attempts to develop architectures based on redundant arithmetic. To this end, we have developed new algorithms and architectures for basic digital arithmetic operations. These involve fixed point, floating point, and pipelined architectures. The architectures proposed here are better either in speed or area when compared with previously known architectures. We have developed fast (low latency) pipelined architectures for fixed point division, square root, and multiplication using a new scheme to convert radix 2 representation to two's-complement binary. All these units use radix 2 redundant arithmetic internally but two's-complement binary representation for input-output communication. We have also developed a fast architecture for binary addition (with carry-propagation) of 2 two's-complement binary numbers. Fast new division algorithms and architecture based on radix 4 (with maximally redundant quotient digit set) and radix 2 (with over-redundant radix 2 quotient digit set) redundant arithmetic have also been developed. These architectures are faster than previously known division architectures. The radix 4 division architecture incurs some acceptable hardware penalty. We have also developed a new fast radix 2 based square root algorithm for floating point numbers. This architecture allows two different square root algorithms to share the same hardware to accommodate both even and odd exponents. It is faster than previously proposed square root architectures but with acceptable hardware penalty. To demonstrate the capabilities of the proposed algorithms, we have implemented a 1.2 micron CMOS VLSI chip that performs radix 2 division and square root by sharing the same hardware. This chip operates on the mantissas of the IEEE 754 std. 1985 single precision numbers.