A Speed Optimized Multifunctional MAC Architecture
Lin Chuan · Science Technology and Engineering · 2006
An algorithm and an afficient architecture for design of a 40±16×16 multifunctional multiply-accumulate (MAC) unit are presented, which is optimized for speed and can support signed/unsigned/mixed multiply/multiply-accumulate operations with various rounding and saturating options. Basically, this algorithm is developed on the modified Booth’s algorithm and Wallace Tree algorithm by making several improvements to them. It simplifies the generation of the partial products, the signs extension optimizes the connection of the Wallace Tree and the orders of many computing operations. As a result, the parallel MAC unit produced by the proposed algorithm outperforms those not improved in our experiment. These findings have already been implemented as a part of a 16-bit general high performance DSP and now prepared for its tape-out and succeeding verification on FPGA.