An Efficient Exact Fused Dot Product Processor in FPGA
Luís Fiolhais, Horácio C. Neto · 2018
The objective of this work is to develop an efficient hardware processor to perform the fused dot product operation with full accuracy. General purpose processors split the dot product into two independent operations, rounding the full precision results of the multiplication and accumulation to the formats precision, which may significantly affect the overall accuracy. Therefore, many systems must rely on software approaches to achieve numerical accuracy, which imposes a significant performance overhead. This work proposes a fused dot product processor which is able to output a partial result per cycle with exact operand arithmetic. The processor uses a long-accumulator with full fixed-point precision to achieve full accuracy. Signed addition support is achieved using a Generalized Signed-Digit redundant numeric representation, which avoids sign extending between segments every cycle. In addition, the proposed architecture uses a novel autonomous carry propagation unit, which searches, selects and propagates stray carries between segments, while the accumulation process is executing. The proposed processor has been implemented in a Zynq 7020-1 FPGA, where a single-precision core can execute with a throughput of one partial result per 90 MHz clock cycle and occupies about 5K LUTs. Full accuracy has been demonstrated using a set of hard to correctly solve benchmarks.