Modified Fused Multiply and Add for Exact Low Precision Product Accumulation
Nicolas Brunie · 2017
The implementation of the Fused Multiply and Add (FMA) operation has been extensively studied in the literature on standard and large precisions. We suggest re- visiting those studies for 16-bit precision. We introduce a variation of the Mixed precision FMA targeted for applications processing low precision inputs (such as machine learning). We also introduce several versions of a fixed point based floating- point FMA which performs an exact accumulation of binary16 numbers. We study the implementation and area footprint of those operators in comparison with standard FMAs.