Optimized Fused Floating-Point Many-Term Dot-Product Hardware for Machine Learning Accelerators

Himanshu Kaul, Mark A Anders, Sanu K. Mathew, Seongjong Kim, Ram Kumar Krishnamurthy · 2019

This paper describes optimizations for the critical maximum exponent and alignment operations, with scalability for many-term fused floating-point dot-product units. The impact of these optimizations is quantified for up to 32-term BFloat16 weight/activation inputs with single-precision dot-product output, targeted for machine learning accelerators. Area and energy efficiency results are compared across performance targets, design parameters, and data statistics.

Read the paper · More papers on PaperTik