FPGA implementation of a binary32 floating point cube root
Carlos Minchola Guardia, Eduardo Boemo · 2014
This paper presents the implementation of a sequential hardware core to compute a single floating point cube root compliant with the current IEEE 754-2008 standard. The design is based on Newton-Raphson recurrence, reciprocal and cube root units are implemented. Optimal performance requires two iterations for reciprocal and one for cube root units obtaining an accurate approximation of +/- 3 least significant bits. Our proposal is able to be performed up to 149 Mhz over Virtex5. The hardware cost occupies 230 Slices and 12 Dsp48s taking a latency of 19 clock cycles.