Implementation of a floating-point matrix-vector multiplication on a reconfigurable architecture
Fabio Garzia, Claudio Brunelli, Davide Rossi, Jari Nurmi · Proceedings - IEEE International Parallel and Distributed Processing Symposium · 2008
This paper describes the implementation of a floating-point 4times4 matrix-vector multiplication on a reconfigurable system. The 4times4 matrix-vector multiplication is meant to be used to perform two steps (transformation and perspective projection) of a 3D graphics application. The target system is based on a bus architecture with a general purpose core as master and the reconfigurable array as main accelerator. The system has been prototyped on a FPGA device. The matrix-vector multiplication has been successfully implemented on the reconfigurable block. Compared to the general purpose implementation it is convenient if the number of vectors to process is higher than seven. If hundreds of vectors are processed, the speed-up achievable reaches 3times.