Column Storage for FPGA-accelerated Data Analytics
David Sidler · Repository for Publications and Research Data (ETH Zurich) · 2013
Data appliances became very popular in recent years.To decrease the network traffic between processing nodes and the disk, some modern systems use smart storage engines to off-load filtering query operators, such as selection or projection.This reduces the amount of data that has to be transferred from the storage engine to the processing nodes.Ibex is an intelligent storage engine that is able to off-load query operators to a FPGA.This work presents a column storage implementation on a FPGA which replaces the existing row storage in Ibex.Memory on a FPGA is limited therefore it is only possible to load small data chunks from the disk (e.g. 4 KB).Since we are reading multiple columns which are stored at different locations on the disk, we have to deal with random access.Using a SSD and native command queuing (NCQ) it is possible to achieve high throughputs despite reading small data blocks.NCQ returns the data out of order, therefore a complex buffering system was implemented to reorder the data on the FPGA.We present an economical hardware design which is able to adapt to the number of columns processed, thereby optimizing hardware utilization and improving performance.Additionally we implemented two simple statistical operators which are applied to column-shaped data passing the FPGA.Due to the parallelism in FPGAs no additional latency is added to the data processing.i iii 5.4.5 Resource consumption . . . . . . . .