Data Integration Tasks on Heterogeneous Systems Using OpenCL
Clayton J. Faber, Anthony M. Cabrera, Orondé Booker, Gabe Maayan, Roger D. Chamberlain · 2019
In the era of big data, many new algorithms are developed to try and find the most efficient way to perform computations with massive amounts of data. However, what is often overlooked is the preprocessing step for many of these applications. The Data Integration Benchmark Suite (DIBS) [1] was designed to understand the characteristics of dataset transformations in a hardware agnostic way. While on the surface these applications have a high amount of data parallelism, there are caveats in their specification that can potentially affect this characteristic. Even still, OpenCL can be an effective deployment environment for these applications.