GPU-Based Data Science for High Productivity
Tirthajyoti Sarkar · Apress eBooks · 2022
In the last two chapters, you learned about various tools and frameworks for doing out-of-core and distributed/parallelized data science. The central goal has always been the same: enhancing the productivity of the data science pipeline. Productivity is often directly related to the speed of execution of various DS tasks including numerical processing, data wrangling, and feature engineering. When it goes to the advanced machine learning stage, depending on the modeling complexity, the matter of speed and performance assumes even a critical role. The history of machine learning has clearly demonstrated that the use of specialized hardware like the graphics processing unit (GPU) played a significant role in the early success of ML. While the use of GPUs and distributed computing is widely discussed in academic and business circles for core AI/ML tasks, they get less coverage when it comes to their utility for regular data science and data engineering tasks. However, the important question remains: can we leverage the power of GPUs for regular data science jobs (e.g., data wrangling, descriptive statistics) too? The answer is not trivial and needs some special consideration and knowledge sharing. In this chapter, I will focus on a specialized suite of tools called RAPIDS that helps any data scientist take advantage of GPU-based hardware for a wide variety of data science tasks.