BigThrill: MPI-based Data Processing Engine
Anastasia Khartikova, Denis Shaikhislamov, Ilya Timokhin, Roman Olegovich Kostromin, Vladislav Muratov, Aleksey Demakov, Maxim Belov, Aleksey Teplov · 2024
Several engines have been proposed already to analyze data coming from multiple sources, known as Big Data, with the most common approach being the Map Reduce model, which is heavily utilized in the state-of-the-art solutions, such as Spark, and in various ab initio written tools, such as the Thrill Project. Here we report our enhancement to the Thrill framework, where the network communication, memory and API layers have been rewritten to utilize hybrid MPI-OpenMP model. Moreover, we present our engine as part of a greater framework, with dynamic and static process management (MPI) and input optimization, making it a tool that can be incorporated into any major platform. Our objective was to provide better parallelization and scalability, as well as to utilize the power of MPI in the process interaction and operation. According to the data collected while testing in nearly real-life scenarios on multiple clusters, we reached the performance boost up to 35% compared to the pure TCP solution, and increased scalability. The further research includes testing novel fault tolerance solutions and tweaking MPI libraries.