Distributed Data Processing for Large-Scale Simulations on Cloud
Tianjian Lu, Stephan Hoyer, Mengqing Wang, Lily J. Hu, Yifan Chen · 2021 IEEE International Joint EMC/SI/PI and EMC Europe Symposium · 2021
The computational challenges encountered in the large-scale simulations are accompanied by those from data-intensive computing. In this work, we proposed a distributed data pipeline for large-scale simulations by using libraries and frameworks available on Cloud services. The building blocks of the proposed data pipeline such as Apache Beam and Zarr are commonly used in the data science and machine learning community. Our contribution is to apply the data-science approaches to handle large-scale simulation data for the hardware design community. The data pipeline is designed with careful considerations for the characteristics of the simulation data in order to achieve high parallel efficiency. The performance of the data pipeline is analyzed with two examples. In the first example, the proposed data pipeline is used to process electric potential obtained with a Poisson solver. In the second example, the data pipeline is used to process thermal and fluid data obtained with a computational fluid dynamic solver. Both solvers are in-house developed and finite-difference based, running in parallel on Tensor Processing Unit (TPU) clusters and serving the purpose of data generation. It is worth mentioning that in this work, the focus is on data processing instead of data generation. The proposed data pipeline is designed in a general manner and is suitable for other types of data generators such as full-wave electromagnetic and multiphysics solvers. The performance analysis demonstrates good storage and computational efficiency of the proposed data pipeline. As a reference, it takes 5 hours and 14 mins to convert simulation data of size 7.8 TB into Zarr format and the maximum total parallelism is chosen as 10,000.