Spearheading Big Data Solutions: Optimizing Data Pipelines For Enhanced Efficiency And Performance
Kiran Polimetla, Farah Jenny · 2024
For a big data solution to work effectively, the following needs to be addressed: Infrastructure to store and process vast amounts of data. Geographic disorientation of experts over petascale data pipelines, thereby stunting the development of real-world use cases. By enabling data scientists to practice data science, optimizing data processing pipelines, and harmonizing tools/algorithms for scaling data structures, it is necessary to avoid reinventing the big data solution wheel for each problem. Real-world use cases illustrate how the advent of Google's Big Query as a more democratized create-sarge maintain-update-a-data-warehouse-and-run-queries software for shared HEP CERN data samples has significantly boosted confidence in adopting the big data philosophy. In big data solutions, emphasis must be placed on harmonization and correctness within the data workflows. CERN's data sample requirements, gathered under different research programs, result in petascale datasets. This usually leads to petascale databases with infrastructure requirements specializing in storage or machine learning processes using the data. Google's BigQuery resolves this dichotomy by allowing scientists to construct machine learning models using big datasets and perform sub-second queries. This paper addresses what has happened in mapping big data technologies to petascale data, the importance of successful and efficient implementation of a data workflow, large-scale interactive analysis, and the divergence of needs in exact-exact-extract adoption. Results are shown on petascale CERN data samples collected as part of the previous collaboration with the Compact Muon Solenoid (CMS) experiment to increase the use of the workspace in the CMS.