Enhancing Efficiency in Large Scale Data Processing: Optimizing Cluster Compute and Storage Resources

Venkateswarlu Chennareddy, Rama C Koppula, Pavan Kumar Patibandla · 2024

In an era dominated by Artificial Intelligence (AI) and Big Data, managing and generating foundational data for large-scale data models presents significant challenges in terms of resources and costs. This paper introduces a novel method aimed at optimizing cluster compute and storage resources, critical for AI applications requiring extensive data processing. By utilizing a custom Stratified Sampling algorithm along with a fully automated, independent, and self-adaptive data comprehensive test suite module, this approach efficiently produces a minimized dataset while maintaining essential coverage of business scenarios. This method not only ensures the integrity and representativeness of the data required for business use cases, but also promises resource reduction, with an estimated 80-85% savings in cluster compute and storage resources. Addressing a pressing need in data processing, this method emerges as a crucial asset for entities seeking to enhance their data management capabilities while ensuring data quality and comprehensive in an AI-centric landscape.

Read the paper · More papers on PaperTik