Leveraging Cloud-Native Data Engineering for Big Data Analytics
Shubham Gupta, Munikrishnaiah Sundararamaiah, Geeta Geeta · 2025
Recent evolution of traditional ETL systems into a cloud-native data engineering design for big data analytics capable of elastic and low-cost scaling. The proposed pipeline combines Apache Spark on Kubernetes, AWS Glue, and Delta Lake by leveraging microservices, containerization, serverless computing, and distributed orchestration. Based on the Taxi Trip dataset, the framework allows up to several times reduction in query execution time, processing throughput, and cost efficiency as compared to typical ETL approaches. Thus, the experimental results show that the query response is 75.9% faster and the cost is reduced by 64.5%. This research highlights the advantages of cloud-native architectures in maximizing real-time analytics and details in future work auto-scaling driven by AI and hybrid edge-cloud models.