Big Data Analytics using Artificial Intelligence: Apache Spark for Scalable Batch Processing
Himanshu Gupta · International Journal of Innovative Science and Research Technology (IJISRT) · 2024
The rapid proliferation of data in the digital age has made big data analytics a critical tool for deriving insights and making informed decisions. However, processing and analyzing large datasets, often reaching hundreds of terabytes, presents significant challenges. This paper explores the use of Apache Spark, a powerful distributed computing framework, for batch processing in big data analytics using artificial intelligence (AI) techniques. We evaluate the scalability, efficiency, and accuracy of AI models when applied to massive datasets processed in Spark. Our experiments demonstrate that Apache Spark, coupled with machine learning and deep learning techniques, offers a robust solution for handling large-scale data analytics tasks. We also discuss the challenges associated with such large-scale processing and propose strategies for optimizing performance and resource utilization.