Apache Kafka on Big Data Event Streaming for Enhanced Data Flows
K. Padmanaban, Tummala Ranga Babu, K. Karthika, Balachandra Pattanaik, K Dhanabhavithra, Charru Pooja Srinivasan · 2024
Apache Kafka's distributed architecture and message queuing capabilities offer significant improvements in real-time and batch data processing efficiency and reliability. This research aims to optimize Kafka setups, data partitioning, and Kafka Connect integration to create a robust and scalable data streaming infrastructure. The study focuses on enhancing data input, processing, and dissemination across systems and applications. By optimizing Kafka configurations and leveraging its capabilities, the research seeks to achieve significant improvements in data processing speed, real-time analytics, and scalable data pipelines. The evaluation of Kafka Event Stream Throughput Over Time and Latency Distribution Across Brokers demonstrates the system's performance and efficiency. The results indicate a throughput of 750-1340 events per hour and a latency distribution of 6-15 milliseconds. Additionally, Consumer Lag Over Time analysis reveals consistent performance with values ranging from 70-140. This research contributes to the advancement of big data processing by demonstrating the effectiveness of Apache Kafka in creating a robust and efficient data streaming infrastructure. The findings provide valuable insights for organizations seeking to optimize their data pipelines and leverage the power of real-time analytics.