A Comparative Study On Real Time Data Streaming For Fraud Detection Using Kafka With Apache Flink And Apache Spark
Rayhan Prawira Daksa, Ade Putera Kemala · Procedia Computer Science · 2025
The Increasing volume and velocity of digital financial transactions have heightened the need for responsive and scalable fraud detection systems. This paper presents a comparative study of two stream processing frameworks. Apache Flink and Apache Spark within a Kafka based real-time fraud detection architecture. Both pipelines were implemented under identical conditions and integrated with a pretrained Random Forest model which was selected based on its superior offline performance among five evaluated machine learning classifiers. Performance benchmarking focused on latency and throughput under two different data ingestion rates, 10 and 40 transactions per second. With evaluation focusing on latency and throughput. Experimental results show that Apache Spark achieved consistently a lower average latency 0.8 seconds at 10 TPS and 0.9 seconds at 40 TPS compared to Apache Flink, which recorded 1.7 and 1.6 seconds respectively. Throughput for both frameworks remained stable underload, demonstrating scalability. These findings provide empirical insights to support informed decision making in selecting stream processing technologies for real-time, missing critical fraud detection use cases.