Stream Big Data Processing

Yusuf Aytas · 2021

This chapter examines stream data processing technologies in depth. Stream processing can help a modern Big Data platform from many different perspectives. The chapter employs a new set of solutions with stream Big Data processing. Some of these solutions are as follows: fraud detection, anomaly detection, alerting, real-time monitoring, and instant machine learning updates. The chapter discusses stream processing through message brokers. Message brokers can validate, store, and route messages from one publisher to many subscribers. The chapter covers popular message solutions implemented over Kafka, Pulsar and presents brokers that adopt AMQP. Kafka Streams use the local state for implementing stateful applications. Stream processing engines provide distributed, scalable, and fault-tolerant scaffolding for users to build streaming applications. The chapter presents widely adopted stream processing engines such as Flink, Storm, Heron, and Spark Streaming. A streaming application needs checkpoints to store the state in a fault-tolerant storage system to recover from failures.

Read the paper · More papers on PaperTik