Big Data Architectures
Dominik Ryżko · 2020
This chapter discusses the most common computation models applicable for processing large data sets mostly in the batch mode. These include MapReduce, Directed Acyclic Graph models, and All-Pairs. Publish-subscribe systems provide loosely coupled methods for data transmission; one of the popular publish-subscribe systems is Apache RabbitMQ. The chapter is devoted to a stream processing, which is gaining enormous interest both from the research community and some major commercial vendors. It also discusses some higher level big data architectures like Spark and Lambda. The chapter shows the overview of the system involved in big data processing. Scribe is a persistent, distributed messaging system and is responsible for collecting, aggregating, and delivering high volumes of log data to real-time and batch systems. It presents an example of an actor-like model for big data processing. The system was designed to handle the variety out of the big data 3V: volume, velocity, and variety.