Distributed Data Streams Processing Based on Flume/Kafka/Spark
Jun Wang, Wenhao Wang, Renfei Chen · 2015
Designed and implemented a distributed data streams processing system based on Flume, Kafka and Spark, fetch and analyze datastreams and mining business intelligence information efficiently, real-timely and reliably, With high scalability and high reliability of Flume, the data of multiple sources can be collected accurately and extended easily.Kafka's characteristics of high throughput, scalability, distribution meet the distribution requirements of massive data.Spark Streaming provides a set of efficient, fault-tolerant and real-time large-scale stream processing frame.Thereby services and strategyof enterprise can be improved.