A Study of a Scalable Distributed Stream Processing Infrastructure Using Ray and Apache Kafka
Kasumi Kato, Atsuko Takefusa, Hidemoto Nakada, Masato Oguchi · 2018
The spread of various sensors and the development of cloud computing technologies enable the accumulation and use of many live logs in ordinary homes. In addition, deep learning technologies have been widely used for image and speech recognition processing. However, a key issue for deep learning is heavy processing loads. To operate a service that utilizes sensor data, those data are transmitted from sensors in ordinary homes to a cloud and analyzed in the cloud. However, services that involve moving image analysis require large amounts of data to be transferred continuously and high computing power for the analysis; hence, it is difficult to process them in real time in the cloud using a conventional stream data processing framework. First, we perform preliminary experiments using Apache Spark [3] (hereinafter called Spark), which is a representative cluster computing platform that is designed to be fast and versatile, and Ray [4] , which is a distributed execution framework. We investigate the characteristics of their distributed recognition processing and demonstrate that Ray enables scalable distributed processing. Next, We implement a prototype system of the proposed distributed stream processing infrastructure using Ray and Apache Kafka [1] (hereinafter called Kafka), which is a distributed messaging system, and demonstrate its performance.