Scheduling edge and in-transit computing resources for stream processing applications
Zamani Zadeh Najari · 2019
The exponential growth of digital data sources has the potential to transform all aspects of society and our lives. However, to achieve this impact, the data has to be processed in a timely manner to extract insights that can drive decision-making. With the increasing availability of Internet of Things (IoT) devices, and potential applications that make use of data from such devices, there is a need to better identify appropriate data processing techniques that can be applied to this data. The computational complexity of these applications and the complexity of the requirements on the data processing techniques often derive from the capabilities of current IoT devices and the need to integrate data streams across multiple IoT devices. This results in larger data sizes and loads on the computing infrastructure. While cloud computing is able to provide flexible computing and storage services, important factors such as bandwidth provisioning between different components, leveraging network computing resources along the data path, and utilizing heterogeneous resources based on their geographic location to deploy the workflows are not supported in current cloud computing approaches and models. In fact, due to the data movement costs, traditional approaches that rely on moving data to remote data centers for processing are no longer feasible. Instead, new approaches that effectively leverage distributed computational infrastructure and services are necessary. Specifically, these approaches must seamlessly combine resources and services at the edge, in the core, and along the data path as needed.To address these challenges, this dissertation explores various approaches to eliminate or alleviate the impact of limited resources to process large amounts of data. First, programming network resources using Software Defined Network (SDN) capabilities is added to the cloud federation to gain control over networking infrastructure. Second, having control over network computing resources (in-transit resources) enables us to provision latent resources at network nodes and process data using such resources by considering location and network properties of the resource and the flow of data. Finally, a novel data delivery approach is developed that uses heterogeneous resources located at the edge of the network and along the data path to process big data streaming applications and deliver the processed data to users while considering users constraints.The main contribution of this dissertation is the integration of SDN and cloud federation that enables the provisioning of the in-transit network resources and provides more information about networking resources for deployment. Another contribution of this dissertation is the design and development of a large-scale publish/subscribe messaging system for data movement using CDN nodes to seamlessly route the data between components and automatically deploy stream-oriented workflows on geo-distributed heterogeneous nodes.The validation of these works is done through a series of experiments based on real scenarios and applications. Three main applications have been considered as use cases of these works: (1) Data gathered from instrumented build environment applications, (2) Video analytics applications for data collected by video surveillance cameras (3) Real-time data from Ocean Observatory Initiatives (OOI). Heterogeneous geographically-distributed resources with different capabilities, availability, and network condition are utilized to prove our hypothesis and show the effectiveness of our approach. The results demonstrate the potential impact of SDN, edge/in-transit processing, and approximate computing on various cloud federation aspects such as completion ratio of the jobs, quality/accuracy of the processed data, and utilization of resources.