Cloud computing–based big data processing and intelligent analytics
Fang Dong, Chenshu Wu, Shangce Gao · Concurrency and Computation Practice and Experience · 2019
Cloud computing-based big data processing and intelligent analyticsCloud and big data have become the big things today in many systems especially regarding of information processing and intelligent analytics.Big data analytics is the use of advanced analytic techniques against very large, diverse data sets, and it allows analysts, researchers to make better decisions using data that was previously inaccessible or unusable.Due to the urgent demand on high capacity of computation and storage resources, cloud computing has been acknowledged as the primary computing paradigm for massive data storage, processing under various circumstances and different requirement.1 Moreover, edge computing pushes the cloud frontier to the edge of the network and extends cloud computing to be able to address more application scenarios.This special track plans to solicit novel and original manuscripts in the above topics with an emphasize on ''Cloud Computing-based Big Data Processing and Intelligent Analytics.''From those submitted papers for the 6th International Conference on Advanced Cloud and Big Data (CBD 2018) held in Lanzhou, China on August 12 to August 14, 2018, nine papers are selected that target the following research issues in cloud computing and big data:• Service deployment and task scheduling in cloud computing and edge computing.• Cloud storage system design and optimization.• Case studies of big data in cloud-based system.• Network performance optimization in data centers.• Approximate big data analysis and processing.• Security threats and solutions in cloud computing and big data processing.Recently, mobile edge computing (MEC) has become a fascinating technology trend for future computing paradigm, while at the same time, it is also facing a great challenge that is how to make full use of edge resources to provide a seamless support for compute-intensive latency-sensitive applications.Most of the existing works assume that tasks can be executed upon every edge server, but the assumption does not hold in practical scenarios because a specific application task often corresponds to a certain service that provides the corresponding running environment.How to decide service deployment of so many types of services among multiple edge servers is also a big challenge.To address the challenge, Zhou et al 2 study dynamic service deployment for latency-sensitive applications and first model the long-term budget-constrained latency minimization problem as a multi-slot latency minimization problem based on the Lyapunov framework.Furthermore, the task scheduling optimization is also considered here, which makes every edge server be fully utilized in an even more efficient collaborative manner.How to repair data blocks in the erasure coding storage system is of great challenge as considering the incast problem introduced at the new node.Current solutions mainly rely on path planning and resource allocation and waste a large amount of storage and bandwidth resources unavoidably.Xia et al 3 propose the incast problem to be resolved economically via the in-network aggregation, where a set of in-network methods to repair a failed data block in the erasure coding storage systems.Compared with the existing methods, the in-network methods are capable of reducing the bandwidth consumption as well as achieving higher repair speed.With the continuous development of intelligent transportation systems (ITSs), various sensor data can be used to detect traffic information, but there is a lack of practical value to guide public traveling.Xu et al 4 propose a real-time traffic index model of expressways by using a traffic index to evaluate the actual conditions of expressways.The model considers the actual situation of floating and non-floating vehicles on expressways.Included is the realization of the complete calculation model of real-time traffic index estimation, including highway section division, spatial topology map matching, driving route calculation, and road congestion status judgment.For roads without floating car coverage, the weighted-moving-average time-series prediction method is used to predict the traffic index, so that the running condition of all roads in the network can be analyzed completely.A spotlight has shined on scientific workflows in recent years, as a result of their enormous impact on big data-related scientific areas.Large-scale scientific workflow scheduling across global data centers requires the scheduling framework to optimize data movement cost by leveraging these distributed data centers.However, challenges regarding of data-intensive workflow execution in multiple geo-distributed data centers still exist, such as data dependency, intermediate data placement, etc. Scientific workflow's data and task co-scheduling aim to solve the above problem, which is known to be NP-hard.Zhang et al 5 propose a novel approach based on the multilevel graph coarsening and un-coarsening framework, together with a specialized hybrid genetic algorithm having distinctive graph partition driven features of repair and local improvement, for scheduling data-intensive scientific workflows in geo-distributed data centers and optimizing the cross-data center data transfer volume.