Big data cloud and the frontier of computer science and technology
Keqiu Li, Hongyi Wu, Zhiyang Li · Concurrency and Computation Practice and Experience · 2015
This special issue presents the recent advances in cloud computing and software-defined network, which \were selected out of the significantly extended versions of accepted papers in the Fifth IEEE International Conference on Big Data and Cloud Computing (BDCloud 2015) 1, the Ninth International Conference on Frontier of Computer Science and Technology 2, and a large number of open submissions. The selection has been very rigorous, and only the best papers were selected. Wei et al. 3 observe that most existing message forwarding algorithms in delay-tolerant networks prefer to deliver messages to the nodes with a higher popularity or centrality. This forwarding scheme can achieve high delivery ratio and low end-to-end delay but is prone to cause unfair load distribution and further lead to network congestion. To tackle this, they first track the evolution of communities through a novel distributed community detection approach. The second one is to develop a congestion avoidance mechanism to divert load away from congested areas and further present a congestion-aware message forwarding algorithm where messages can avoid being transmitted to the congested nodes Since the rapid growth of large-scale online services, massive amounts of the generated traffic have been seen in the data center network. Li et al. 4 study the emerging congestion problem in the software-defined data center network. Note that existing approaches are either hard to be implemented in hardware or unable to obtain the optimal solutions. The authors first propose a heuristic algorithm for efficiently compute a timeslot allocation for the coming packets. Further, they model the path selection as a bin-packing problem. By seamlessly combining the timeslot allocation and path selection, each data packet will not suffer queuing and waiting in the data center network. Anomaly detection is an effective approach to enhance availability and reliability of cloud infrastructures. Hong et al. 5 study the anomaly detection problem in cloud computing systems without the need for prior knowledge about normal or anomalous behaviors. They propose an unsupervised online anomaly detection scheme based on hidden Markov model. In order to achieve high scalability, their proposed algorithm runs in a distribution manner among multiple computing machines in the cloud. They also perform extensive experiments based on real data sets to validate the high detection accuracy for their proposed algorithm. Because of the benefits of reducing the communication overhead, distributed data-centric storage in wireless sensor networks have received considerable attention. Xu et al. 6 focus on the big data storage problem in wireless sensor network with the nonuniform node distribution. Note that most existing distribution methods can significantly consume more energy and are unable to deal with the case of nonuniform sensor nodes distribution. To address this issue, they propose an efficient storage retrieval algorithm to estimate the real distribution of the sensor nodes and the real addresses of these nodes. Based on this algorithm, they further take the data redundancy among sensor nodes into account and exploit an efficient routing mechanism. Incorporating cloud computing into vehicular networks is a promising solution to the collection, storage, and analysis of big traffic-related data but can lead to new challenges to the allocation and management for cloud resources in road-side cloudlet. Yao et al. 7 study a VM migration problem with the goal of minimizing the total network cost, by making the decisions on which VM should be migrated and where the VM shall be migrated. They further formulate an optimization for the static off-line VM placement problem and then propose a heuristic algorithm with polynomial time to solve the optimization. NoSQL systems, replicating and partitioning data over many servers for improving the performance, are widely used for storing big data. Conventional radon virtual nodes and manual configuration methods for consistent hashing can significantly lead to imbalanced data partition. Huang et al. 8 study the performance degradation problem caused by the imbalanced data partition. They first propose a novel imbalance coefficient of data distribution. They further propose a dynamic programming algorithm to compute the position of the new coming node in the consistent ring. Finally, they conduct comprehensive simulations based on a benchmark Yahoo Cloud Serving Benchmark (YCSB) to show the benefit of their proposed algorithm. Data centers are increasingly deploying the NUMA architecture. Zhu et al. 9 focus on the performance degradation problem when running multi-threaded programs on such NUMA systems. Note that the existing works mainly use the single-threaded multi-programming workloads to study the performance of NUMA on the resource contention and data locality. To solve the performance lagging problem, they propose a novel scheduler—symmetric scheduler, which can balance the number of costly remote shared data accesses for threads on NUMA systems. Finally, they perform extensive simulations on the PARSEC benchmark, and their proposed schedulers can significantly outperform Linux kernel scheduling mechanism. As the number of functionally equivalent services in the cloud grows, collaborative service QoS prediction has recently garnered increasing attention. Tang et al. 10 propose a collaborative QoS prediction method with location-based data smoothing, for addressing the data sparsity issue and improving the QoS prediction accuracy. Note that existing solutions, simply exploring the historical QoS information generated by interactions between users and services, however, can significantly suffer from the data sparsity issue. To address this issue, the authors firstly compute neighborhoods of users and services based on their locations, which provide a basis for data smoothing. They further combine user-based and service-based collaborative filtering techniques to make QoS predictions. Finally, they conduct comprehensive experiments on real service invocation dataset to validate the performance of their proposed QoS prediction method. We hope that you will enjoy reading these papers in this special issue. We would like to thank the authors for contributing their papers to this issue and thank all the reviewers for their time and constructive reviews. Finally, we would like to thank the editors of Concurrency and Computation: Practice and Experience for providing this opportunity to publish this special issue.