Big data and smart computing: methodology and practice
Chunming Rong, Lu Liu, Guolong Chen · Concurrency and Computation Practice and Experience · 2016
The field of Big Data is particularly challenging in practical aspects of data collection, data analysis and data storage and has massive consequential implications for the future. The scope and scale of Big Data frequently necessitate the processing power attributed to High Performance Computing (HPC) and the scalability of Cloud Computing for processing and storage. Source topics for Big Data are continually expanding and include diverse topic areas of finance, medical, particle physics and social interaction and everything in between. Effective Big Data analysis encompasses the need to minimise data sets by identifying, analysing and storing only significant data; recognising and separating the relevant and irrelevant data provide many research opportunities in itself. The outcomes of Big Data analysis may explain past events and trends, suggest controls that are necessary for the present or predict a future world in terms of preparation, planning or positioning. This special issue will further publicise and promote this immensely import field of research with goals of creating a public record of achievements so far and providing inspiration for even greater developments in the coming years. Computational costs associated with Cloud Computing can be significant. In the paper ‘Online optimization scheduling for scientific workflows with deadline constraint on hybrid clouds’ 1, Bing Lin, Wenzhong Guo and Xiuyan Lin examine scheduling strategies and propose ‘hierarchical iterative application partition’ (HIAP) as an algorithm to increase the number of workflows completed within a given timeframe, thereby reducing overall costs. Ensuring the integrity of transmitted data is of paramount importance for data analysis, the paper ‘A MapReduce based Parallel K-Means Clustering for Large Scale CIM Data Verification’ 2 discusses the topic and presents a parallel K-means clustering algorithm for large scale Common Information Model (CIM) data verification. The paper concludes that time saving is achievable using parallel K-means while generating a high level of precision in data verification. ‘Bursty Event Detection from Microblog: A Distributed and Incremental Approach’ 3 is a practical example of the use of Big Data analysis. The researchers propose a method of bursty event detection, BEE+, as a means of tracking ‘topic drift’ from a microblog dataset of over 6 million posts. The amount of data collected and necessary storage rate are frequent considered to be problems for Big Data systems. ‘Performance Evaluation of a Distributed Storage Service in Community Network Clouds’ 4 compares the write and read capability of Tahoe-LAFS storage system is when deployed on community clouds and commercial systems. The paper concludes that write speeds are comparable, while read speeds were better in the commercial system. The information content of Big Data can have a significant commercial value. The cost to analyse data can be very high while the act of data collection can be both expensive and time consuming; loss of data to a competitor or invalidation because of falsification could result in commercial collapse of a business. ‘Secure Cryptographic Functions via Virtualization-based Outsourced Computing’ 5 considers the use of cryptography to protect data and, more fundamentally, suggests a method for the protection of the cryptographic system and process. In ‘Towards an Autonomous Decentralised Orchestration System’ 6, the authors propose distributed execution engines which exploit the benefits of parallel computation in the workflow to improve overall execution time. The paper provides an evaluation of the system and demonstrates the scalability benefit of the decentralised system. The limitations of Cloud related simulation tools are the subject of ‘Multi-layered simulations at the heart of workflow enactment on Clouds’ 7. The authors suggest that a multistage approach is advantageous in resolving some of the issues of scalability and scope without adversely affecting the performance of the workflow execution simulation. The continuing expansion of the Internet has presented ever greater challenges to crawler services used to collect information for indexing. ‘A Task Scheduling Strategy based on Weighted Round-Robin for Distributed Crawler’ 8 presents an implementation of a multithread distributed crawler which is scalable and fault tolerant. The paper includes experiments which indicate that the system exhibits good load balancing performance. The ability to adapt to changes of tenant requirements and cloud services is investigated in ‘Cross-Clouds Services Autonomic Management Approach based on Self-Organising Multi-Agent Technology’ 9. The research proposes a method where cloud services are managed by a series of autonomous agents which interact with each other to obtain macro-level service aggregation. The efficiency and usability of the proposed approach are confirmed with experimental results using public data sets. ‘Bilinear-map Accumulator based Verifiable Intersection Operations on Encrypted Data in Cloud’ 10 investigates the problem of conducting set-intersection operations on the ciphertext sets in the Cloud without the capability of decryption. The research opposed a model, called VIOEDC, to address this problem. The correctness and the security properties of the model have been approved in this paper. [Correction added on 07 June 2016, after first online publication: this paragraph has been added.] The papers presented in this special issue show that the subject arena of Big Data and Smart Computing continues to offer many diverse opportunities for research and investigation. It is anticipated that this special issue papers will provide a foundation for further work in the future by the authors and by many other researchers.