QoS oriented mapReduce optimization for hadoop based bigData application
Burhan Ul, Islam Ullah Khan, Rashidah Funke Olanrewaju · 2014
Due to increase in data load on cloud infrastructure the maintenance of quality services for BigData applications has become a major issue [1]. The data storage of structured as well as unstructured kinds are also a big concern to be considered while operating with certain search engine or data exploring applications[2]. On the other hand data intensive computation for cloud services is also increasing with vast pace. The predominant requirement for quality services in cloud infrastructure are optimized resource utilization, efficient processing in terms of execution time and data portability [1]. To develop such a cloud framework a number of researches have been done and still going on. Hadoop framework is one of the most successful cloud frameworks for cloud applications. This is the matter of fact that Hadoop is a potential candidate to operate with hundreds of Petabytes of data on cloud, but with increase in further loads, it still needs certain optimization in terms of execution speed and fair load balancing or job scheduling across encompassing nodes in cloud [3]. Hadoop encompasses two components, Hadoop distribution file system (HDFS) and MapReduce. HDFS represents an interface of nodes, task and user assigned jobs. But in fact a revolutionary system optimization can be carried in terms of MapReduce only [4]. HDFS possesses little scope of enhancements. Dealing with unstructured data and data intensive applications, data locality and its mining becomes a very difficult scenario for cloud infrastructure [5]. Therefore, generic load balancing and scheduling cannot be advocated for such huge data processing applications. Even individual enhancements of MapReduce which itself is a combination of Mapping, Sampling and Reducing, might cause higher and of course unwanted overheads on cloud, resulting into QoS degradation. MapReduce can provide certain scope for further enhancement [6]. In Hadoop based cloud compute system, the MapReduce framework functions for collecting data from various clusters or nodes and then processing for Map process which is followed by Reduce phase. Thus in between these two phases along with the resource retrieval and process of resource allocation at proper cluster nodes might be a huge fraction that can be further enhanced to yield better results. The predominant issue with MapReduce is that this framework is fundamentally batch-processing oriented and whenever processing is initialized, it’s updation for input data cannot be done with the expectation of similar output [6]. This is the main reason that makes this framework function poor in case of its real time application. Similarly, the data collection, process for Mapping and the shuffling which is in later stage converted to the intermediate data is a time consuming process, which is required to be optimized. Again, the reduce process of the intermediate data with the key based approach is also a tedious task compared to the conventional table driven or SQL based approach, as the data categorization on the basis of certain rank is a mammoth task. The replacement or allocation of these data efficient nodes is a great deal to be taken care of. So, considering these all aspects of functional MapReduce technique, it can be found that there is a huge space for algorithmic optimization for MapReduce. Therefore a scheme must be developed that could effectively eliminate the issue of resource utilization and latency in cloud infrastructure. The optimization with intermediate storage in MapReduce can give improved results and the selection of right data location for specific data type can reduce the computational complexity and execution time can be enhanced. Thus scheduling of MapReduce components with data location and availability of sensitive mode can be much fruitful. It can further be optimized if along with MapReduce job assignments a fair load balancing across comprising nodes is formed. On the other hand, because of the dynamic characteristics of cloud and its heterogeneous behaviour existing between the central servers and storing disks there must be something like a parallel architecture that could enhance the processing speed and data retrieval rate in MapReduce. Similarly, the load must be distributed uniformly across the network, so that the situation of uneven data distribution can be eliminated [7].