Performance Enhancement in Big Data handling

Himadri Sekhar Ray, Swastik Mukherjee, Nandini Mukherjee · 2020

In the contemporary world, with the emergence of information technologies, Big data is one of the rising IT trends. An overwhelming amount of data and information is generated every day throughout the world. Storing and processing this huge volume of data is named by a ubiquitous term: Big Data Management. The traditional database architectures are not designed to face the challenges associated with huge data. The Apache Hadoop is the well-known software library framework that allows distributed processing of large data sets across clusters of low-end computers using simple programming models. Hadoop is designed to scale up from single servers to thousands of machines, each offering local computation and storage. When HDFS takes in data, it breaks the raw data into separate blocks and distributes them to different nodes(DataNode) in a cluster for parallel processing. But there is no scope to determine, which block is stored in which DataNode. Thus if there is a need to run a query on a specific item/keyword/key, the Hadoop runs the query or analytics program to all its cluster, which somehow increase the overall cost. To overcome these issues, We propose a scheme where we store the data in such a way that the system responses faster to each query and overall processing time decreases. Our motivation is to deliver an efficient and reliable storage and retrieval system in the health-care industry.

Read the paper · More papers on PaperTik