Improvised Distributions framework of Hadoop: A review
Baydaa Hassan Husain, Subhi R. M. Zeebaree · International Journal of Science and Business · 2021
HADOOP is an open-source virtualization technology that allows the distributed processing of large data sets across standardized server clusters. With two modules, HADOOP Distributed File System (HDFS) and MapReduce framework, it is designed to scale single servers to thousands of computers, providing local computation and storage. Over a decade after HADOOP emerged on the forefront as an open system for Big Data analysis. Its growth has prompted several improvisations for particular data processing needs, based on the type of processing conditions at various periods of computation. This paper, through reviewing several kinds of research provides the basic HADOOP system structure and the description of the MapReduce, HDFS Efficiency. Explaining how the HADOOP framework can overcome the “5Vs” challenges in Big Data. However, in addition to the many benefits of the HADOOP system, like fault tolerance, reliability, high availability, scalable, decreases execution time, reduces latency, improve the security issues, improving the quality of data analysis, better scheduling model, and cost-efficiently. On the other hand, there were some barriers and challenges regarding adjusting data regularly, security issues, and load balancing. Finally, the certainly benefit and challenges of the HADOOP system have been represented paving the way for the future research to find solutions to these challenges.