A Proposal for Optimization of Data Node by Horizontal Scaling of Name Node Using Big Data Tools
Chandrima Roy, Manjusha Pandey, Siddharth Swarup Rautaray · 2018
The data which are beyond the storage space of the server and beyond to the processing power is called Big Data. It is not manageable by traditional RDBMS or conventional statistical tools. Big data Increases the storage capacities as well as the processing power. Horizontal scaling or sharding is needed to divide the data set and distributes the data over multiple servers. Redundancy and fault tolerance is achieved by horizontal scaling. Optimization of horizontal scaling is an important aspect of Big Data technology. Instead of using vertical scaling that means upgrading to fancier computers when the current system becomes inadequate, we have to add more node (computers) to a cluster. It increases the parallelism, rather than the performance of any one node. This paper presents the fundamentals of big data analytics but directing towards an analysis of various optimization techniques used in the big data environment.