Catalyzing big data excellence: the enhanced Data Placement Policy (EDPP) revolution
Jeffin Gracewell J, Ponnada Naga Ramya, S. Venkatesh Babu · Egyptian Informatics Journal · 2025
The Enhanced Data Placement Policy (EDPP) offers a novel approach to optimising data storage in Hadoop-based Big Data systems that utilise the Hadoop Distributed File System (HDFS). Efficient data placement is essential for maximising data retrieval speed and overall system performance. EDPP leverages the combined power of MapReduce (MR) and Particle Swarm Optimisation (PSO) to address this challenge effectively. In these systems, files are divided into chunks, which can be either duplicates (pointing to existing data copies) or unique (requiring strategic placement). EDPP uses MR to identify the most suitable Data Nodes (DNs) for storing unique chunks, enhancing data retrieval efficiency. Furthermore, EDPP addresses the issue of maintaining balanced data distribution in heterogeneous clusters. By utilising PSO, it intelligently selects DNs based on their storage capacity and response times, ensuring optimal resource utilisation and load balancing. This research presents a practical and efficient solution to the complex problem of data placement in Big Data systems, offering significant benefits to organisations with diverse hardware and software configurations in their Hadoop clusters. Finally, experimental results demonstrate that the Enhanced Data Placement Policy (EDPP) significantly improves performance in Hadoop-based Big Data systems. The EDPP reduces data retrieval time by an average of 33 % compared to traditional methods and achieves a throughput of 35.03 MB/s with a 98 GB dataset, more than doubling the performance of the existing methods. Additionally, it ensures even data distribution, preventing node overload and optimising resource utilization.