Review of Approaches to Optimize Small File Handling in Hadoop Ecosystem
2025
The Hadoop ecosystem, widely used for big data storage and processing, encounters significant performance and scalability challenges due to the proliferation of small files.These small files, often generated by log data, sensors, or data ingestion processes, lead to inefficient storage utilization, excessive memory consumption in the Name Node, and increased latency in processing.This review provides a comprehensive analysis of recent advancements and methodologies proposed to address small file problems in Hadoop.It examines techniques such as file merging strategies, metadata management optimizations, archive-based storage approaches, machine learning-based prediction models, and hybrid frameworks integrating HDFS with NoSQL or cloudnative systems.The paper categorizes these methods based on their implementation layers-HDFS-level, MapReduce-level, and application-level-and evaluates their effectiveness in terms of performance, scalability, fault tolerance, and implementation complexity.Special attention is given to recent solutions that leverage dynamic repositories, hash-based archives, and intelligent compaction strategies for real-time analytics.Furthermore, this review highlights research contributions aimed at enhancing data deduplication, access speed, and metadata compression.The comparative analysis reveals that no single solution universally addresses all performance bottlenecks, thus motivating future work towards adaptive, workload-aware frameworks.Finally, the paper identifies open research gaps such as integration with modern data lake architectures, security-aware file optimization, and the impact of AI-driven automation in file management.This review aims to guide researchers and practitioners in choosing appropriate solutions for optimizing small file handling in Hadoop-based environments.