Developing Scalable and Fault Tolerant Distributed Architecture for Big Unstructured Data Systems
G Saranya, Kumaran K, V Pratheep, S R Dhanush · 2025
In the era of big data, efficiently processing and retrieving insights from unstructured data presents a critical challenge. This paper introduces a scalable leader-worker distributed data pipeline designed to handle large sets of unstructured data blobs. Leveraging data parallelism, the system distributes workload across worker nodes, ensuring high performance for embarrassingly parallel tasks with minimal inter-process communication. Fault tolerance is achieved through an optimized next-in-line leader re-election strategy, addressing cascading failures and mitigating the thundering herd effect. The use of Apache Zookeeper for service discovery ensures seamless coordination within the cluster while maintaining a loosely coupled architecture for enhanced scalability. Google Protocol Buffers are employed for secure data transmission, offering tamper resistance against eavesdropping and man-in-the-middle attacks. The system also integrates the TF-IDF algorithm for efficient information retrieval, demonstrating the ability to achieve data parallelism while preserving task uniformity across nodes. Performance evaluation highlights significant reductions in response times with increasing cluster size, validating the system's scalability, efficiency, and robustness. The proposed architecture is well-suited for real-world applications requiring secure, fault-tolerant, and scalable solutions for unstructured data analysis.