Big Data Infrastructures Using Apache Storm for Real-Time Data Processing

Chitra Sabapathy Ranganathan, Rajeshkumar Sampathrajan, P L Kishan Kumar Reddy, S Parkavi, Balasubramanian Meenakshi, S. Naga Nandini Sujatha · 2025

Apache Storm-based Big Data infrastructures strive to provide a scalable and fault-tolerant platform that will transform real-time data processing. The goal is to use Apache Storm's features to handle and analyze massive amounts of data in real-time, guaranteeing correct insights that are delivered promptly. The objective is to provide a solid foundation that can manage data moving at high speeds from many sources, allowing for analytics in real-time and continuous computing. The goal is to maximize throughput while reducing delay via architectural optimization and efficient data flow methods. The goal is to improve Apache Storm's capabilities and make sure data processes smoothly across dispersed systems by combining it with other big data technologies. The end goal is a robust system that can provide analytics on data in real-time, which will aid in decision-making across several industries, including banking, telecoms, and social media. Efficiency, scalability, and dependability in managing massive data streams are key considerations. The data flow topology dataset for 5 distinct components and 5 streams shows throughput values ranging from 2500 to 4100 tuples per second. Similarly, the fault tolerance and reliability dataset for 5 distinct nodes and 5 different intervals shows throughput values ranging from 700 to 1050 tuples per second. All these results are derived from real-time sensor data used for traffic monitoring.

Read the paper · More papers on PaperTik