Big Data Processing System for Job and Task Scheduling associated with Multi Job without Prior Information

Gaurav Dimri, Rajinder Tiwari, Harin Kumar Mallela, Virendra Pal Singh, Madan Mohan Sati, Jitendra Kumar · 2024

Big data is really popular right now since it has shown to be very successful in a lot of different areas, like social media and online shopping. Big data refers to the set of tools and technology required to quickly and efficiently collect, organise, store, share, and Examine datasets ranging in size from gigabytes to petabytes. Big data sets can be unstructured, semi-structured, or structured. In big data processing systems, job and task scheduling is critical to enhancing overall system performance. In this work, we develop and implement a workable and effective task scheduler for huge data processing structures, enabling improved performance even in the event that the job's size is unknown in advance. The proposed job scheduler's exceptional performance stems from its multiple level priority queue design, which assigns jobs to lower priority queues when their total service consumption surpasses a certain threshold. This study integrated new task scheduler into YARN, a well-liked resource management used by Hadoop/Spark, to showcase its functionality. This study also verified its performance with large-scale trace-driven simulations and experiments on real datasets. The findings of this research's experiments and simulations have amply validated the design's efficacy: the new job scheduler that has been suggested can cut the Fair scheduler's average task response time by up to 50%.

Read the paper · More papers on PaperTik