Faster The Slow Running RDBMS Queries With Spark Framework

Hariteja Bodepudi · International Journal of Scientific and Research Publications · 2020

Data is increasing day by day due to increase of the advanced internet technology, online browsing, Internet banking and online shopping.Modern Technology has helped the humankind to access and communicate anywhere in world very easily.This advanced technology drives the increase of data day by day.Most of the companies traditionally maintain the transactional data in the Relational Data Base systems like Oracle, SQL Server, MYSQL etc.Every Organisation will have the daily reporting, weekly reporting, and monthly reporting on the collected data.Due to increase in the volumes of data from Terabytes to Petabytes, the processing of the complex queries for reporting in the RDBMS was really time taking and slow.The usage of Spark to process these reporting complex queries will be 10 times faster than the actual processing of the data in the RDBMS.In General , the RDBMS uses the single node of the cluster to process the query and it is really time taking to process the complex queries used for daily ,weekly and monthly reporting as it got complex joins and multiple aggregate functions involved in it.But spark leverage all the nodes in the clusters to process the data very quickly by using the partitions in the table to break into small chunks and achieve the level of Parallelism to process 10 times faster than the RDBMS.This Paper talks about how the RDBMS reporting scripts performance can be improved by incorporating the Spark framework without changing the existing queries in the RDBMS.

Read the paper · More papers on PaperTik