Comparative Analysis of SQL and Apache Spark: Time Evaluation in Loan Portfolio Analysis

Hanif Han, Teddy Mantoro, Handri Santoso · 2024

This paper provides a detailed comparative analysis of SQL and Apache Spark within the context of Loan Portfolio Analysis (LPA), focusing on the data processing efficiency of each approach across different dataset sizes. With the financial industry increasingly dependent on fast, large-scale data processing, understanding the strengths and limitations of these tools is essential. While SQL remains a longstanding solution for data management and querying, Apache Spark's distributed computing capabilities offer substantial advantages, particularly for managing large and complex datasets. Through a series of practical case studies, this study examines the time required for data aggregation and analysis using both SQL and Spark. The findings highlight notable performance differences, with Spark consistently showing increased efficiency when handling larger datasets and complex queries. Such insights are crucial for financial institutions aiming to enhance both operational efficiency and the accuracy of decision-making. This research provides a clear framework for tool selection based on dataset size and complexity, thus supporting institutions in making better-informed decisions regarding data processing practices. Ultimately, this comparative analysis serves as a guide to optimize data handling in the financial sector, aligning processing tools with data needs for improved results and strategic agility.

Read the paper · More papers on PaperTik