Performance Analysis via Hadoop: A Study of between text and.orc file extensions
Luana Thamiris Da Silva de Oliveira, Maristela Terto de Holanda, Marcio de Carvalho Victorino · 2023
In the age of Big Data, access to open data contributes to increasingly in-depth studies on different themes, whether social, economic, or political. A practical example is the analysis on the distribution of the Emergency Aid in Brazil, approved by the National Congress in 2020, due to the pandemic of the new coronavirus. However, as will be seen throughout the text, there are challenges to processing massive databases. Tools capable of processing a large volume of data in a short period of time are increasingly needed. Apache Hadoop is an important framework that allows distributed processing of massive databases, reducing query times. To elucidate this issue, this paper presents a comparison between two file extensions (.txt and.ORC) and checks their performance in processing Hive queries on the Emergency Relief database. The results showed significant differences in average execution time when adopting the.ORC file format. On the other hand, what was thought to be discrepant, the average value of the aid, was not proven, since all States received similar values.