Efficient time compression earthquake database using hadoop Hive ORC format
Pramod Ravindra Patil, Vivek Kshirsagar · 2017
Today is an age of Big data. Big data is the normally unstructured data. Apache Hive is largely used for analysis in process of huge data. Because it is like SQL so easy to get analytical report. The main problem is that unstructured data loading and storage as well as Fast and timely analysis of large amount of data. There are data Compression columnar format like ORC(Optimized Row And Columnar) and Parquet columnar format. In this paper we used USGS (United States Geological Survey) Earthquake dataset. USGS provides the multi-Dimension dataset of earthquake of every day, week and month. We applied hadoop Hive's ORC format On monthly USGS earthquake dataset. ORC format Stored dataset efficiently without lose so that the most important data without losing stored on HDFS. We compare result of ORC Sorted and Unsorted dataset on the basses of time required to load the dataset on HDFS.