Introduction to Big Data
David Lauzon · 2012
This paper presents an introduction to some of the most popular Big Data technologies. Big Data is a term coined to describe volume of data too high to be served by standard RDBMS or OLAP technologies, thus requiring alternative approaches. These approaches are usually regrouped as NoSQL technologies. This paper presents an overview of the Hadoop ecosystem, and some of its main components (HDFS, Hive, MapReduce, Impala, HBase). While HDFS provides the distributed lesystem; MapReduce provide the core framework to query HDFS data. Hive enables to query HDFS in batch using a subset of the SQL language. Impala does the same as Hive, but attemps to do it in real time. HBase is a distributed column-based storage, extremely ecient in retriving information based on a subset of a column key. We will then cover some uses cases when it is appropriate to use these technologies.