Advanced Analytics-Technology and Tools
EMC Education Services · 2015
This chapter addresses several aspects of collecting, storing, and processing unstructured and structured data. It presents some key technologies and tools related to the Apache Hadoop software library, a framework that allows for the distributed processing of large datasets across clusters of computers using simple programming models. The chapter focuses on how Hadoop stores data in a distributed system and how Hadoop implements a simple programming paradigm known as MapReduce. The power of MapReduce is realized with the use of the Hadoop Distributed File System (HDFS) to store data in a distributed system. HBase is one example of the Not only Structured Query Language (NoSQL) data stores that are developed to address specific Big Data use cases. Along with other Hadoop-related tools, Pig, Hive, Mahout, and HBase are briefly covered in this chapter. To illustrate the power of Hadoop in handling unstructured data, the chapter provides several Hadoop use cases.