Advanced Analytics-Technology and Tools

EMC Education Services · 2015

This chapter addresses several aspects of collecting, storing, and processing unstructured and structured data. It presents some key technologies and tools related to the Apache Hadoop software library, a framework that allows for the distributed processing of large datasets across clusters of computers using simple programming models. The chapter focuses on how Hadoop stores data in a distributed system and how Hadoop implements a simple programming paradigm known as MapReduce. The power of MapReduce is realized with the use of the Hadoop Distributed File System (HDFS) to store data in a distributed system. HBase is one example of the Not only Structured Query Language (NoSQL) data stores that are developed to address specific Big Data use cases. Along with other Hadoop-related tools, Pig, Hive, Mahout, and HBase are briefly covered in this chapter. To illustrate the power of Hadoop in handling unstructured data, the chapter provides several Hadoop use cases.

Read the paper · More papers on PaperTik