Scaling up through Parallel and Distributed Computing
Huy T. Vo, Claudio T. Silva · 2020
This chapter provides an overview of techniques that allow us to analyze large amounts of data using distributed computing. It also provides a conceptual and practical framework to manage large amounts of data that may not fit in memory or may require too much time to analyze on a single computer. The chapter focuses on one such framework, called MapReduce, to perform large-scale data analysis distributed across multiple computers. It describes the MapReduce framework, work through an example using it, and highlights one implementation of the framework, called Hadoop, in detail. Hadoop requires a distributed cluster of machines to operate efficiently. Hadoop is written entirely in Java, thus it is best at supporting applications written in Java. However, Hadoop also provides a streaming application programming interface that allows arbitrary code to be run inside the Hadoop MapReduce framework through the use of UNIX pipes.