Design and Realization of Hadoop Distributed Platform
Jing Wang, Haibo Luo, Yongtang Zhang · 2016
Hadoop provides a reliable shared storage and analysis system to meet the needs of the mass data of Internet users.We have designed the architecture of Hadoop distributed development platform.And under the environment of Ubuntu, the Hadoop distributed development platform is realized.The aim is to provide a parallel cloud computing system to a user in a template based and reusable way.Through the experimental test, the platform has the reliability and practicability, which provides the basis for the development of large data or cloud computing application services. Hadoop IntroductionThe rapid development of Internet has led to huge amounts of data distributed storage and computing needs.In 2004, Google published two papers "The Google File System" and "MapReduce: Simplied Data Processing on Large Clusters", which is to extend and improve its own search System.In 2005, Doug Cutting [1] took examples by technologies of Google two papers, and build a layer on Nutch to control the distributed processing, redundancy, automatic failure recovery and load balancing problem, that is the Hadoop [2].At present, Hadoop can be regarded as the defacto standard in the field of big data in industry, while it is mainly represented by Yahoo, Facebook, Adobe, EBay, IBM, Last.Fm, and LinkedIn at aboard, and took Baidu, Tencent, Alibaba, Huawei, China Mobile, and Pangu search as primary at home.Hadoop is an open-sourcing efficient platform for cloud computing infrastructure, it is not only widely used in the field of cloud computing, and it can also support the search engine services, as the underlying infrastructure system of search engine.At the same time it is more and more favored in huge amounts of data processing, data mining, machine learning, scientific computing, and other fields.In essence, the rapid development of the Internet has led to huge amounts of data distributed storage and computing needs, and Hadoop just provides a very good solution for these needs [2]. Hadoop distributed platform architecture Hadoop work pattern and overall architectureHadoop uses PC cluster for storage and computing, and it uses the Master/Slave architecture, running NameNode, JobTracker on the Master, deploying DataNode and TaskTracker on the Slave, using DataNode process in NameNode process guidance cluster to manage data storage of its own nodes.JobTracker processes command and running management and TaskTracker processes on the cluster to realize the operation control and scheduling, therefore it objectively form the working mode of command and management on Master, and implementation of computing and storage on Slave.Hadoop's overall architecture is shown in Fig. 1.Computer cluster generally run on Linux.Hadoop is based on Java, relying on Java virtual machine.The programming framework of HDFS and MapReduce are two cores of Hadoop.HDFS is responsible for data storage on cluster, which is the data reading and writing foundation of NameNode and DataNode.MapReduce programming framework is a programming model, which provides runnable and manageable programming for JobTracker and TaskTracker.SecondaryNameNode is a snapshot of NameNode, which is generally not at the same computer with Master.Hadoop can view running state, operation progress and cluster information through browser.