A system of job log analyzing for Hadoop
Xiaogang Zhao, Zhiqiang Ma, Liu Xu · 2012
Handling the huge amount of history logs produced by Hadoop distributed computing platform is really a troublesome task and always these history files looks useless. But if we want find out the health degree of cluster platform we must analyze the huge history logs produced by the running jobs. It seems that a single-machine analyzing program cannot satisfy you because of its low speed, high demand of memory and CPU. In this thesis we tried to solve this problem in a distributed way with the Map/Reduce calculation model. We also built a data platform(hive and MySQL) to store these data. From the experiment we can see the distributed way to process log files get good performance when data log files are huge.