Applying MapReduce algorithm to performance testing in lexical analysis on HDFS
Ognjen V. Joldzic · 2013
This paper presents an overview of distributed data processing technology, and explores the possibilities and advantages of using this technology in lexical analysis of Cyrillic text. A detailed overview of one of the most widely used framworks for processing large datasets - Apache Hadoop - is presented, along with a recommendation for planning and deployment of such systems. The paper also analyzes results obtained by running lexical analysis programs on a small Hadoop cluster and the effect of various configuration parameters on total execution times of the test programs.