Comparison Study between MapReduce (MR) and Parallel Data Management Systems (DBMs) in Large Scale Data Analysis
Miriam Lawrence Mchome · HIMALAYA · 2011
As the quantity of structured and unstructured data increases, data processingexperts have turned to systems that analyze data using many computers in parallel.This study looks at two systems designed for these needs: MapReduce and paralleldatabases. In the MapReduce programming model, users express their problem interms of a map function and a reduce function. Parallel databases organize data as asystem of tables representing entities and relationships between them. Previouscomparison studies have focused on performance, concluding that these twosystems are complimentary. Parallel databases scored high on performance andMapReduce scored high on flexibility in handling unstructured data. Both systemsoffer a querying language: Pig Latin for MapReduce systems and SQL for paralleldatabases. This study compares the operations, query structure and support foruser defined functions in these languages. The findings offer data processingexperts insights into how data organization and querying structure affects dataanalysis.