Applying Hadoop's MapReduce framework on clustering the GPS signals through cloud computing
Wichian Premchaiswadi, Walisa Romsaiyud, Sarayut Intarasema, Nucharee Premchaiswadi · 2013
Year by year, we are considerably witnessing a dramatic increase in the size of data gathered from machines or human interactions. Typically, the data generated by machines is massive, complex and comes from different varieties including sensors collecting climate information, posts being shared in social media sites, videos being posted online, digital pictures, transaction records of online purchases, cell phone GPS signals and so on. Not surprisingly, the amount of data generated by machines is greater than the data generated by human elements. Sensor data (obtained from transportation, logistics, retail, utilities, and telecommunications) has continuously been generated from fleet GPS trans-receivers, RFID tag readers; smart meters, to cell phones. Such data has frequently been used in numerous parallel processing methods so as to optimize operations and drive operational business intelligence (BI) systems scrutinizing immediate business opportunities. Appropriately, MapReduce is a programming model designed for expressing distributed computations on massive datasets and an execution framework for large-scale data processing on clusters of commodity servers. In this paper, we enhanced the Hadoop MapReduce for data-intensive computing on massive datasets of GPS signals. We developed an execution framework for large-scale data processing through the cloud system - in order to reduce the execution time of the cluster systems - as well.