Massive Data Processing and Management in Cloud Computing: A Survey

Kyong-Ha Lee, 최현식, Chungyondohn · 2011

Cloud computing, an old idea whose time has finally come [1], has been gaining much interest from both IT industry and academia since it promises a cheaper and better way of computing by leasing their computing from the data centers of IT companies which are big enough to realize the economy of scale. In cloud computing, data processing is parallelized across tens of thousands of node computers which compose certain data centers. Although traditional distributed and parallel computing models such as parallel DBMS have already been discussed for many years, there are distinct differences between cloud computing and the conventional parallel processing. In addition, managing data in the cloud immensely differs from the conventional DBMS techniques in terms of the aggressive use of data replication for high performance and availability, less strong consistency model, simpler access patterns and the support of less-rigid data model. This paper gives readers a brief survey of techniques for processing and for managing data in cloud computing. Including the current techniques adopted in cloud computing, authors give guidelines about how to select relevant techniques with given work type and workload. Finally, we suggest research issues and opportunities in cloud data processing and management.

Read the paper · More papers on PaperTik