Is There a Cloud in your Future? Applications of "Cloud Computing" to Web-scale Problems
Jimmy Lin · Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign) · 2008
IBM and Google recently committed a total of $30 million over two years to an initiative on “cloud computing”, in collaboration with six universities across the country (see references). They are: Berkeley, Carnegie Mellon, MIT, Stanford, the University of Maryland, and the University of Washington. I am the leader of this initiative at the University of Maryland, and to my knowledge the only participant from an iSchool (the rest are lead by faculty in computer science departments). “Cloud computing ” refers to technology for exploiting large computer clusters to tackle “Web-scale ” information processing problems, where immense quantities of data make traditional sequential processing impractical. Specifically, this initiative focuses on Google’s MapReduce programming paradigm, which was specifically designed for processing extremely large data sets (and indeed used by Google itself for much of its production operations). Programs written in the MapReduce functional style are automatically parallelized and executed on a large cluster of commodity machines. The run-time system takes care of the details of partitioning the input data, scheduling the program’s execution across a set of machines, handling machine failures, and managing the required inter-machine communication. Hadoop is an open-source implementation of the MapReduce framework. As a part of this initiative, IBM and Google are making Hadoop clusters available to the university collaborators, with the simultaneous goal of advancing research and education. For the past two months, the Computational