An introduction to information retrieval and the use of map-reduce for text processing

Daniel Taipala · Journal of computing sciences in colleges · 2015

According to projections by the IDC (International Data Corporation), in their Digital Universe study released in 2012, the amount of data on the planet that is stored and accessible in some fashion will grow to exceed 44 zettabytes by the year 2020 [1]. To give some sense of the scale of a zettabyte, it is one billion terabytes. Most of this data is unstructured meaning that it is not stored or managed in traditional data management or database solutions. Much of this unstructured data is textual or is described with textual meta data meaning that the need for technologies that are able to search for and find useful and meaningful data within this massive universe of data is increasingly important and the demand for IT professionals who understand and can develop solutions capable of searching and analyzing such data is high. This tutorial will introduce basic concepts of text search and analytics. Participants will use map-reduce processes to develop text search and analytics solutions. Applications for search and other text analytics will also be discussed.

Read the paper · More papers on PaperTik