Efficient and Flexible Information Retrieval Using MonetDB/X100

Sándor Héman, Marcin Żukowski, Arjen P. de Vries, Peter A. Boncz, Gerhard Weikum, Joseph M. Hellerstein, Michael R Stonebraker · Centrum Wiskunde & Informatica (CWI), the national research institute for mathematics and computer science in the Netherlands · 2007

Today's large-scale IR systems are not implemented using general-purpose database systems, as the latter tend to be significantly less efficient than custom-built IR engines. This paper demonstrates how recent developments in hardwareconscious database architecture may however satisfy IR needs. The advantage is flexibility of experimentation, as implementing a retrieval system on top of a DBMS boils down to relational query formulation, rather than system programming. We demonstrate in the context of the TeraByte TREC efficiency task that our experimental MonetDB/X100 database system provides highly competitive results both regarding precision and speed. We analyze the two innovations in MonetDB/X100 that most contributed to this successful application of DB technology in IR, namely vectorized incache processing and the use of two new light-weight compression schemes that work between the RAM and CPU cache memory levels.

Read the paper · More papers on PaperTik