Entity Search Track submission by Yahoo! Research Barcelona

Roi Blanco, Peter Mika, Hugo Zaragoza · 2010

The data has been indexed using the distributed indexing method describedin [3], i.e. implementing distributed indexing for MG4J [2] using Hadoop. Werefer the reader to our paper for the details on the indexing process, the variousalternative index structures we have implemented, the cost of creating theseindices and the size of the resulting indices.As a rst step of processing, we grouped triples about the same subject intovirtual documents using Hadoop. We indexed the data primarily using the ver-tical alternative, but for the Entity Search Track we also added the URI of theobject as an index eld as shown in Figure 2. We have preselected 300 datatypeproperties to be indexed based on the frequency of the properties. (We were re-quired to cap the number of indexed elds due to memory requirements duringindex building.) Object-properties and their values have not been indexed. Wehave segmented literal values into tokens using MG4J’s FastBu eredReader. Wehave blacklisted popular terms and also ignored terms longer than four char-acters that contained only numbers. As the data set contains text in Asianlanguages that would not be queried for and would have required special tok-enization, we also ignored terms with non-ASCII characters. Lastly, we ignoreddocuments longer than 10,000 triples.In addition to the vertical index, we have also used the token eld of ahorizontal index (see Figure 1). Note that the two indices do not necessarilyhave the same contents, because the horizontal index contains the values for alldatatype-properties. We used the horizontal index solely for acquiring globalterm frequency information (see section 3) and document sizes.We have also built a document collection for our own testing using MG4J’sSimpleCompressedDocumentCollection. The building of this collection took con-1

Read the paper · More papers on PaperTik