Scaling to Large³ Data: An Efficient and Effective Method to Compute Distributional Thesauri

Martin Johannes Riedl, Chris Biemann · 2013

We introduce a new highly scalable approach for computing Distributional Thesauri (DTs).By employing pruning techniques and a distributed framework, we make the computation for very large corpora feasible on comparably small computational resources.We demonstrate this by releasing a DT for the whole vocabulary of Google Books syntactic n-grams.Evaluating against lexical resources using two measures, we show that our approach produces higher quality DTs than previous approaches, and is thus preferable in terms of speed and quality for large corpora.

Read the paper · More papers on PaperTik