Towards terascale knowledge acquisition

Patrick Pantel, Deepak Ravichandran, Eduard H. Hovy · 2004

Although vast amounts of textual data are freely available, many NLP algorithms exploit only a minute percentage of it. In this paper, we study the challenges of working at the terascale. We present an algorithm, designed for the teraxale, for mining is-a relations that achieves similar performance to a state-of-the-art linguistically-rich method. We focus on the accuracy of these two systems as a function of processing time and corpus size.

Read the paper · More papers on PaperTik