Ngram Search Engine with Patterns Combining Token, POS, Chunk and NE Information

Satoshi Sekine, Kapil Dalwani · 2010

We developed a search tool for ngrams extracted from a very large corpus (the current system uses the entire Wikipedia, which has 1.7 billion tokens).The tool supports queries with an arbitrary number of wildcards and/or specification by a combination of token, POS, chunk (such as NP, VP, PP) and Named Entity (NE).It outputs the matched ngrams with their frequencies as well as all the contexts (i.e.sentences, KWIC lists and document ID information) where the matched ngrams occur in the corpus.It takes a fraction of a second for a search on a single CPU Linux-PC (1GB memory and 500GB disk) environment.

Read the paper · More papers on PaperTik