Applications of n‐grams in textual information systems

Alexander M. Robertson, Peter Willett · Journal of Documentation · 1998

This paper provides an introduction to the use of n‐grams in textual information systems, where an n‐gram is a string of n, usually adjacent, characters extracted from a section of continuous text. Applications that can be implemented efficiently and effectively using sets of n‐grams include spelling error detection and correction, query expansion, information retrieval with serial, inverted and signature files, dictionary look‐up, text compression, and language identification.

Read the paper · More papers on PaperTik