Approximating matrix multiplication for pattern recognition tasks

Edith Cohen, David D. Lewis · 1997

Many pattern recognition tasks, including estimation, classification, and the finding of similar objects, make use of linear models. The fundamental operation in such tasks is the computation of the dot product between a query vector and a large database of instance vectors. Often we are interested primarily in those instance vectors which have high dot products with the query. We present a random sampling based algorithm that enables us to identify, for any given query vector, those instance vectors which have large dot products, while avoiding explicit computation of all dot products. We provide experimental results that demonstrate considerable speedups for text retrieval tasks. 1 Introduction In pattern recognition tasks, a database of instances to be processed (images, signals, documents,...) is commonly represented as a set of a vectors x 1 ; : : : ; xn of numeric feature values. Examples of feature values include the number of times a word occurs in a document, the coordinates...

Read the paper · More papers on PaperTik