Retrieval of optimal subspace clusters set for an effective similarity search in a high-dimensional spaces

Ivan Viсtorovich Sudos · 2012

High dimensional data is often analysed resorting to its distribution properties in subspaces. Subspace clustering is a powerfull method for elicication of high dimensional data features. The result of subspace clustering can be an essential base for building indexing structures and further data search. However, a high number of subspaces and data instances can conceal a high number of subspace clusters some of which are difficult to analyse within search algorithm. This paper presents a model of generic indexing approach based on detected subspace clusters and the way to find an optimal set of clusters to have an acceptable tradeoff between search speed and relevance.

Read the paper · More papers on PaperTik