Retrieval of optimal subspace clusters set for an effective similarity search in a high-dimensional spaces
Ivan Viсtorovich Sudos · 2012
High dimensional data is often analysed resorting to its distribution properties in subspaces. Subspace clustering is a powerfull method for elicication of high dimensional data features. The result of subspace clustering can be an essential base for building indexing structures and further data search. However, a high number of subspaces and data instances can conceal a high number of subspace clusters some of which are difficult to analyse within search algorithm. This paper presents a model of generic indexing approach based on detected subspace clusters and the way to find an optimal set of clusters to have an acceptable tradeoff between search speed and relevance.