Data Clustering by Large-Scale Adaptive Agent Systems

Elth Ogston, Maarten R. van Steen, Frances M.T. Brazier, Benno J. Overeinder · Digital Academic REpository of VU University Amsterdam (Vrije Universiteit Amsterdam) · 2005

Finding items with specific characteristics in large distributed systems can be problematic.Directories, the most common way of enabling capability-based search in distributed systems, gather in a central location data which is often inherently decentralized.This paper considers an alternative decentralized method of enabling search, which leaves data objects in place at their natural location within a physically distributed network.This form of decentralized search is likely to gain importance as computer systems become more widely distributed and the autonomy of their components increases.Autonomy plays a key role in this development since it precludes the global use of a single comparison method for determining the similarity of objects.With this in mind, this work explores the ability of explicitly autonomous peer-to-peer agents to form themselves into groups within an overlay network, based on their locally perceived similarity, thus providing a structure that can facilitate search.We introduce a decentralized clustering procedure, designed to discover small clusters (of 500 items or less) in very large, widely distributed sets of data.This procedure is shown to scale well in the number of clusters.The paper further demonstrates that for 2D spatial data it produces clusterings that compare well to those produced by central clustering algorithms, especially for data sets that contain widely varying cluster characteristics.

Read the paper · More papers on PaperTik