uCLUST - A NEW ALGORITHM FOR CLUSTERING UNSTRUCTURED DATA

D.V. Prasad, Sathya Madhusudanan, Suresh Jaganathan · 2015

Data that resides in a fixed field within a record or file is called structured data and have a defined schema. Unstructured Data refers to information that either does not have a pre-defined data model and does not fit well into relational tables. Clustering gains importance in the fields of Libraries (book ordering), Insurance (identifying groups and identifying frauds), WWW (document classification and clustering weblog data). Available clustering algorithms work only with structured data and use medoids as parameter for clustering. Clustering big data is not feasible, as they are mostly unstructured. It is not possible to label large collection of objects and identifying subspace clusters in unstructured data is a difficult task because of time complexity. In this paper, we proposed and designed a new algorithm called uCLUST, which identifies clusters in unstructured data as traditional distance functions cannot capture the pattern similarity among the objects. The proposed algorithm is applied in 6 different datasets and results are tabulated.

Read the paper · More papers on PaperTik