Survey on mining clusters using new k-mean algorithm from structured and unstructured data
T. Nelson Gnanaraj, K. Ramesh Kumar, N. Monica, M. Tech Scholar · 2014
Big data is the popular term used in the current era for extracting knowledge from large datasets. Bigdata is the collection of large and complex dataset. The challenge in big data is volume, variety and velocity (3V’s).variety can be classified into structured, unstructured data and semi structured. Structured data are the identifiable data, which is organized in some structure. Data stored in the relational database are example of structured data. Unstructured data are the data without identifiable structure, audio, video and images are few examples. All web and bioinformatics data comes under semi structure data which does not have any regular structure, it is neither structured nor semi structured. Clustering the one of the best technique in knowledge extraction process. It is nothing but grouping of similar data to form a clusters. The distance between the data in one clusters and other should not be less. Many algorithms are practiced for clustering, in that k-mean clustering is the one of the popular term for cluster analysis. The main aim of the algorithm is to partition the dataset into k clusters based on some computational value. The limitation of k-mean clustering is that it can be applied to either structured or unstructured, not in combination of both. This project overcomes that limitation by proposing new k –mean algorithm for extracting hidden knowledge by forming clusters from the combination of both structure and unstructured dataset.