Rough sets used in the measurement of similarity of mixed mode data

Sarah Coppock, Lawrence J. Mazlack · 2004

Similarity is important in knowledge discovery. Cluster analysis, classification, and granulation each involve some notion or definition of similarity. The measurement of similarity is selected based on the domain and distribution of the data. Even within a specific domain, some similarity metrics may be considered more useful than others. There is an amount of uncertainty in quantitatively measuring the similarity between records of mixed data. The uncertainty develops from the lack of scale that both nominal and ordinal data have. Rough set theory is one tool developed for handling uncertainty. Rough sets can be used in dissimilarity analysis of qualitative data. It would seem that rough sets could be applied in measuring similarity between records containing both quantitative and qualitative data for the purpose of clustering the records.

Read the paper · More papers on PaperTik