Rough sets similarity-based learning from databases
Xiaohua Tony Hu, Nick J. Cercone · 1995
Many data mining algorithms developed recently are based on inductive learning methods. Very few are based on similarity-based learning. How-ever, similarity-based learning accrues advan-tages, such as simple representations for con-cept descriptions, low incremental learning costs, small storage requirements, etc. We present a similarity-based learning method from databases in the context of rough set theory. Unlike the pre-vious similarity-based learning methods, which only consider the syntactic distance between in-stances and treat all attributes equally important in the similarity measure, our method can anal-yse the attribute in the databases by using rough set theory and identify the relevant attributes to the task attributes. We also eliminate superflu-ous attributes for the task attribute and assign a weight to the relevant attributes according to their significance to the task attributes. Our sim-ilarity measure takes into account the semantic information embedded in the databases.