DATA CLEANINGTOOL: USAGEOFFUZZYROUGHSETTHEORY AS MACHINE LEARNINGPRE-PROCESSING

B Hameed, Ahmed A. Elfetouh, M Abu_Elkheir · International journal of intelligent computing and information sciences/International Journal of Intelligent Computing and Information Sciences · 2015

Real-world data is often incomplete, inconsistent, and/or lacking in certain behaviors ortrends, and is likely to contain many errors. Data preprocessing is a crucial phase in the data miningprocess that involves techniques toresolve such issues. Feature selection is a popular datapreprocessing procedure that is focused on omitting attributes from decision systems while stillmaintain the ability of those decision systems to distinguish different decision classes. A popular way toevaluate attribute subsets with respect to this criterion is based on the notion of dependency degree. Inthis paper, we conduct an experimental study using the generalized classical rough set framework fordata-based attribute selection and reduction, based on the notion of fuzzy decision reducts to evaluatethe viability of using Fuzzy rough subset feature. Experimental results shows that, general optimizationcan be achieved under average accuracy reduction, ±10.7 %, against high reduction rate overattributesranging from 36% to 97% and over instances from 1.7% to 44%.

Read the paper · More papers on PaperTik