Web Document Classification Based on Fuzzy-Rough Set

Ding Hua-fu · Computer Technology and Development · 2010

The diversity and variability of network information brings great difficulty to information management and information filtering.Put forward a method to Web document classification based on fuzzy-rough set in order to improve the speed and accuracy of network information classification and use machine learning method for training and testing Web document.In the training process,firstly,representing preprocessed Web documents by vector space model,forming initial attribution features space and conducting weight value computing.Then,conducting information filtering and reducing attribution feature space by fuzzy-rough set algorithm,forming classification rules.Finally,classifying documents by multiple knowledge bases.In the testing process,matching key attributes directly and computing weight value by the fuzzy-rough factor,then classifying document by space distance method.The experiment results and the comparison with others show that this Web document classification has good classification performance.

Read the paper · More papers on PaperTik