Bayes Discriminator for BBS Documents Based on Latent Semantic Analysis
Chang Liu · Chinese Journal of Computers · 2004
With the rapid development of Internet, the abuse and misuse of BBS become a social problem of information pollution and call on the demand to the discrimination techniques for BBS document. Borrowing the techniques from data mining, probability-statistics and Natural Language Understanding, this paper proposes a new discrimination method for BBS document, called Bayes Discrimination based on Latent Semantic Analysis(BDLSA). The main steps of the new method includes following steps: (1)Makes typical phrase set by extracting the typical sentences from training documents in preprocessing stage with natural language understanding techniques.(2)Applies synonymy reduction on typical phrases by Latent Semantic Analysis.(3)Discovers the association rules between typical phrases to increase the independency of phrases so that the traditional Bayes discriminator works efficiently.(4)Discriminates BBS document by Bayes classifier. The algorithms to construct typical phrase set and to reduce synonymy are proposed and implemented. The experiment is based on real document form Web, with training data of 583 documents and test-data of 308 documents, the correctness is up to 75%. This shows the effetiveness and validation of the new method.