Probabilistic model for a distributed feature selection method
Zsolt Berényi, István Vajk · 2009
When building topic based document classifiers, feature selection is a key step: features not holding any information about the topic of a document introduce only unnecessary noise during the classification. In a distributed environment, when the nodes are interacting, the locally retrieved features and the their attributes must be shared to have at every node a more accurate estimation of the global classifier. When expanding the knowledge of the local classifiers, to reduce costs, the network traffic should be kept to a minimum. We propose a probabilistic model for a keyword selection method which makes a more thorough analysis possible and can be used as a baseline when sharing information in a distributed environment. It can be used for incrementally building up the distributed classifiers ensuring minimal network traffic. This model can be refined later on by sending more content-related information to achieve higher performance. This probabilistic model together with experimental results are presented in this paper.