A strategy for automatic moderation of a large data set of users comments
Marcos Rodrigues Saúde, Marcelo de Medeiros Soares, Henrique Gomes Basoni, Patrick Marques Ciarelli, Elias Silva de Oliveira · 2014
The increase use of social media and Web 2.0 are daily drawing more people to participate and express their point of views about a variety of subjects. However, there are a huge number of comments which are offensives and sometimes non-politically corrects and so must be hindered from coming up online. This is pushing the services providers to be more careful with the contents they publish to avoid judicial claims. This work proposes the use of automatic classification techniques to identify and only allow to go online harmless comments. We applied various techniques regarding with data processing, such as weighting of terms and the dimensionality reduction. All these techniques have been studied in order to model algorithms to be able to mimic well the human decisions regarding to the comments. The results indicate that we are able to mimic experts decision on 96.78% in the data set used.