Identification of spam comments using natural language processing techniques
Cristina Radulescu, Mihaela Dînșoreanu, Rodica Potolea · 2014
The high popularity of modern web is partly due to the increase in the number of content sharing applications. The social tools provided by the content sharing applications allow online users to interact, to express their opinions and to read opinions from other users. However, spammers provide comments which are written intentionally to mislead users by redirecting them to web sites to increase their rating and to promote products less known on the market. Reading spam comments is a bad experience and a waste of time for most of the online users but can also be harming and cause damage to the reader. Research has been performed in this domain in order to identify and eliminate spam comments. Our goal is to detect comments which are likely to represent spam considering some indicators: a discontinuous text flow, inadequate and vulgar language or not related to a specific context. Our approach relies on machine learning algorithms and topic detection.