Natural Language Processing For Content Analysis in Social Networking

A. A. Sattikar, R. V. Kulkarni · 2012

After going through review of almost twenty-five researches on security and privacy issues of social networking, researchers observed that users have strong expectations for privacy on Social Networking web sites as in Social Networking blogs or posts content is surrounded by a very high degree of abusive or defaming content .Social networking has emerged as the most important source of communication in the world. But a huge controversy has continued in full force over supervising offensive content on internet stages. Often the abusive content is interspersed with the main content leaving no clean boundaries between them. Therefore, it is essential to be able to identify abusive content of posts and automatically rate the content according the degree of abusive content in it. While most existing approaches rely on prior knowledge of website specific templates and hand-crafted rules specific to websites for extraction of relevant content, HTML DOM analysis and visual layout analysis approaches have sometimes been used, but for higher accuracy in content extraction, the analyzing software needs to mimic a human user and under- stand content in natural language similar to the way humans intuitively do in order to eliminate noisy content. In this paper, we describe a combination of HTML DOM analysis and Natural Language Processing (NLP) techniques for rating the blogs and posts with automated extractions of abusive contents from them.

Read the paper · More papers on PaperTik