A Perspective on Text Classification, Clustering, and Named-entity Recognition in Social Media
Kia Jahanbin, Fereshte Rahmanian, Vahid Rahmanian, Masihollah Shakeri, Heshmatollah Shakeri, Zhila Rahmanian, Abdolreza Sotoodeh Jahromi · AMBIENT SCIENCE · 2019
Introduction:Knowledge Discovery is the process of extracting implied effective, and fresh valuable information from data (Fayyad ., 1996;Frawley ., 1992).Data Mining has used specif ic algorithms for inventive patterns from data.KDD objects at discovering concealed patterns and connections in the data.Knowledge discovery from the text (KDT or Text Mining) was f irst introduced by Feldman & Dagan (1995), refers to the process of extracting high quality of information from structured; such as RDBMS data (Akbari ., 2018; Chen ., 1996), semi-structured; such as XML and JSON, and unstructured text resources; such as word documents, videos, and images (Pouriyeh & Doroodchi, 2009;Pouriyeh , 2010).After the emergence of social media, an enormous amount of data started to generate.The researchers call this, the Big Data which has three key features, known by "3V": Volume, Variety, and Velocity.Some researchers add two other characteristics: Value and Veracity, and thus speak about "5V" (Jahanbin & Mehrjoo, 2017;Lomotey & Deters, 2014).Mukkamala .(2014) and Nguyen .(2015) use interchangeably the terms Big Social Data and Social Big Data to refer the overall data created by social media.To demonstrate the huge amount of data generated by social media aff irms that through Facebook, 10 million photos are uploaded every day.Marrin (2014) highlights the fact that more than 300 million tweets are sent to Twitter every day too, and 3000 photos are uploaded from Flickr every