A novel approach for clustering hindi tweets

Ashish Varghese, Pooja Awate, Akshata Inde, Tejas Jadhav, Kavita Moholkar · 2017

Social media online platforms have critical data. The information accessible on the Social Networking platforms isn't just enormous in volume, yet in addition in assorted variety of subject and quality. Such factors make it infeasible to productively rummage helpful data from it. Twitter is an online news and person to person communication benefit where user posts and co-operate with messages, “Tweets”. Tweets on Twitter contains diverse news and events. Thus, this paper intends to order these posts into various classes utilizing WSD (Word Sense Disambiguation) with the goal that the user will have the capacity to get to the class of tweets he needs to search for. For grouping the tweets, we utilize knn (K - Nearest Neighbour) algorithm. Word sense disambiguation is a process of automatically figuring out the intended meaning of such words when used in a sentence. For this purpose, we used LESK algorithm based on two assumptions. First assumption is, when two words are used in closed proximity in a sentence. Second, if one sense each of the two words can be used to talk of the same topic.

Read the paper · More papers on PaperTik