A survey on online tweet segmentation for linguistic features

R.P. Narmadha, G.G. Sreeja · 2016

Tweet classification has been used for classifying the tweets based on different class labels contained in the tweets. NLP models have to be constructed in order to learn the tweets with linguistic feature. Before feature extraction of the data; tweets are pre-processed with stop word removal and Stemming process. Initially tweets will be organized into meaningful segments based on Local and global context of the tweet information. Tweet Summarization is proposed in order to avoid the overload problems due to diversity among the sentences, large number of tweets is meaningless, irrelevant and redundant. Further, tweets are strongly correlated with their posted time and new tweets tend to arrive at a very fast rate. Classification technique which establishes the optimal cluster of a tweet is carried in terms of tweet cluster vector. Additionally tweet vector cluster is established as potential sub-topic delegates and maintained dynamically in memory during stream processing. Data structure is used to store and organize cluster snapshots at different moments, thus allowing historical tweet data to be retrieved by any arbitrary time durations.

Read the paper · More papers on PaperTik