Multi-lingual sentiment analysis of Twitter data by using classification algorithms
Ankit Kumar Soni · 2017 Second International Conference on Electrical, Computer and Communication Technologies (ICECCT) · 2017
Big data is a term that defines data set are so large or complex that traditional data processing applications are inadequate. It is also defined by the 3V's i.e. Volume, Velocity and Variety. By volume we mean the enormous amount of data that the organizations collect from variety of sources. Velocity here stands for unprecedented speed via which the data streams in and needs to be handled within proper time limits. Lastly, data comes in variety of formats: multi-language, structured, unstructured, email etc. Challenges faced by the organizations with this huge amount of data includes analysis, capture, data curation, search, sharing, transfer, visualization, querying and information privacy. The need of big data doesn't revolve around what amount of data you have, but what you want to do with it. You can fetch the data from any of the source and analyze it to examine and find answers that enable 1) cost reductions, 2) time reductions, 3) new product development and optimized offerings, and 4) smart decision making. Not all the data that is collected is important for the user, so there's a need to refine this data in order to filter out the useful information as filtering the data can make your results more efficient. The aim of this paper is to use the classification techniques to develop a system for filtering out the useful information from the raw data and to analyze the sentiments carried out by Twitter micro blogging services. Million of tweet posted daily which contain opinion and sentiment of users around the world. This tweet having more than one language. Sentimental classification can benefit companies by analysis customer sentimental but to analysis Multilanguage dataset of tweet is the main challenge. Till today not a single solution is proposed for above mention problem.