Text Mining Analysis in Turkish Language Using Big Data Tools

Mehmet Çakır, Seren Guldamlasioglu · 2016

Gigantic amount of content is generated every second on the web. Blog posts, tweets, videos, uploaded images maintain the rich and increasing content of the social media. Great amount of information is obtained from social media through text mining techniques which are being improved and developed continuously. Due to the increasing variety of social media, data is flooding into every business and techniques of gathering information from unstructured data becomes a challenge. Processing and analyzing of large collections of social media text is difficult by using traditional single computer methods. Fortunately, with each passing day, the development and improvements in Big Data infrastructure are promising and the analysis capabilities are also increasing. This infrastructure supports language independent researches on text mining. There can be diversified searches with language-specific definitions and assumptions. Since many of the studies on social media text mining analysis focus on English language, this paper would help enhance text mining efficiently in Turkish language which is a word order free agglutinative language. The linguistic aspects of Turkish language make text mining complicated. Suffixes can change the word meanings easily. Word stemming becomes elaborated due to indigenous letters. The intent of this article is to demonstrate the issues, methods and applications of qualitative approaches that enable text mining in Turkish language considering its compelling features with covering preprocessing, feature selection and topic clustering on text mining considering Turkish linguistic features.

Read the paper · More papers on PaperTik