Querying the Greek Web in Greeklish

Paraskevi Tzekou, Sofia Stamou, Νικόλαος Ζώτος, Lefteris Kozanidis · 2007

In this paper, we experimentally study the problem of querying the web in a hybrid language, namely Greeklish. Greeklish is the transliteration of Greek in Latin characters of the ASCII code. Although Greeklish emerged as a convenient mean for the creation and distribution of digital data at a time when Unicode Transformation Format was not supported for the Greek alphabet, nevertheless it is still being utilized as a matter of habit or need. Today, a considerable amount of the Greek web data contains pages written in Greeklish. Although, these are less official web pages and they appear mainly in blogs or forums, their contents may be of good quality and usefulness to the Greek online information seekers. However, the paradox of searching the Greek web is that search engines perceive Greeklish as a totally different language form Greek and as such they do not return Greek pages in response to Greeklish queries. As a consequence, users who issue Greeklish queries (sometimes for technical reasons) are systematically deprived of information that would otherwise be valuable to their search intentions. In an analogous manner, searching the web via Greek queries excludes from the search results pages of valuable content simply because they are written in Greeklish. In this paper, we study the phenomenon of Greeklish web searches and we propose a model that treats Greek and Greeklish web data in a uniform manner. Our aim is to improve the usability of Greek search engines and ameliorate the user experience, regardless of the preferred query alphabet.

Read the paper · More papers on PaperTik