Information retrieval on Turkish texts

Fazlı Can, Seyit Koçberber, Erman Balçık, Cihan Kaynak, Huseyin Cagdas Ocalan, Onur M. Vursavas · Journal of the American Society for Information Science and Technology · 2007

Abstract In this study, we investigate information retrieval (IR) on Turkish texts using a large‐scale test collection that contains 408,305 documents and 72 ad hoc queries. We examine the effects of several stemming options and query‐document matching functions on retrieval performance. We show that a simple word truncation approach, a word truncation approach that uses language‐dependent corpus statistics, and an elaborate lemmatizer‐based stemmer provide similar retrieval effectiveness in Turkish IR. We investigate the effects of a range of search conditions on the retrieval performance; these include scalability issues, query and document length effects, and the use of stopword list in indexing.

Read the paper · More papers on PaperTik