Frequently Asked Questions Retrieval for Croatian Based on Semantic Textual Similarity
Mladen Karan, Lovro Żmak, Jan Šnajder · Meeting of the Association for Computational Linguistics · 2013
Frequently asked questions (FAQ) are an efficient way of communicating domainspecific information to the users. Unlike general purpose retrieval engines, FAQ retrieval engines have to address the lexical gap between the query and the usually short answer. In this paper we describe the design and evaluation of a FAQ retrieval engine for Croatian. We frame the task as a binary classification problem, and train a model to classify each FAQ as either relevant or not relevant for a given query. We use a variety of semantic textual similarity features, including term overlap and vector space features. We train and evaluate on a FAQ test collection built specifically for this purpose. Our best-performing model reaches 0.47 of mean reciprocal rank, i.e., on average ranks the relevant answer among the top two returned answers.