QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set Operations

Chaitanya Malaviya, Peter Shaw, Ming‐Wei Chang, Kenton Lee, Kristina Toutanova · 2023

Formulating selective information needs results in queries that implicitly specify set operations, such as intersection, union, and difference.For instance, one might search for "shorebirds that are not sandpipers" or "science-fiction films shot in England".To study the ability of retrieval systems to meet such information needs, we construct QUEST, a dataset of 3357 natural language queries with implicit set operations, that map to a set of entities corresponding to Wikipedia documents.The dataset challenges models to match multiple constraints mentioned in queries with corresponding evidence in documents and correctly perform various set operations.The dataset is constructed semi-automatically using Wikipedia category names.Queries are automatically composed from individual categories, then paraphrased and further validated for naturalness and fluency by crowdworkers.Crowdworkers also assess the relevance of entities based on their documents and highlight attribution of query constraints to spans of document text.We analyze several modern retrieval systems, finding that they often struggle on such queries.Queries involving negation and conjunction are particularly challenging and systems are further challenged with combinations of these operations. 1 * Work done during an internship at Google.retrieving an exhaustive document set, instead lim-65 iting annotation to the top few results of a baseline 66 information retrieval system.67 To analyze how well retrieval systems handle 68 such queries, we present QUEST, a dataset with 69 natural language queries from four domains, that 70 are mapped to relatively comprehensive sets of en-71 tities corresponding to Wikipedia pages.We use 72 Wikipedia categories and their mapping to entities 73 in Wikipedia as a building block for our dataset 74 construction approach, but do not allow access to 75 this semi-structured data source at inference time, 76 to simulate text-based retrieval.Wikipedia cate-77 gories represent a broad set of natural language 78 descriptions of entity properties and often corre-79 spond to selective information need queries that 80 could be plausibly issued by a search engine user 81 ([At least 90% of the time based on our filtering?]).82 The correspondence between property names and 83 document text is also often subtle and requires so-84 phisticated reasoning to determine relevance, rep-85 resenting the natural language inference challenge 86 inherent in the task, while the knowledge of cate-87 gory membership allows us to construct relatively 88 comprehensive sets of candidate entities for atomic 89 categories and their combinations.90 Our dataset construction process is outlined in 91 Figure 1.The base queries in our dataset are 92 semi-automatically generated using Wikipedia cat-93 egory names.To construct queries, we sample 94 category names and compose them into complex 95 queries by using pre-defined templates (for exam-96 ple, A \ B \ C).Next, we ask crowdworkers to 97 paraphrase these automatically generated queries, 98 while ensuring that the paraphrased queries are 99 fluent and clearly describe what a user could be 00 looking for.These are then validated for natural-01 ness and fluency by a different set of crowdworkers, 02 and filtered according to those criteria.Finally, for 03 a large subset of our dataset, we collect scalar rel-04 evance labels based on the entity documents, and 05 textual attributions mapping query constraints to 06 spans of document text, to aid the development of 07 systems that can make precise inferences based on 08 trusted sources.09 Performing well on this dataset requires sys-10 tems that can match query constraints with cor-11

Read the paper · More papers on PaperTik