Use of Patterns for Detection of Likely Answer Strings: A Systematic Approach.

Martin M. Soubbotin, Sergei M. Soubbotin · Text REtrieval Conference · 2002

The paper describes the Question Answering approach applied first at TREC-10 QA track and developed systematically in TREC 2002 experiments. The approach is based on the assumption that answers can be identified by their correspondence to formulas describing the structure of strings carrying certain (generalized) semantics, supposed by the question type. These formulas, or patterns, are like regular expressions but include elements corresponding to predefined lists of terms. Complex can be constructed from blocks corresponding to such entities as persons' or organizations' names, posts, dates, locations, etc. Using various combinations of blocks and intermediate syntactic elements allows to build a great variety of patterns. Exact position of elements corresponding to the exact was localized within the structure of each pattern. Each is characterized by a generalized semantics, thus the pattern-matching string must be checked for correlation with the question terms and/or their synonyms/substitutes. Essentials of the Approach In 2002 TREC QA track tests we have further developed the approach described in [Soubbotin, 2001]. In general, our method lies in the domain of approaches examining the potential of extraction for question answering tasks [Srihari, Wei Li, 1999; De Boni, 2001]. The evolution of IE systems, as represented, in particular, at Message Understanding Conferences (MUCs), shows a certain shift from deep text analysis based on computational linguistic and NLP methods to surface techniques [Eagles, 1998]. Our approach can be considered as being in line with this tendency. More specifically, our approach is based on the use of formulas describing the structure of strings likely bearing certain information. For example, string Director Louis Freeh can be recognized, according to one of such formulas, as likely bearing the following information: a person represented by his/her first and last names occupies a (leading) post in an organization. The formula for this string is: a word composed of capital letters; an item from the list of posts in an organization; an item from the list of first names; a capitalized word. We can mark two first items in this formula as exact if we want to get answer to the question Who is Louis Freeh?, and two last items, if the question is Who is FBI head? (question 1583 at TREC 2002). First used at TREC-10 QA track, formulas of such kind were called patterns [Soubbotin M.M and Soubbotin S.M, 2001]. The term pattern is widely used in the field of Information Extraction. Our concept of as structural formulas for strings is obviously different from that in traditional IE field, but keeping this difference in mind, we consider it convenient to use this term. Each is characterized by a certain generalized semantics, because the formulas' items refer to certain (e.g., posts) and not to specific units (e.g., president, head, director). Therefore, after a string corresponding to a formula is recognized, the next step is to identify the question terms (or their synonyms/substitutes) within it or in its surrounding. To increase the likelihood of getting the right answer, the surrounding of the found string must be checked for the presence of expressions negating its semantics (e.g., former, -elect, deputy, etc., located before or after the term from the list of posts). After a question's type is defined (e.g., question about a person occupying certain post in an organization, question about husband/wife/relative of a person, question about acronym, etc.), a set of formulas, prepared for this type, is applied to match the strings in question-relevant passages. Our approach does not need to distinguish linguistic entities in the text. We handle the source text strictly as string, i.e. consisting only of characters. The used in our QA approach are aimed only at recognizing sequences of elements that correspond to the predefined formulas. As surface patterns, our formulas for strings are similar to wrappers [Adams, 2001; Kushmerick, 2000] and look like regular expressions. However, used by the wrapper techniques are mostly resource-specific, they relate to the document formats rather than the ways is presented in written texts per se. As for difference from regular expressions, it is worth noting that patterns, that we use, include elements referring to the lists of predefined words/phrases. Currently, increased attention is seen on surface approaches in QA. In some recent publications surface similar to those used by us were discussed [Magnini, et al., 2002; Brill, et al., 2002; Brill, et al., 2001; Ravichandran and Hovy, 2002; Hovy et al., 2002]. Patterns and Question Types The IE task, as presented at its main forum the Message Understanding Conferences (MUCs), is focused on certain topics, or domains (Terrorism, Management Successions, Natural Disasters, Outbreaks of Infectious Diseases, etc.). The QA task requires another way to categorize the addressed Information. The usual praxis of TRECs' QA tracks participants is to predefine a set of potential question types. The questions accumulated from several TRECs represent a good source for defining question types on a more or less detailed basis. The paradigm of information categories defined by question types (in contrast to topic/domain paradigm) allows to create systematically a variety of patterns, basing on potential relationships inside each question category. So, for the question type Who is person X? we can presuppose among the main alternative possibilities that this person is known for the (top-level) position he/she occupies in a organization, company or government; for his/her contributions as author, inventor, founder, etc.; as outstanding figure in a professional area; as wife/husband/relative of a well-known person; as involved in well-known event (e.g., as a criminal/perpetrator). In each case, a relationship is established between two or more entities: person, post, and organization/company; author and work; etc. The same entities are present if the Who-questions refer to posts, authors, etc. (e.g., Who occupies the post Y in the organization Z?.) For most Where-questions, we can suggest geographical items as answers. This is achieved by constructing structural formulas like: item from the list of cities/towns/counties, etc.; comma; item from the list of countries/states. There are question types suggesting as answers combinations of digits with units of measurement or currencies names. Completeness of lists corresponding to semantic elements is evidently important (e.g., the list of currencies must include not frequently used words, such as dlrs). The type of the processed question is defined basing both on its interrogative and on the presence of words/expressions that are included in the list of characteristic terms for the corresponding question type.

Read the paper · More papers on PaperTik