The role of automated word classification in the summarization of the contents of sets of documents

Susan Gauch · 2003

In future digital libraries, even perfect retrieval will typically return too much material for a user to cope with. One way to deal with this problem is to produce automated summaries tailored to the user's requirements. One of the prime purposes of a summary of a collection of documents is to collapse together all of the important information elements that are common to the collection. This requires some method of discovering classes of similar items, e.g., word classes. This paper describes automated techniques for placing words in similarity classes. To do this, each target word is described by a composite vector that records the occurrence of words positioned near any occurrence of the target. Target words with similar contexts are grouped together by a clustering algorithm. We describe how such classifications can be used in information retrieval and for the summarization of biological literature. The dilemma of perfect retrieval In retrieving documents or portions of full-text documents, recall is the percentage of the desired documents that are retrieved and precision is the percentage of all retrieved documents that are of the desired type. No matter how good future systems become, even if they achieved 100% recall and precision, the amount of information that will be on line will be so large that the user will still be overwhelmed. It will rarely be the case that one returned paragraph or even one entire document will answer the user's questions. The information the user wants is typically scattered throughout the documents simply because none of the documents were written (nor could they have been written) to satisfy the interests that one particular user would have at some later time. The user could look for a review of the topic, but again, there would probably not be a review focused on the user's interests, much less one that was as up-to-date as the literature itself. Because the information desired is scattered across many documents, ranking the documents in order of relevance does not solve the problem. One solution: Summarizing document sets Retrieval systems could help to avoid the dilemma above if they could automatically produce a summary of the relevant documents tailored to the user's interests, particular query, level of expertise and adjusted to some particular length (from a paragraph to many pages). There has been work on extracting information from single sentences, from paragraphs (Zadrozny & Jensen, 1991), work on summarizing the arguments in whole documents (Alvarado, 1990) and work on automatic abstracting (Paice, 1990). Extensions of these techniques can be applied to summarizing the contents of sets of documents. Manual analysis of reviews in the biological and computer science literature reveals the strategies authors use to summarize large collections of literature. One of the primary devices is to generate that describe items of a given class, citing the appropriate sources. The listings could be sets of genes or enzymes in biological articles or sorting algorithms or network protocols in computer science. Discovering the set of items in a given class in a document collection would need to be automated for this strategy to succeed. It is not appropriate to say that the system should refer to some standard listing of the items of a given class, because new terms are constantly being introduced in rapidly moving fields such as computer science or biology. Furthermore, terminology and use is often specific to a given subfield. As an example, we would like an automated summary system to produce tables such as the following for a biological topic, Term Context and source λ repres sor OL and OR each contain a series of nonidentical binding sites for the λ repressor... [Stryer, 1975] Pages 35-39 of Genetic Switch [Ptashne, 1992]. lac repres sor The repressor of the lactose operon [Stryer, 1975] In a navigation (hypertext) environment the user could select any of the items in the table for expansion. In order to select the terms that should be grouped together in tabular summaries as in the example above (λ and lac), word classification must be done. This is described in the next section. Describing and quantifying word contexts To discover word classes, we describe the context of a word (the target word) by the preceding two context words and the following two context words. Each context position is represented by a vector containing the joint frequencies of the 150 highest frequency words in the corpus, giving a 600-dimensional context vector. The entries in the context vectors are converted to mutual information measures, with smoothing. The similarities of the resultant context vectors for the 1,000 highest frequency words are computed from the normalized inner products of their context vectors (cosine rule). The resulting set of 500,000 similarities is used as the basis of a hierarchical clustering algorithm, a bottom-up approach producing binary trees with a similarity at each node, -1.0 ≤ ≤ 1.0. The method was inspired by (Finch & Chater, 1992) and is described in more detail in (Futrelle & Gauch, 1993). Near the leaves, the words were found to be grouped by both semantic and syntactic similarity. Further up the tree, the larger classes retained only syntactic similarity. 1 Prof. Gauch's current address: Dept. of Computer Science, U. Kansas, Lawrence, KS. Some examples from the biological literature The corpus used for this analysis was the 220,000 words of text in 1,700 abstracts that completely cover the field of bacterial chemotaxis since its inception in 1965. Bacterial chemotaxis is a phenomena in which single bacteria move toward higher concentrations of chemical attractants such as sugars (and away from repellents). One of the classes of terms that is constantly being added to by biologists is genetic mutant designators. One class of these the system discovered consists of ten items: motB, tar, tsr, cheB, cheZ, cheY, cheA, flaA, flaE, There are two apparent anomalies in this list, tar and double, both common words in other contexts. The utility of the classification method is that it is sensitive to the particular use of these words in this specialized field. tar means taxis towards aspartate in this field and double is used to describe mutants which have two lesions in the same or different genes. Thus, if a table of mutants were constructed to summarize this set of papers it should include all ten items. The following class contains compounds that are attractants used in chemotaxis studies, aspartate, maltose, galactose, ribose, serine These could usefully be placed in a list summarizing the major compounds of interest. The word classes also include , which are fundamental to the understanding of living systems, chemotaxis, taxis, sensing, motility, rotation, behavior, movement, transport, uptake Again, a tabulation of these along with excerpts describing them or references to articles devoted to them would be useful as part of a summary. Note that the word classes shown above are both syntactically and semantically homogeneous. The examples above contain only nouns. The homogeneity is easily seen from some other classes generated by the system, adjectives: higher, lower, greater, less other, several, many molecular, structural nouns (physical units): degrees, min, s, mM, microM, nm

Read the paper · More papers on PaperTik