Text Summarization Using XML-Tagged Documents
Kenneth C. Litkowski · 2003
CL Research’s participation in the Document Understanding Conference extended the framework used in the TREC 2003 question-answering track, in which texts are parsed and processed into XML-tagged documents where sentence elements are marked with discourse, syntactic, and semantic attributes. This extension was made primarily to test the viability of using XML-tagged documents for summarization. The extension of the Knowledge Management System was able to take advantage of these attributes in implementing various text summarization capabilities. While implementation of these capabilities made little use of current summarization technologies, the CL Research system performed at a higher than expected level, finishing first in mean length-adjusted coverage for summaries against a provided viewpoint. The system performed less well on this measure in event summarization (tenth), novelty summarization (fifth), and headline generation (eleventh), but performed well on quality measures (finishing first among teams participating in all tasks)) and relevance (finishing first on each summarization task, with all sentences in these tasks judged relevant to the topic). The system’s performance arises primarily from the use of an “antecedent” tag attached to referring expressions (such as pronouns) within a document. In particular, when accumulating word frequencies, the antecedent was used instead of the referring expression; thus, instead of treating a pronoun as a word in the frequency count, its antecedent was used. The system’s performance demonstrates the basic viability of using XML-tagged documents. Many options were explored in setting up the summarization capability, indicating considerable flexibility in examining documents from many perspectives and considerable potential in possible further improvements in the system. The system can indicate not only that a concept appears frequently within a document, but also how it is used (e.g., as subject, verb object, or prepositional object). More specifically, the availability of considerable structural information within documents permits a relatively simple examination of phenomena that have been used in text summarization, as well as the creation of a document’s semantic network.