Ambiguity management in natural language generation
Martin Kay, Hadar Shemtov · 1997
This dissertation is concerned with ambiguity in natural language generation processes. The core of this work is a view of generation as a many to many relation, between multiple meanings and multiple expressions of these meanings. A possible task for this generator is machine translation, for which it can provide the following advantages: Admitting multiple meanings eases the requirement of full disambiguation; Generating multiple paraphrases allows target language considerations to affect the realizations; Combining the two ideas, choices of meaning and of realization can be mutually constraining. The dissertation investigates the notion of ambiguity preservation. This notion refers to cases where the source and the translation have equivalent ambiguities. For example: (a) The man saw the girl on the hill with a telescope. (b) L'homme a vu la fille sur la colline avec une lunette. (c) Der Mann sah das Madchen auf dem Hugel mit dem Fernrohr. When ambiguities coincide in this manner, determining the actual intended meaning can be avoided. Careful analysis of data reveals that opportunities for preserving ambiguities are abundant, especially if partial preservation is envisaged. To achieve this, the dissertation proposes new generation algorithmS that support the view of generation as a many to many process. The algorithms are based on the parsing technique of packing multiple descriptions and they extend its advantages to the realm of generation. Instead of enumerating meanings and pursuing each of them independently, these generators operate on packed representations of the meaning, effectively pursuing all of the possibilities in parallel. The output they produce is also a packed representation of expressions. The key advantage of this is that common subexpressions are generated and represented only once. This approach allows an exhaustive exploration of the space of possible meanings, without the combinatorial explosion that is associated with such processes. Although the application worked out in detail is ambiguity preservation, this dissertation actually presents a general framework for exploring relationships between source and target texts. Other possible applications of it are finding paraphrases with particular properties, allowing target language readers to comprehend ambiguous source texts and more generally, enabling effective human-machine interactions in the translation process.