Towards Syntactically Contrained Statistical Word Alignment
Greg Hanneman · 2008
In statistical machine translation, the fundamental problem of word alignment is the process of finding word-to-word connections (i.e. translations) across languages given a sentence in one language and its translation in another. In more formal terms, given a source-language sentence F of n words (f1, f2, ..., fn) and a target-language sentence E of m words (e1, e2, ..., em), an alignment is a mapping between subsets of F (elements of the power set 2 ) and subsets of E (elements of 2). Instead of a mapping between full subsets, an alignment is usually indicated as a collection of links, each of which connects some fj (1 ≤ j ≤ n) to some ei (1 ≤ i ≤ m). The total collection of links makes up the alignment for the given sentence pair. In the general case, the total number of possible alignments, called the alignment space, is extremely large. With no restrictions in place and an n-word sentence pair, there are n possible alignment links and 2 2 possible alignments. If a one-to-one constraint is enforced, such that one word in F may only align to one word in E, this exponential space can be reduced to n!. Additional constraints may further restrict the alignment space or lead to related spaces (Cherry and Lin, 2006a). The natural goal of constrained alignment is to restrict the alignment space in such a way that “bad” or linguistically very unlikely alignments are ruled out while “good” or linguistically sound alignments remain possible or are preferred. Word alignment is most commonly carried out within the scope of a parallel sentence represented as a flat stream of plain-text words or as a flat stream of sets of feature–value pairs. However, in the realm of natural language, it is also possible to represent the structure inherent in a sentence; further, the structure can provide useful information about what alignments are “good” and what alignments are “bad” beyond what information can be extracted from a flat string. In this paper, we will consider a number of techniques for representing different levels of syntactic structure in the alignment process and examine the benefit of the information it provides. First, in Section 2, we briefly describe basic statistical alignment models that do not take into account any overt representation of the syntax of the sentence they are aligning. A number of published extensions to or replacements for the base models will be discussed in Section 3; these approaches all explicitly model some level of structure on one or both sides of the parallel sentence pair. Section 4 considers tradeoffs that these models introduce, compares their expressive and restrictive powers, and concludes the paper with some possible avenues for future alignment research.