Large-Scale Thai Statistical Machine Translation

Glenn Slayden, Mei-Yuh Hwang, Lee Schwartz · 2010

Thai language text presents unique challenges for integration into large-scale multi-language statistical machine translation (SMT) systems, largely stemming from the nominal lack of punctuation and inter-word space. We review our independent solutions for Thai character sequence normalization, tokenization, typed-entity identification, sentencebreaking, and text re-spacing. We describe a general maximum entropy-based classifier for sentence breaking, whose algorithm can be easily extended to other languages such as Arabic. After integration of all components, we obtain a final translation BLEU score of 0.19 for English to Thai and 0.21 for Thai to English. 1 MT System Description Our existing multilingual SMT system is based on hybrid generative/discriminative models, and we extend this approach for both English-to-Thai and Thaito-English MT. Given a sequence of foreign words, f, its best translation is the sequence of target words, e, that maximizes e∗ = argmaxe p(e|f) = argmaxe p(f|e)p(e) = argmaxe { log p(f|e) + log p(e)} where the translation model p(f|e) is computed on at least a dozen features, as briefly described in Section 1.1. The target language model (LM), p(e) , is represented by a smoothed n-grams (Chen 1996) and sometimes more than one LM is adopted in practice. To achieve the best performance, the log likelihoods evaluated by these features/models are linearly combined. After p(f|e) and p(e) are trained, the combination weights λ are tuned on a held-out dataset to optimize an objective function, which we set to be the BLEU score (Papineni et al. 2002): {λ∗} = max{ } BLEU({e∗}, {r}) e∗ = argmaxe { λ log p (f|e) + λ log p (e)} where {r} is the set of gold translations for the given input source sentences. To learn λ we use the algorithm described by Och (2003), where the decoder output at any point is approximated using n-best lists, allowing an optimal line search to be employed. 1.1 English-Thai Treelet Translation Since we have a real-time rule-based English parser available, we based our English-Thai system on the “treelet” concept suggested in Menezes and Quirk (2008). This approach first parses the source language into a dependency tree which includes part-of-speech labels. A by-product of parsing is word breaking. For example, John’s is parsed into John (noun) and ’s (possessive). During training we also break Thai sentences into tokens (Sections 2 and 3). Source and target sentences are lower-cased, to be able to share the training data between different casings. Next we apply worddependent Hidden Markov Model (WDHMM) alignment (He 2007) to learn the mapping between the source and target token streams. Given a parallel corpus with these word alignments and dependency parses, we extract both treelet translation pairs and order templates. From each aligned sub-graph of the source tree, we extract that sub-graph along with its translation, and collate these across the whole dataset to produce a set of treelets with their translations. Likewise from each source dependency tree node along with its full daughter set, we also extract an order template that specifies the relative order of those daughters in the target language, as well as any inserted or deleted words. Order templates are also collated across the whole dataset to produce a distribution of possible reordering and word insertion or deletion. Taking a parsed source sentence, a derivation is a set of treelet translation pairs and order templates that completely cover the input dependency parse; the target side of this derivation may be read off to produce a target language string. To translate a given source sentence, we first parse the sentence, then find the best derivation for that sentence according to the log linear combination method as described above. This system employs eleven features altogether. The first two features are the log probabilities of each source treelet given its target side treelet and each target treelet given its source side using relative frequency estimates. The next two features are again conditional log probabilities, estimated using Model 1 (word-based probabilities) instead (Brown et al. 1993). Another feature is the log probability of the target side of each order template given its source side, again estimated using relative frequency. In the case that no template matches a given source configuration, we generate an order template that preserves the source order; we include a feature that counts the number of such source order templates. We also in1 Although Thai does not use orthographic case, this operation affects Roman alphabet text which may appear in Thai input texts. clude features to count the number of words inserted or deleted by order templates. The log probability of the target translation according to an n-gram language model acts as another feature. Finally, we add a word count feature to counteract any language model preference for shorter outputs, and a phrase count feature to bias the search toward derivations with fewer and thus larger treelets. After the treelet decoder outputs the translated Thai word sequences, we re-space the output to comply with Thai convention (see Section 5) based on linguistic rules and transfer of syntactic information between source and target dependency trees. 1.2 Thai-English Phrasal Translation Lacking a Thai parser, we use a purely statistical phrasal translator after Pharaoh (Koehn 2004) for Thai-to-English translation. As with the treelet system, Thai sentences are normalized, lower-cased, and tokenized. English words are broken mainly by whitespace, with a few rules to break contractions, and then lower-cased. Again WDHMM is used to align the parallel corpus to obtain phrase pair mappings. Similar 11 features are combined linearly in the log probability domain to evaluate a given translation. 2 Thai Character Sequence Normalization The four Thai tone marks and some Thai vowel glyphs are superand/or sub-scripted with respect to a base character. These combining marks are represented by individual Unicode code points. When two or more of these combining marks are present on the same base character, the ordering of these code points in memory should be consistent so that orthographically identical entities are recognized as equivalent by the MT system. However, many computer word processors do not enforce the correct sequence and/or do not properly indicate incorrect sequences to the user visually. 2 This often results in documents with invalid byte sequences. Correcting these errors is desirable for SMT input. For this task, we developed a Finite State Transducer (FST) which detects and repairs a number of sequencing errors which render Thai text either linguistically invalid, or not in a correct Unicode sequence. For example, a superscripted Thai tone mark should follow a superor sub-scripted Thai vowel when they both apply to the same consonant. When the input has the tone mark and the vowel glyph swapped, the input can be fully repaired: อ า ◌ น ➜ อ ◌ า น ➜ อาน อ ◌ ◌ น ➜ อ ◌ ◌ น ➜ อ น Figure 1. Two unambiguous repairs Other cases are ambiguous. The occurrence of multiple adjacent vowel glyphs is an error where the original typist’s intention is not clear. We retain the first2 Windows Vista helps the situation by visually distinguishing certain incorrect character sequences. appearing glyph, unless it is a pre-posed vowel, in which case we retain the last-appearing instance. These two treatments are contrasted in Figure 2.

Read the paper · More papers on PaperTik