Contribution des basses fréquences à l'alignement sous-phrastique multilingue : une approche différentielle

Adrien Lardilleux · HAL (Le Centre pour la Communication Scientifique Directe) · 2010

The goal of this thesis dissertation is to show that, contrary to preconceived ideas, one can efficiently take advantage of low frequency words in natural language processing. We put them to use in sub-sentential alignment, which constitutes the first step of most data-driven machine translation systems (statistical or example-based machine translation). We show that rare words can be used as a foundation in the design of a multilingual sub-sentential alignment method, using differential techniques similar to those found in example-based machine translation. This method is truly multilingual, in that it allows the simultaneous processing of any number of languages. Moreover, it is very simple, anytime, and scales up naturally. We compare our implementation, Anymalign, to two statistical tools proven in the domain. Although its current results are in average slightly behind those of state of the art methods in phrase-based statistical machine translation, we show that the intrinsic quality of our lexicons is actually superior to that of lexicons produced by state of the art methods.

Read the paper · More papers on PaperTik