Latent variable models for machine translation and how to learn them
Philip Schulz · UvA-DARE (University of Amsterdam) · 2020
This thesis concerns itself with variation in parallel linguistic data and how to model it for the purpose of machine translation.It also reflects the paradigm shift from phrase-based to neural machine translation in that it addresses the variation phenomena in both frameworks.Machine translation is the task of automatically translating text between different languages.As such, any trainable machine translation system needs to be exposed to a lot of parallel training data, i.e. sets of sentences in two or more languages that we know to be translations of one another.This data is expensive to generate since translations need to be produced by human translators first.Of course, not all human translators perform their task identical.On the contrary, their translation outputs may vary wildly.This has to do with the personal style that every translator injects into their work as well as the translators proficiency in a particular language or domain.A more experienced translator will likely produce more accurate results than a newly trained one.Similarly, a translator who specialises in sports news may not be qualified to translate legal documents.Besides these differences between translators, a translator's performance can vary from day to day depending on factors such as motivation, fatigue, stress and the like.Finally, there is variation between languages.Many romance languages allow for the omission of pronouns under certain circumstances while the use of pronouns is mandatory in English.German and many Slavic languages employ grammatical gender, a concept unknown (and deeply confusing) to the anglophone parts of the world.For machine translation research this presents the following problem: the data is not homogeneous and a verbatim translation that may be appropriate in one context may be wrong in another.Moreover, translation systems are usually trained on a variety of documents from different sources which means that they encounter different linguistic styles.Users of machine translation systems these days expect not only an accurate translation that carries all the information of the original text, they also expect it to be grammatical and well-sound.I have v therefore taken to modelling at least some of the variation found in translation data.My main contention is that this improves the output translation as it relaxes the assumption of data homogeneity which we know to be false.Measuring translation quality with the commonly used BLEU I show experimentally that this indeed the case.This thesis begins with a short motivating introduction in Chapter 1.It then provides the necessary mathematical background in Chapter 2. Since the probabilistic models presented in this thesis necessitate the use of approximate inference techniques, a particular emphasis is places on these methods.The chapter also provides the reader with an introduction to phrase-based and neural machine translation.Chapter 3 introduces a new latent variable model to handle variation in word alignment.Word alignment is the first step in the phrase-based machine translation pipeline.It connects words across two parallel sentences which are likely translations of each other.These word-level translations are expanded to phrases in a later step which are in turn memorised by the translation system.A common assumption of many word alignment models is that each word in one of the languages needs to have a counterpart on the other side.This is of course false, since languages vary in how they express concepts.As mentioned earlier, some languages omit pronouns while others may omit prepositions.The reason that a pronoun occurs in sentence A and not in sentence B is thus entirely due to the grammatical requirements of language A and has nothing to do with translation.I therefore present a latent variable model that is a mixture of a classical alignment models and a language model component.The language model can account for grammatically induced words and thus prevents the alignment models from producing erroneous alignments.Experiments show that the resulting alignments lead to improved translations.In model presented in Chapter 4 approaches variation phenomena more holistically as it is embedded into an end-to-end neural machine translation system.The hypothesis underlying that model is that the sources of variation in translation are too numerous to annotate explicitly.The model therefore attributes all variation at a given word position in the translation data to a common noise source.The innovation here is that the noise sources evolves together with the translation.Noise is modelled on a word (or sub-word) level and changes according to the hitherto produced translation.The model is an instance of a deep generative model, in particular a variational autoencoder, and uses recent variational inference techniques that allow for gradient flow through stochastic computation graphs.Not only does the model outperform its baselines, it is also shown to produce different but accurate translations when the noise source is varied stochastically.The thesis concludes with Chapter 5.That chapter also provides an outlook on future research avenues for which I hope to have provided some of the groundwork.vi SamenvattingDeze dissertatie gaat over variatie in parallelle taalkundige data en over hoe dit gemodelleerd kan worden ten behoeve van machinevertaling.De dissertatie laat ook de paradigmaverschuiving zien van frase-gebaseerde naar neurale machinevertaling door het fenomeen van variatie te beschouwen vanuit beide paradigma's.Machinevertaling is het automatisch vertalen van tekst tussen verschillende talen.Om een machinevertalingssysteem goed te trainen, moet het blootgesteld worden aan veel parallelle trainingsdata, d.w.z., verzamelingen zinnen in twee of meer talen waarvan we weten dat ze vertalingen zijn van elkaar.Het is duur om deze data te genereren omdat zulke vertalingen door menselijke vertalers geproduceerd moeten worden.Natuurlijk produceren niet alle menselijke vertalers precies dezelfde vertalingen.Integendeel, hun vertalingen kunnen enorm uiteenlopen.Dit heeft zowel te maken met de persoonlijke stijl van iedere vertaler als met verschillen in expertise over een bepaald domein of in een bepaalde taal.Een meer ervaren vertaler produceert over het algemeen nauwkeurigere resultaten dan een minder ervaren vertaler.Ook is een vertaler die gespecialiseerd is in sportverslaggeving wellicht bijvoorbeeld niet gekwalificeerd om wetgeving te vertalen.Bovendien kan de kwaliteit van vertalingen verschillen van dag tot dag afhankelijk van factoren zoals de motivatie van een vertaler, vermoeidheid, stress, enzovoorts.Ten slotte is er ook nog variatie tussen talen.Veel Romaanse talen staan bijvoorbeeld toe dat voornaamwoorden weg worden gelaten in bepaalde gevallen, terwijl dit niet mag in bijvoorbeeld het Engels.Het Duits en veel Slavische talen maken gebruik van grammaticaal geslacht, wat een concept is dat onbekend (en hoogst verwarrend) is in Engelstalige delen van de wereld.Voor onderzoek naar machinevertaling levert dit de volgende uitdaging op: de data is niet homogeen en een letterlijke vertaling die gepast is in de ene context kan verkeerd zijn in de andere.Bovendien zijn vertalingssystemen vaak getraind op diverse documenten van verschillende bronnen, wat betekent dat ze verschillende taalkundige stijlen tegenkomen.Gebruikers van moderne machinevertalingssystemen verwachten niet alleen een accurate vertaling die alle informatie vii van de oorspronkelijke tekst bevat, maar ze verwachten ook dat de vertaling grammaticaal is en goed klinkt.Daarom heb ik een model gemaakt van de variatie in vertalingsdata, of althans van een deel hiervan.Mijn hoofdstelling is dat dit vertalingen verbetert, omdat het de aanname dat de data homogeen is versoepelt (we weten tenslotte dat deze aanname niet klopt).Ik laat experimenteel zien dat dit inderdaad verbeterde vertalingen oplevert, middels de algemeen gebruikte BLEU I-maatstaf.De dissertatie begint met een korte motiverende introductie in Hoofdstuk 1. Daarna beschrijft het de noodzakelijke wiskundige achtergrond in Hoofdstuk 2. Omdat de probabilistische modellen in deze dissertatie gebruik maken van benaderende inferentietechnieken wordt er een bijzondere nadruk gelegd op deze technieken.Hoofdstuk 2 biedt ook een introductie in frase-gebaseerde en neurale machinevertaling.Hoofdstuk 3 introduceert een nieuw latente-variabele-model om om te gaan met variatie in woorduitlijning.Woorduitlijning is de eerste stap in de frasegebaseerde machinevertalingsprocedure.Het verbindt woorden, tussen twee parallelle zinnen, die waarschijnlijk vertalingen zijn van elkaar.Deze vertalingen op woordniveau worden dan uitgebreid naar frases in een volgende stap, die vervolgens door het vertalingssysteem onthouden worden.Een veelgebruikte aanname van veel woorduitlijningsmodellen is dat ieder woord in een van de talen een tegenhanger moet hebben in de andere taal.Dit is natuurlijk niet waar, omdat talen verschillen in hoe ze concepten uitdrukken.Zoals eerder genoemd laten sommige talen voornaamwoorden weg terwijl andere talen voorzetsels weglaten.De reden dat een voornaamwoord voorkomt in zin A en niet in zin B is daarmee volledig afhankelijk van de grammaticale structuur van taal A en heeft niets te maken met vertaling.Ik introduceer daarom een latente-variabele-model dat een combinatie is van een klassiek uitlijningsmodel-en een taalmodelcomponent.Het taalmodel kan grammaticaal geïntroduceerde woorden verklaren en voorkomt hiermee dat de uitlijningsmodellen een verkeerde uitlijning produceren.Experimenten laten zien dat de resulterende uitlijningen leiden tot verbeterde vertalingen.Het model dat geïntroduceerd wordt in Hoofdstuk 4 benadert variatiefenomenen op een meer holistische manier omdat het ingebed is in een integraal neuraal machinevertalingssysteem.De hypothese die onder dit model ligt is dat de bronnen van variatie in vertaling