Lexical disambiguation in machine translation with latent semantic analysis
Elizabeth E. Davis, Simon D. Levy · Journal of computing sciences in colleges · 2007
This undergraduate honors thesis focuses on the problem that machine translators face in choosing the correct translation of a polysemous English noun in a foreign language. For instance, general-purpose internet translators such as Google Translate cannot distinguish between the meanings of the English noun bat, generating the French word for the baseball implement rather than that for the flying mammal - even when provided with such contextual hints as wings and cave. Latent Semantic Analysis (LSA; Foltz & Laham 1998) offers a solution for this issue. LSA is a well-developed technique and theory for relating words and meanings by analyzing text corpora. A multidimensional vector represents each word and each sentence or other contextual block, and similarity or disparity of meaning can then be calculated by the relative angles of these vectors. LSA uses singular value decomposition (Golub & vanLohn 1996) on matrices constructed from words and passages in the learning texts to reduce the number of dimensions in these vectors. Such a dimensional reduction has been shown to be capable of producing simulations of human contextual associations significantly better than simple proximity frequencies.