Monolingual and bilingual concept visualization from corpora
Dominic Widdows, Scott Cederberg · 2003
uce a meaningful diagram of results related to a particular word or query, we perform two extra steps. Firstly, we restrict attention to a given number of closely related words (determined by cosine similarity of word vectors), selecting a local group of up to 100 words and their word vectors for deeper analysis. A second round of Latent Semantic Analysis is then performed on this restricted set, giving the most significant directions to describe this local information. The 2 most significant axes determine the plane which best represents the data. (This process can be regarded as a higher-dimensional analogue of finding the line of best-fit for a normal 2-dimensional graph.) The resulting diagrams give a summary of the areas of meaning in which a word is actually used in a particular document collection. This is particularly effective for visualizing words in more than one language. This can be achieved by building a single latent semantic vector space incorporating words from two la