Combining Modern Machine Translation Software with LSI for Cross-Lingual Information Processing
Roger B. Bradford, John Pozniak · 2014
The growing internationalization of business and social interactions poses significant challenges in implementing multilingual information systems. For applications requiring retrieval, clustering, and categorization of multilingual document collections, cross-lingual application of latent semantic indexing (LSI) has a number of characteristics that make it potentially attractive. However, this technique is dependent upon the availability of applicable parallel corpora. Historically, such corpora have been quite limited in size and scope. In this paper, we provide new results regarding implementation of cross-lingual LSI text processing systems employing parallel corpora produced using modern machine translation (MT) products. We present measurements using the Reuters 21578 test set to demonstrate three key points regarding this combined LSI/modern MT approach: (1) for some languages, this approach can create parallel corpora of sufficient fidelity to support effective multilingual and cross-lingual LSI applications, (2) the technique is not particularly sensitive to details of LSI parameters, and (3) multiple languages can be represented in a single LSI space with little degradation in performance.