Cross-view Embeddings for Information Retrieval
Parth Gupta · 2017
In this dissertation, we deal with the cross-view tasks related to information retrieval using embedding methods.We study existing methodologies and propose new methods to overcome their limitations.We formally introduce the concept of mixed-script IR, which deals with the challenges faced by an IR system when a language is written in different scripts because of various technological and sociological factors.Mixedscript terms are represented by a small and finite feature space comprised of character n-grams.We propose the cross-view autoencoder (CAE) to model such terms in an abstract space and CAE provides the state-of-the-art performance.We study a wide variety of models for cross-language information retrieval (CLIR) and propose a model based on compositional neural networks (XCNN) which overcomes the limitations of the existing methods and achieves the best results for many CLIR tasks such as ad-hoc retrieval, parallel sentence retrieval and cross-language plagiarism detection.We empirically test the proposed models for these tasks on publicly available datasets and present the results with analyses.In this dissertation, we also explore an effective method to incorporate contextual similarity for lexical selection in machine translation.Concretely, we investigate a feature based on context available in source sentence calculated using deep autoencoders.The proposed feature exhibits statistically significant improvements over the strong baselines for English-to-Spanish and English-to-Hindi translation tasks.