An Effective Approach of Extracting Local Documents from the Distributed Representation of Text using Document Embedding and Latent Semantic Analysis

Vikas Chib, Ahsan Jafri · 2019 International Conference on Smart Systems and Inventive Technology (ICSSIT) · 2019

Document retrieval is the process of extracting the relevant documents on the basis of the defined query. The main problem is on finding the related document based on the local representation of the query. Here local representation means, for e.g. if you want to search for documents related to “bank”. “Bank” will have two representation “bank of a river” (distributed representation) or “Saving Bank” (local representation). But we need documents related to “Saving Bank” which belongs to the local representation of the query “bank”. This paper proposes a new model named Latent Semantic Analysis (LSA) and Document Embedding's to find the relevant documents. This is the initial attempt to combine the document embedding's vectors with LSA All neural embedding models learn distributed representation of text and match the results in the latent semantic space on a given query, but searching documents from the distributed representation will lose the relevance of local representation of a given query. We propose a novel information retrieval system, which uses doc2vec model to give top N similar documents with a relevant ranking using Latent Semantic Indexing to give the top K (documents score is greater than a soft threshold) documents which are the local representation of given query. We can use these K documents to find the most similar ones. We can show that this `dual' combination performs better than other traditional information retrieval algorithm or recently developed neural network models.

Read the paper · More papers on PaperTik