Query Answering for Kisan Call Centerwith LDA/LSI
Sandeep Kumar Mohapatra, Anamika Upadhyay · 2018 International Conference on Advances in Computing, Communication Control and Networking (ICACCCN) · 2018
With the enhanced storage and processing capacity of the computers in today's world we have lots of unprocessed data in every domain. This data has gained interest of data scientist of all over the world to get some meaningful information from the raw data. We are here processing kisan call center data. First the preprocessing of data is done which is then converted to a similarity matrix which we save in a database using mongoDb. This prepared data will now be fed to the gensim TI-IDF model for text transformation. After this we have used another two methods of topic modelling to be used with TI-IDF model. We are concerned in getting the query answers efficiently from the data collected by usingLatent Dirichlet allocation (LDA) and Latent Semantic Indexing (LSI) in pipeline to the TI-IDF model. One algorithm which will be used with pipeline to the TI-IDF is LDA. LDA is a technique that automatically discovers topics that these documents contain. Another model which we are exploring is the LSI which also will be used with TI-IDF model for preparing the model for getting query answers efficiently. On these parameter we will be building the model training and finally will compare the models for finding which model works best in terms of efficiency, accuracy and complexity.