A Personal Conversation Assistant Based on Seq2seq with Word2vec Cognitive Map

Maoyuan Shen, Runhe Huang · 2018

WeChat is one of social network applications that connects people widely. Huge data is generated when users conduct conversations, which can be used to enhance their lives. This paper will describe how this data is collected, how to develop a personalized chatbot using personal conversation records. Our system will have a cognitive map based on the word2vec model, which is used to learn and store the relationship of each word that appears in the chatting records. Each word will be mapped to a continuous high dimensional vector space. Then we will adopt the sequence-to-sequence framework (seq2seq) to learn the chatting styles from all pairs of chatting sentences. Meanwhile, we will replace the traditional one-hot embedding layer with our word2vec embedding layer in the seq2seq model. Furthermore, we trained an autoencoder of seq2seq architecture to learn the vector representation of each sentence, then we can evaluate the cosine similarity between model generated response and the pre-existing response in test set, and we can also display the distance with principal component analysis (PCA) projection. As a result, our word2vec embedded seq2seq model significantly outperforms the one-hot embedded one.

Read the paper · More papers on PaperTik