Generative Indonesian Conversation Model using Recurrent Neural Network with Attention Mechanism
Andry Chowanda, Alan Darmasaputra Chowanda · Procedia Computer Science · 2018
This research contributes to data collection for Indonesian Language conversation corpus or dataset collected from conversation in the movies through 1961 subtitles with 1678320 unique words. The dataset collected thus trained with deep learning algorithm, a dual encoder Recurrent Neural Network (RNN), Long-Short Term Memory (LSTM) with Attention mechanism with the best result was 2.37% unknown words, loss rate was 1.66, and perplexity of 4.96, trained with vocabulary size of 24000. The dataset could be used for many machine (deep) learning for Natural Language Processing Problem in Indonesian Language, while the pre-trained model can be integrated into several system such as agent-based system, be they a chat-bot or an Embodied Conversational Agent (ECA). For future work, more data will be collected not only from the movie conversation but also natural human-human conversation in Indonesian Language.