Improve Language Modelling for Code Completion by Tree Language Model with Tree Encoding of Context (S)

Yixiao Yang, Xiang Chen · Proceedings/Proceedings of the ... International Conference on Software Engineering and Knowledge Engineering · 2019

In last few years, using a language model such as LSTM to train code token sequences is the state-of-art to get a code generation model.However, source code can be viewed not only as a token sequence but also as a syntax tree.Treating all source code tokens equally will lose valuable structural information.Recently, in code synthesis tasks, tree models such as Seq2Tree and Tree2Tree have been proposed to generate code and those models perform better than LSTMbased seq2seq methods.In those models, encoding model encodes user-provided information such as the description of the code, and decoding model decodes code based on the encoding results of user-provided information.When applying decoding model to decode code, current models pay little attention to the context of the already decoded code.According to experiments, using tree models to encode the already decoded code and predicting next code based on tree representations of the already decoded code can improve the decoding performance.Thus, in this paper, we propose a novel tree language model (TLM) which predicts code based on a novel tree encoding of the already decoded code (context).The experiments indicate that the proposed method outperforms state-of-arts in code completion.

Read the paper · More papers on PaperTik