Human perceiving behavior modeling in evaluation of code generation models
Sergey Valerevich Kovalchuk, Vadim Lomshakov, Artem Aliev · 2022
In this study, we evaluated a series of code generation models based on CodeGen and GPTNeo to compare the metric-based performance and human evaluation.For a deeper analysis of human perceiving within the evaluation procedure, we implemented a 5-level Likert scale assessment of the model output using a perceiving model based on the Theory of Planned Behavior (TPB).Through this analysis, we demonstrated an extension of model assessment as well as a deeper understanding of the quality and applicability of generated code for answering practical questions.The approach was evaluated with several model settings in order to assess diversity in the quality and style of answer.With the TPBbased model, we showed a different level of perceiving of the model result, namely, personal understanding, agreement level, and readiness to use the particular code.With this analysis, we investigate a series of issues in code generation, namely, natural language generation (NLG) problems observed in the context of programming and question-answering with code. 1 https://stackoverflow.com/code generation evaluation were developed, such 40 as CodeBLEU (Ren et al., 2020), RUBY (Tran et 41 al., 2019), and others.Still, recent studies 42 (Evtikhiev et al., 2022) show that the direct 43 application of metrics often leads to issues in code 44 generation evaluation.45 With this in mind, investigated the applicability 46 of human evaluation widely spread in NLG 47 problems (De Mattei et al., 2021; Hämäläinen & 48 Alnajjar, 2021) to assess an alternative approach to 49 code evaluation and a deeper understanding of 50 human perceiving of code generation (and NLG 51 output in general).The study poses two research 52 questions.First, how is human perceiving reflected 53 by the text-or code-oriented NLP metrics?Second, 54 what is the structure of human perceiving in the 55 human evaluation procedure in the question-56 answering scenario?Here we consider perceiving 57 as an act of becoming subjectively aware and 58 conscious of the observed information.59 The structure of the paper is as follows.The next 60 section describes the datasets used in the study.The 61 following section presets the details of code 62 generation models' selection and preparation.63 Section 4 describes human evaluation solutions 64 and procedures.Section 5 discusses the study 65 results and evaluation results.Finally, Sections 6 66 and 7 provide a discussion and concluding remarks 67 respectively.68 2 Dataset 69 Within the study, we focused on question 70 answering (QA) with a generation of short snippets 71 as answers to real-world problems such as 72 questions asked in Stack Overflow 1 (SO).For 73 consistency, we added the following restrictions to 74 the questions and answers considered within the 75 study, taking SO as a reference for the analysis.