Exploring Large Language Models for Method Name Prediction
Hanwei Qian, Tingting Xu, Ziqi Ding, Wei Liu, Shaomin Zhu · 2024
High-quality method names are crucial for program comprehension and maintenance. However, method naming is a challenging task, especially for inexperienced developers. To alleviate this difficulty, various deep-learning techniques have been proposed to predict appropriate names for given method code bodies. The recent advent of Large Language Models (LLMs) has showcased outstanding performance in Natural Language Processing (NLP), leaving a profound impression on the software engineering (SE) community. LLMs have been applied to several SE tasks, e.g., code summarization, demonstrating their significant potential. However, there has not yet been a systematic and comprehensive investigation into how to adapt LLMs to method name prediction (MNP) tasks and their performance on such tasks. To fill this gap, in this paper, we conduct the first empirical study to understand the capabilities of LLMs in MNP tasks. Firstly, we examine LLMs’ comprehension abilities and interaction patterns in the MNP task, testing various prompts and temperature parameters. Results indicate that LLMs excel under 0 temperature and a few-shot prompt template. Following this, we compare the performance of prompt-guided LLMs and fine-tuned PLMs, with GPT-3.5 performing best among LLMs and CodeT5 leading among PLMs, with minimal performance disparity. SentenceBert with Cosine Similarity (SBCS) has been introduced to quantify the semantic differences between predicted and ground-truth names. Lastly, the manual review highlights human evaluations outperforming F1metrics for LLMs, with issues such as global information lack and dataset quality affecting performance identified through case analysis.