MMCS: A Code Summarization Approach Based on Multi-Modal Feature Enhancement
Zhaolei Zhou, Lei Liu · 2024
In the process of program development and maintenance, with the continuous upgrading and iteration of the program, the code structure becomes more and more complex, and the code comments are frequently out-of-date, inconsistent, or completely missing, which significantly lowers the productivity of software development and maintenance by requiring developers to spend a lot of time understanding the code. Therefore, many studies have tried to introduce deep learning techniques into automatic code summarization, all of which have achieved good results, but there are still some shortcomings that need to be improved, e.g., some studies directly input Abstract Syntax Tree (AST) sequences into the encoder, which results in insufficient structural information extraction; some studies perform graph embedding of ASTs or directly extract the control-flow graph of the code for graph embedding, which can only extract the AST's local structural information, and the semantic information extraction is insufficient. In this paper, we address the above problems and propose a code summarization approach MMCS based on multi-modal feature enhancement. MMCS firstly introduces four types of extended edges to transform AST, then designs multi-modal encoder, multi-modal fusion module, and finally generates the summarization by using the Transformer decoder. MMCS is conducted on Java and dataset and Python dataset with four baseline comparison experiments. The experimental results show that MMCS outperforms CODE-NN, Tree2Seq, Hybrid-DeepCom and MMTrans baselines on BLEU-4, METEOR and ROUGE, and MMCS improves 0.63%, 0.69% and 0.51 % on three evaluation metrics, respectively, compared to MMTrans.