Automated Code Summarization by Training Large Language Models with Crowdsourced Knowledge
Meng Xia, Shradha Maharjan, Myoungkyu Song · 2025
In modern software development, efficient program comprehension is essential for maintaining and evolving software systems. Developers often dedicate over $50 \%$ of their time to understanding code due to the complexity and time demands involved. Code summarization-generating concise natural language descriptions of source code-has emerged as a potential solution. However, existing automated summarization techniques frequently produce summaries that are incomplete or lack accuracy. Moreover, as software evolves, documentation often becomes outdated, leading to discrepancies between code and comments. To address these challenges, we present DeepKnowCode, an automated approach that utilizes a DEEP learning technique based on a large language model and crowdsourced KNOWledge for CODE summarization. This approach aids developers by generating summaries that elucidate (1) the internal behavior of the code, (2) the rationale behind its implementation, and (3) practical guidelines for its use. We implemented a research prototype to assess real-world applicability and rigorously evaluated DeepKnowCode against state-of-the-art approaches. In our evaluation, DeepKnowCode demonstrates performance improvements in BLEU scores of 38.8% and 39.2% over two baselines, with statistical validation underscoring its effectiveness. By incorporating crowdsourced knowledge, DeepKnowCode captures essential elements of code semantics, context, and patterns, enhancing its ability to produce accurate and contextually relevant code summaries that facilitate program comprehension.