Exploring the Effectiveness of Joint Bi-LSTM-GNN vs ChatGPT-LLM in Source Code Retrieval

Nazia Bibi, Ayesha Maqbool, Tauseef Ahmad Rana · 2023

The ability to search for and retrieve important code snippets from huge codebases is a critical task in software engineering. Traditional approaches for retrieving source code rely on keyword-based queries, which might produce erroneous and partial results. As a result, scientists are investigating the use of deep learning approaches to enhance code search. In this study, a Joint Bi-LSTM-GNN (JBLG) model is proposed for source code search. This model combines the features of BiLSTM and GNN. This study also investigates the usage of Large Language Models (LLMs), notably ChatGPT, for source code retrieval in comparison with our proposed JBLG model for source code search. Our results reveal that our proposed JBLG model surpasses standard retrieval approaches and other deep learning models. The JBLG model is tested on CodeSearchNet dataset, which comprises query code pairs from open-source projects in a variety of programming languages. Our joint model obtains an impressive mean reciprocal rank (MRR) score, which represents a considerable improvement over the best-performing baseline model. Overall, our findings show that the ChatGPT model does not perform, and this might be due to ChatGPT being a language model developed for natural language processing tasks rather than code retrieval. To further enhance the effectiveness of the model, future approaches for this study include examining the usage of attention processes and more sophisticated deep learning techniques. The suggested approach might also be expanded to include more programming languages and software engineering functions including code summary and code completion.

Read the paper · More papers on PaperTik