CSFIR: Leveraging Code-Specific Features to Augment Information Retrieval in Low-Resource Code Datasets
Zhenyu Tong, Chenxi Luo, Tiejian Luo · 2024
From search engines like Google to advanced applications such as Retrieval Augmented Generation integrating Large Language Model (LLM), Information Retrieval (IR) serves a crucial role. To facilitate the development of increasingly large and complex code program projects, researchers introduce IR systems into the code domain. Unfortunately, although IR systems achieve significant success in retrieving natural language corpus and query, they face challenges when tasked with retrieving corpus consisting of code sequences. Primarily, in practical applications, most code sequences lack corresponding natural language annotations, which are known as low-resource scenarios, hindering the training of neural network-based retriever in IR systems. Furthermore, modern IR systems often overlook the structural features of code sequences, which may be beneficial for understanding these sequences. Additionally, the length of most code sequences exceeds that of equivalent natural language expressions, complicating the processing of relationships between code and natural language sequences. To address these challenges, we propose a novel IR system CSFIR, which leverages Code-Speclflc Features to augment IR. For the prevalent issue of unlabeled code sequences in low-resource scenarios, we employ a supervised fine-tuned LLM as a generator to generate natural language queries for unlabeled code sequence. Subsequently, we extract structural features from the abstract syntax tree of code sequence using graph convolution networks and integrate these features to enhance the original retriever. Finally, given the adverse effects of lengthy code sequences on generators, we propose a subtree segmentation algorithm, which reduces the length of code sequences without compromising their original meaning, thereby enhancing the quality of queries generated by the generator. We conduct comparative experiments to ascertain the efficacy of our method. Regarding Recall@100, our CSFIR system improves from 90.12 in a traditional IR system to 96.18. Our code is available at https://github.com/tzy3141S/CSFIR.git.