Code Ranking with Structure Awareness Contrastive Learning
Hailin Huang, Liuwen Cao, Jiexin Wang, Tianchen Yu, Yi Cai · 2025
Large language models (LLMs) have revolutionized the field of programming for developers by automatically generating code based on natural language intent (NL intent). In numerous cases, LLMs can produce correct programs after several trials. As a result, a major challenge for this task is to select the most appropriate program from the multiple samples (also called code ranking) generated by LLMs. Recent popular approaches for code ranking involve the ranker-based methods, in which we train a ranker to classify the error in code using execution results (correct or error types) of code as supervised signals select the best program. However, existing rankerbased code ranking approaches rely on classification labels, which are highly sensitive to label distribution and show weak generalization ability to other distributions. In this paper, we introduce SACL-CR to address this challenge, a novel structureaware contrastive learning framework for code ranking. This approach effectively addresses the generalization issues of existing ranker-based methods by integrating both code sequence and structural information. Encoders trained with this method can effectively identify errors in code, enhancing the model's ability to differentiate between correct and incorrect code. Our research demonstrates that SACL-CR significantly enhances the pass@k accuracy of several code generation models, including CodeLlama and DeepseekCoder, on the HumanEval and MBPP datasets. The open-source code will be released at https://github.com/Iced-Americano2001/SACL-CR.