Code Membership Inference for Detecting Unauthorized Data Use in Code Pre-trained Language Models

Sheng Zhang, Hui Li, Rongrong Ji · 2024

Code pre-trained language models (CPLMs) have received great attention since they can benefit various tasks that facilitate software development and maintenance.However, CPLMs are trained on massive open-source code, raising concerns about potential data infringement.This paper launches the study of detecting unauthorized code use in CPLMs, i.e., Code Membership Inference (CMI) task.We design a framework BUZZER for different settings of CMI.BUZZER deploys several inference techniques, including signal extraction from pre-training tasks, hard-to-learn sample calibration and weighted inference, to identify code membership status accurately.Extensive experiments show that CMI can be achieved with high accuracy using BUZZER.Hence, BUZZER can serve as a CMI tool and help protect intellectual property rights.

Read the paper · More papers on PaperTik