Unlearnable Examples: Protecting Open-Source Software from Unauthorized Neural Code Learning
Zhenlan Ji, Pingchuan Ma, Shuai Wang · Proceedings/Proceedings of the ... International Conference on Software Engineering and Knowledge Engineering · 2022
The vast volume of "free" code maintained on open-source code management systems significantly simplifies the process of producing and sharing open-source software.Recently, we have seen a growing trend in which these open-source software is being used for neural code learning without authorization.Note that open-source software does not necessarily imply "unrestricted usage," e.g., software under the BSD license requires users to retain the copyright notice and credit the software's developers.The unauthorized use of software for (commercial) neural code learning models has raised copyright concerns.This paper, for the first time, provides approaches for protecting opensource software from unauthorized neural code learning via unlearnable examples.Our proposed technique applies a set of lightweight transformations toward a program before it is open-source released.When these transformed programs are used to train models, they mislead the model into learning the unnecessary knowledge of programs, then fail the model to complete original programs.The transformation methods are sophisticatedly designed to ensure that they do not impair the general readability of protected programs, nor do they entail a huge cost.We focus on code autocompletion as a representative downstream task of unauthorized neural code learning.We demonstrate highly encouraging and cost-effective protection against neural code autocompletion.