Robust and Undetectable White-Box Watermarks for Deep Neural Networks.
Tianhao Wang, Florian Kerschbaum · 2019
Watermarking of deep neural networks (DNN) can enable their tracing once released by a data owner. In this paper we generalize white-box watermarking algorithms for DNNs, where the data owner needs white-box access to the model to extract the watermark, and attack and defend them using DNNs. White-box watermarking algorithms have the advantage that they do not impact the accuracy of the watermarked model. We demonstrate a new property inference attack using a DNN that can detect watermarking by any existing, manually designed algorithm regardless of training data set and model architecture. We then use a new training architecture and a further DNN to create a new white-box watermarking algorithm that does not impact accuracy, is undetectable and robust against moderate model transformation attacks.