Black-Box Steganography for Large Language Models
Xinxin Li, Zichi Wang, Xinpeng Zhang · IEEE Transactions on Circuits and Systems for Video Technology · 2025
In recent years, the rapid development of deep learning has brought new opportunities for steganography. However, the current advanced white-box model steganography methods are not suitable for large language models. Since the parameter scale and complexity of large language models are far beyond that of ordinary models, retraining them to hide secret data is extremely challenging. Moreover, the cover parameters or structures of the embedded data are vulnerable to detection by attackers. To enhance practicality, we propose a black-box steganographic scheme for large language models, which embeds secret data into the third-party pre-trained large language models using backdoor techniques without knowing the internal complex structure and parameters of the large language models. Specifically, the sender first encodes the secret data into trigger labels and then uses a certain proportion of trigger samples and clean samples to fine-tune the third-party large language model to embed the secret data without significantly reducing the model performance. The receiver uses trigger samples to extract the secret data by interacting with the large language model, thereby achieving covert communication of the secret data. Experiments demonstrate the effectiveness of the proposed scheme in terms of embedding capacity, robustness, and security.