TokenFree: A Tokenization-Free Generative Linguistic Steganographic Approach with Enhanced Imperceptibility
Ruiyi Yan, Tian Song, Yating Yang · 2024
Since tokenization serves a fundamental preprocessing step in numerous language models, tokens naturally constitute the basic embedding units for generative linguistic steganography. However, tokenization-based methods face challenges including limited embedding capacity and possible segmentation ambiguity. Despite existing character-level (one tokenization-free type) linguistic steganographic approaches, they face the problem of generating unknown or out-of-vocabulary words, potentially compromising steganographic imperceptibility. In this paper, we focus on both embedding capacity and imperceptibility of tokenization-free linguistic steganography. First, we suggest that unknown words mainly result from low-entropy distributions and rigid coding rules used in candidate pools, thus we propose an entropy-based selection approach to flexibly construct candidate pools. Further, we present a lexical emphasis approach, prioritizing characters within candidate pools capable of forming in-vocabulary words. Experiments show that, across a range of high embedding rates, our approaches achieve considerably higher imperceptibility and text fluency, increase anti-steganalysis capacity averagely by 14.4%, and particularly reduce out-of-vocabulary rate averagely by 88.7%, compared to the existing state-of-the-art character-level steganographic methods.