HIFI-Stego: A High-Fidelity Embedding Audio Steganography Based on Audio Features Decoupling
Sanfeng Zhang, Baiyu Tian, Yang Gao, Xingyu Liu, Wang Yang · IEEE Transactions on Audio Speech and Language Processing · 2025
The higher audio quality of steganography is directly correlated with the increased resistance to steganalysis tools. The advancement of generative AI technologies, particularly those that decouple style features from content, has shown promising developments by facilitating the creation of superior media content. This paper introduces the concept of audio decoupling and presents HIFI-Stego, an embedding audio steganography technique that aims to improve security while maintaining elevated stego audio quality. HIFI-Stego comprises a generator based on the encoder-decoder architecture and a secret message extractor. The encoder of the generator decouples the original audio, yielding the content vector, while the vocoder WORLD is employed to extract the style vector to preserve high-quality audio related features such as timbre and tone. Subsequently, the decoder embeds the secret vector into the decoupled content vector and then couples it with the style vector to generate high-fidelity stego audio. As embedding is not done in traditional time domain or frequency domain, existing analysis tools targeted at traditional steganographies fail to effectively detect the presence of the hidden message. The secret message extractor reuses the generator encoder and augments it with a single-layer convolutional neural network, resulting in a simplified structure suitable for lightweight deployment. Experimental results demonstrate that HIFI-Stego outperforms traditional generative and embedding steganographies in terms of audio quality, steganographic capacity, anti-analysis ability, and concealment.