Analyzing Latent Entropy in Deep Q-Learning

Jacob E. Kooi, Mark Hoogendoorn, Vincent François-Lavet · 2025

Algorithms based on deep learning and the Bellman iteration serve as a basis for most state-of-the-art approaches in the field of reinforcement learning. Deep Q-learning stands out as the predominant example. In scenarios involving high-dimensional input data, like pixel observations, the architecture of a deep Q-Network typically features a sequence of convolutional layers succeeded by a set of linear layers. In this setting, the activations in the network’s final hidden layer can be seen as the latent representation, encompassing all the compressed information. We show that this learned representation is prone to saturation or contraction, leading to vanishing gradients, a reduction of information content and sub-optimal convergence. In addition, the temporal evolution of the latent representation in RL is analyzed by characterizing its entropy. Finally, a set of methods is proposed to alter the latent representation during learning by influencing its entropy. Three entropy-enhancing techniques are compared, which show a strong empirical relation between representation entropy and downstream performance.

Read the paper · More papers on PaperTik