Explaining Internal Representations in Deep Networks: Adversarial Vulnerability of Image Classifiers and Learning Sequential Tasks with Sparse Reward

Igor Farkaš · 2025

Modern artificial intelligence based on deep neural networks has demonstrated in the past decade great achievements in concrete tasks, sometimes even surpassing human performance. On the other hand, there exist fundamental problems in these models, be it image classification or natural language tasks. Since deep networks are inherently black boxes, it is important to design and apply techniques that help shed light on the functioning of the trained models. In the talk, we will discuss two domains. First, in the context of image classification we will illustrate the effect of adversarial examples that can easily fool trained models, hence revealing their lack of robustness. In the second part, we will deal with sequential tasks, such as computer games, with an extremely sparse reward. Introducing the concept of intrinsic motivation, we will describe the neural networks based model that can successfully use reinforcement learning and self-supervised knowledge distillation to solve these tasks thanks to optimized organization of its internal representations. Finally, we briefly mention our current work related to cognitive robotics.

Read the paper · More papers on PaperTik