Learning to Attend with Neural Networks
Jimmy Lei Ba · TSpace (University of Toronto) · 2020
As more computational resources become widely available, artificial intelligence and machine learning researchers design ever larger and more complicated neural networks to learn from millions of data points. Although the traditional convolutional neural networks (CNNs) can achieve superhuman accuracy in object recognition tasks, they brute-force the problem by scanning over every location in the input images with the same fidelity. This thesis introduces a new class of neural networks inspired by the human visual system. Unlike CNNs that process the entire image at once into the current hidden layer, attention allows for salient features to dynamically come to the forefront as needed. The ability to attend is especially important when there is a lot of clutter in a scene. However, learning attention-based neural networks poses some challenges to the current machine learning techniques: What information should the neural network ``pay attention''? Where does the network store its sequences of ``glimpses''? Can our learning algorithms do better than simply ``trial-and-error''? To address these computational questions, we first describe a novel recurrent visual attention model in the context of variational inference. Because the standard REINFORCE or the trial-and-error algorithm can be slow due to its high variance gradient estimates, we show a re-weighted wake-sleep objective can improve the training performance. We also demonstrate the visual attention models outperform the previous state-of-the-art methods based on CNNs in the images and captions generation tasks. Furthermore, we discuss how the visual attention mechanism can improve the working memory of recurrent neural networks (RNNs) through a novel form of self-attention. The second half of the thesis focuses on gradient-based learning algorithms. We developed a new first-order optimization algorithm to overcome the slow convergence of the stochastic gradient descent algorithms in RNNs and attention-based models. In the end, we explored the benefit of applying second-order optimization methods in training neural networks.