Equivariant deep generative models

Bahare Azari · 2022

Deep generative models are powerful tools for learning and interpreting complex real-world data. Generative models are often equipped with latent variables that allow them to learn low-dimensional representations to find the most simplified and compressed description of data, which can then be used for different tasks such as predictions. Among various techniques for constructing and training generative models, variational inference methods are of interest due to their ability to provide fast, scalable, and accurate inference. Models and learning tasks often have inherent symmetry, that is, input transformations leave the output unchanged or the output undergoes similar transformations. However, in existing generative models, symmetries of the problem are not taken into account during the learning process. Consequently, the learned representations for individually transformed inputs may not be meaningfully related. In the first part of this dissertation, we incorporate these symmetries into the designed generative models and deep architectures to achieve enhanced generalization and guaranteed performance under input transformations. We do so by developing appropriate invariant/equivariant neural network structures for different models and tasks. Specifically, we propose a Circular-symmetric Correlation Layer (CCL) based on the formalism of roto-translation equivariant correlation on the continuous group S1 x R, and implement it efficiently using the well-known Fast Fourier Transform (FFT) algorithm. We showcase the performance analysis of a general network equipped with CCL on various recognition and classification tasks and datasets. We then propose an SO(3) equivariant deep dynamical model (EqDDM) for motion prediction that learns a structured representation of the input space in the sense that the embedding varies with symmetry transformations. EqDDM is equipped with equivariant networks to parameterize the state-space emission and transition models. We demonstrate the superior predictive performance of the proposed model on various motion data. In the second part of this dissertation, we investigate the application of generative models in the fields of wireless communications for channel estimation, equalization, and packet decoding and computational neuroscience for emotion categorization. We show how these applications benefit from these modeling frameworks in terms of decoding accuracy and interpretability.--Author's abstract

Read the paper · More papers on PaperTik