Pierre-Aurelien Gilliot, Christophe Andrieu, Anthony Lee, Song Liu, and Michael Whitehouse’s contribution to the Discussion of ‘the Discussion Meeting on Probabilistic and statistical aspects of machine learning’
Pierre-Aurélien Gilliot, Christophe Andrieu, Anthony Lee, Song Liu, Michael Richard Whitehouse · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2023
We would like to thank the authors for their interesting contribution and the clarification brought on a particular aspect of diffusion models concerned with the criterion used in the fitting phase of the procedure. Another aspect, not discussed extensively, is the choice of the neural network used to model the gradients involved in the generative phase. In particular, the implicit prior information this induces on the class of distributions considered may seem mysterious and perhaps not always controllable? Here, we report results from a toy numerical experiment, which seem to raise some questions. We consider 28×28 images, in which all but four pixels have values distributed according to a Beta(10,10), re-scaled between 0 and 0.49. The four remaining pixels have values distributed according to a Beta(10,10), re-scaled between 0.51 and 1, and are organised in two distinct configurations: square or diamond (Figure 1). Squares are uniformly distributed across the whole image, while diamonds are uniformly distributed only in the bottom right quadrant. Four different datasets were created, each comprising 82,000 of such images, with varying proportions of squares and diamonds (Figure 2), and were each used to train four different U-nets approximating the score function following standard practice (Ho et al., 2020; HuggingFace, 2023). Configurations in the training dataset: a) squares and b) diamonds. The grey shading represents the bottom right quadrant. Images are binarised using a threshold set at 0.5 for easier visualisation. Frequency of diamonds. The grey bar represents the frequency of diamonds generated inside the bottom right quadrant while white bar indicates the frequency of diamonds generated outside of it. A generative model which is coherent with this dataset should be expected to generate images with squares or diamonds, respecting pixel boundaries for the diamonds, and matching the frequencies of each configuration. Remarkably, most sampled images ( 94%) correctly consist of either squares or diamonds. Diamonds were not confined to the bottom right quadrant anymore (Figure 2), which could stem from the shift-equivariant property of the U-Net’s convolutional layers (Cohen & Welling, 2016). However, this may not be a desirable feature. Diffusion modelling aims at alleviating the challenges faced by score matching when dealing with disjointed, multimodal distributions such as the mixture square/diamond considered here (Song & Ermon, 2019). However, we encountered difficulties in accurately sampling according to the training set mixture weights: the sampled diamond frequencies did not reflect the training set diamond frequencies across the various datasets (Figure 2). Interestingly, these sampled frequencies differed when using another seed to initialise the U-Net weights (but were still unexpected). While the link between simple properties of the neural network chosen and the class of distributions implicitly defined can sometimes be understood (Yim et al., 2023), a general characterisation is missing from the literature. This phenomenon is further compounded by additional unknown neural network properties. We are curious about the current understanding of this matter.