Wasserstein GANs are Minimax Optimal Distribution Estimators
Arthur Stéphanovitch, Eddie Aamari, Clément Levrard · arXiv (Cornell University) · 2023
We provide non asymptotic rates of convergence of the Wasserstein Generative Adversarial networks (WGAN) estimator. We build neural networks classes representing the generators and discriminators which yield a GAN that achieves the minimax optimal rate for estimating a certain probability measure $μ$ with support in $\mathbb{R}^p$. The probability $μ$ is considered to be the push forward of the Lebesgue measure on the $d$-dimensional torus $\mathbb{T}^d$ by a map $g^\star:\mathbb{T}^d\rightarrow \mathbb{R}^p$ of smoothness $β+1$. Measuring the error with the $γ$-Hölder Integral Probability Metric (IPM), we obtain up to logarithmic factors, the minimax optimal rate $O(n^{-\frac{β+γ}{2β+d}}\vee n^{-\frac{1}{2}})$ where $n$ is the sample size, $β$ determines the smoothness of the target measure $μ$, $γ$ is the smoothness of the IPM ($γ=1$ is the Wasserstein case) and $d\leq p$ is the intrinsic dimension of $μ$. In the process, we derive a sharp interpolation inequality between Hölder IPMs. This novel result of theory of functions spaces generalizes classical interpolation inequalities to the case where the measures involved have densities on different manifolds.