On the Price of Source Anonymity in Heterogeneous Parametric Point Estimation
Weining Chen, I-Hsiang Wang · 2019
Parametric point estimation from anonymous and heterogeneous data is studied. For heterogeneity, we assume n samples are independently drawn, each following one of K possible distributions. For anonymity, we assume the estimator knows the number of samples drawn from each distribution, but which one each sample follows is hidden. In words, samples as a sequence are passed through an unknown permutation prior to being observed. The goal is to find an estimator that minimizes the worst-case statistical risk over all possible permutations. We prove that an optimal estimator depends only on the empirical distribution (type) of samples, and when the risk function is the mean squared error (MSE), it follows a non-trivial Cramer-Rao lower bound. We further characterize its asymptote as n → ∞, assuming the number of samples from each distribution is proportional to n. The lower bound is of the order of 1/n, and the reciprocal of its prefactor is the Fisher information of the mixture of the K distributions.