Efficiency of maximum likelihood estimation for a multinomial distribution with known probability sums
Yo Sheena · Information Geometry · 2025
For a multinomial distribution, suppose we have prior knowledge about the sum of the probabilities of certain categories. This enables the construction of a submodel within the full (i.e., unrestricted) model. The maximum likelihood estimator (MLE) under such a submodel is generally expected to achieve higher estimation efficiency than the MLE under the full model. However, this article shows that this expectation does not always hold. We derive the asymptotic expansion of the risk of the MLE, with respect to Kullback–Leibler divergence, for both the full model and the “m-aggregation” submodel. The results indicate that the second-order term (the order $$n^{-2}$$ term) of the submodel is larger than that of the full model unless the submodel is based on “solid” prior knowledge. From this theoretical finding, we conjecture that in some cases the submodel may yield a higher risk than the full model. We confirm this conjecture by presenting an explicit example in which the use of the submodel increases the risk.