MODEL SELECTION BASED ON QUASI-LIKELIHOOD WITH APPLICATION TO OVERDISPERSED DATA
Yiping Tang · Journal of the Japanese Society of Computational Statistics · 2013
ABSTRACT Overdispersion is a common phenomenon in discrete data analysis with generalized linear models which can never be ignored. In the literature on model selection with overdispersion, the main stream of the methods is to use the maximum likelihood method with a strong assumption of dening an entire distribution of the data to construct information criteria. For example, Fitzmaurice (1997) proposed model selection methods based on Efron's (1986) double-exponential families. However, the assumption of a parametric distribution of the data is sometimes too strong. In this paper,we propose a new model selection strategy based on the quasi-likelihood functions.First, we dene a statistical model with two moments in the form of mean and variance. Second, we dene the semiparametric information based on the semiparametric statistical models. Finally, we propose the semiparametric information criterion (SIC) based on quasi-likelihood estimators with weak assumptions on the rst two moments.Although the proposed semiparametric information criterion is inspired by data with overdispersion, it applies more generally to data no matter whether they are discrete or continuous. It may be useful for any models for which the quasi-likelihood methods are appropriate. We also demonstrate the use of SIC through simulation studies and data application of the well-known data on germination of Orobanche.