Abstract
Sverchkov and Pfeffermann[16] consider Small Area Estimation (SAE) under informative probability sampling of areas and within the sampled areas, and not missing at random (NMAR) nonresponse. To account for the nonresponse, the authors assume a given response model, which contains the outcome values as one of the covariates and estimate the corresponding response probabilities by application of the Missing Information Principle, which consists of defining the likelihood as if there was complete response and then integrating out the unobserved outcomes from the likelihood by employing the relationship between the distributions of the observed and the missing data. A key condition for the success of this approach is the ‘correct’ specification of the response model. In this article, we consider the likelihood ratio test and information criteria based on the appropriate likelihood and show how they can be used for the selection of the response model. We illustrate the approach by a small simulation study.
Keywords
Introduction
There exists almost no survey without nonresponse, but in practice most methods that deal with this problem assume either explicitly or implicitly that the missing data are ‘missing at random’ (MAR). However, in many practical situations, this assumption is not valid, since the probability to respond often depends on the outcome value, even after conditioning on available covariate information. In such cases, the use of methods that assume that the nonresponse is MAR can lead to large bias of parameter estimators and distort subsequent inference.
The case where the missing data are not MAR (NMAR) can be treated by postulating a parametric model for the distribution of the outcomes before non-response and a model for the response mechanism. These two models define a parametric model for the observed outcomes, so that the parameters of these models can be estimated from the observed data. See, for example, Pfeffermann and Sverchkov[7] for details, with overview of related literature.
Modeling the distribution of the outcomes before non-response can be problematic since only the observed data are available. Sverchkov[12] proposes an alternative approach, which allows to estimate the parameters of the response model without postulating a parametric model for the distribution of the outcomes before nonresponse. To account for the nonresponse, Sverchkov[12] assumes a given response model and estimates the corresponding response probabilities by application of the missing information principle (MIP), which consists of defining the likelihood as if there was complete response, and then integrating out the unobserved outcomes from the likelihood, employing the relationship between the distributions of the observed and unobserved data. Sverchkov and Pfeffermann[16] apply this approach for small area estimation (SAE) under informative probability sampling of areas and within the sampled areas, and NMAR nonresponse. We describe the main steps of this approach in Sections 2 and 3.
A key condition for the success of this approach is the ‘correct’ specification of the response model. In section 4 we consider the likelihood ratio test and information criteria based on the appropriate likelihood and show how they can be used for the selection of the response model. Section 5 illustrates the application of the approach by a small simulation study.
Notation and Models
Let
where
In practice, not every unit in the sample responds. Define the response indicator;
Define,
The model (2.2) is again general and all that we state at this stage is that under informative sampling and/or NMAR nonresponse, the population and the respondents’ models differ;
Remark 1. The respondents’ model refers to the observed data and hence can be estimated and tested by standard SAE methods. See Pfeffermann[3] and Rao and Molina[9] for estimation and testing procedures in SAE, with references.
Let
See Sverchkov and Pfeffermann,[16] and Pfeffermann and Sverchkov[8] for details.
Unlike the sampling probabilities, the response probabilities are generally unknown. We assume therefore a parametric model, which is allowed to depend on the outcome and the covariate values;
Under these assumptions, if the missing outcome values were observed, γ could be estimated by solving the likelihood equations:
In practice, the missing data are unobserved and hence the likelihood equations (3.1) are not operational. However, one may apply in this case the missing information principle:
Missing Information Principle (MIP, Cepillini et al.[1], Orchard and Woodbury[2]): Let
See Sverchkov[12] and Sverchkov and Pfeffermann[16] for derivation of (3.2). In these equations, EU,Es,Ere define respectively expectations with respect to the population distribution, the sample distribution and the respondents’ distribution. Notice that the internal expectations in the last expression are with respect to the model holding for the observed data for the respondents.
Remark 2. When the response probabilities
where
Sverchkov and Pfeffermann[16] propose to solve the equations (3.2) by maximizing the log-likelihood leading to them, that is, maximizing,
We distinguish between γ* and γ because by (3.2), the derivatives should only be taken with respect to γ.
We maximize the likelihood (3.4) by replacing ui by ûi, obtained by fitting a model of the form (2.2), and dropping the external expectation-Es. The maximization is carried out iteratively, by maximizing in the (q+1) iteration the expression,
with respect to γ(q+1). The maximization can be carried out, for example, by SAS Proc NLIN. See Sverchkov[14] and the examples following Remark 3 for details.
Remark 3. A fundamental question regarding the solution of the MIP equations is the existence of a unique solution or more generally, the identifiability of the response model. Riddles et al.[10] propose a similar approach to deal with NMAR nonresponse in the general context of survey sampling inference and establish the following fundamental condition for the response model identifiability: the covariates
Clearly, the larger are the absolute values of the random effects, the more they affect the values of the outcome values and hence also the values of the response probabilities. In the simulation study of Sverchkov and Pfeffermann,[16] the authors study the effect of the magnitude of the variance of the random effects on the prediction of the area means. The conclusions from that study is that although the estimators of response model parameters become biased as the variance of the random effects increases, the biases are relatively very small and so are the standard deviations of the estimators. Furthermore, increasing the variance of the random effects has negligible effect on the estimation of the true response probabilities and the predictors of the true small area means remain virtually unbiased in each of the areas.
Riddles et al.[10] prove asymptotic normality of the estimator
Suppose that the model fitted to the observed data of the respondents is the mixed generalized logistic model,
Consider a generic response model,
The components of (3.2) can be written in this case as,
The random effects ui and the logistic probabilities py(xij,u) can be estimated by use of the SAS procedure PROC NLMIX.
In Example 1, the outcomes follow a discrete distribution. In this section, we consider continuous outcomes. The proposed algorithm consists of two parts:
where
Substitute (3.8) and (3.9) into (3.2) and estimate γ by iteratively maximizing (3.5).
There is no direct way to test the appropriateness of a chosen response model because the outcome values, which are part of the model, are unknown for the nonresponding units. If the model for the outcomes before nonresponse was known, one could derive the model holding for the observed outcomes based on this model and the model assumed for the responding units, and test the resulting model by use of standard tests that compare the cumulative hypothesized distribution of the observed data with the corresponding empirical distribution, and/or by testing moments of the assumed model. See, for example, Pfeffermann and Landsman[4] and Pfeffermann and Sikov.[5] However, in the approach described in Section 3, we start with a model fitted to the observed outcomes, which does not include the response model and therefore, we cannot use a similar strategy.
When following the approach proposed in Section 3, the likelihood (3.4) suggests at least two procedures for the selection of the response model in SAE under NMAR nonresponse. (a) Compare different models based on information criteria such as the Akaike information criterion, AIC = –2l() + 2dim(γ), or Schwarz information criterion,
Simulation Study
Simulation Set-up
We start by defining the sample model before nonresponse because as stated in Section 4, our approach for estimating the response model is based on fitting a model to the observed outcomes, which does not include the response model. For convenience, we assume noninformative sampling of areas and within the areas, such that the sample model before nonresponse is the same as the population model. Note that although the sampling design defines the observed model (2.2), once this model is estimated, the sampling design does not affect the estimation of the response probabilities in Section 3.
The simulation study consists of the following steps:
Generate auxiliary values
Consider three unit response models (no selection of areas):
Select 3 sets of respondents:
R1 uses Poisson sampling, independently between the units with response probabilities
R2 is the same as R1 but with response probabilities
R3 is the same as R1 but with response probabilities
The 3 response probabilities yield similar response rates of 65–75 per cent.
The working model for the observed data for the responding units is,
Remark 4: The working model (5.2) is correct for the observed sample R3 that corresponds to MAR nonresponse, but not for R1 and R2, under which the nonresponse is NMAR.
Define three working response models:
Note that pr(yij,xij;γ3) is nested in pr(yij,xij;γ1) and pr(yij,xij;γ2), and pr(yij,xij;γ1) is nested in pr(yij,xij;γ2). The response probability pr(yij,xij;γ3) defines MAR nonresponse and hence, can be estimated by solving (3.3).
Estimate the unknown parameters θ0,θ1,θ2 in (5.2) by SAS Proc NMIX, and then estimate γ by maximizing (3.4), as described in Section 3. The maximization was carried out by use of SAS Proc NLIN under the following 9 scenarios, as defined by the true response model and the assumed working response model:
S1: R1 set of respondents, M1 working response model.
S2: R1 set of respondents, M2 working response model.
S3: R1 set of respondents, M3 working response model.
S4: R2 set of respondents, M1 working response model.
S5: R2 set of respondents, M2 working response model.
S6: R2 set of respondents, M3 working response model.
S7: R3 set of respondents, M1 working response model.
S8: R3 set of respondents, M2 working response model.
S9: R3 set of respondents, M3 working response model.
Select the response model based on:
1. The Likelihood Ratio Test (LRT); test a saturated model [
2. AIC selection criterion: compare the values of the AIC as obtained for competing two models;
3. BIC selection criterion: compare the values of the BIC as obtained for competing two models.
Repeat the whole process independently 500 times.
S1 Vs S2 (R1 – set of respondents, M1 – correct model, M2 – saturated model). Note that although M2 is a saturated model, it is also correct but with an additional term. The LRT selects the model M1 in 368 out of the 500 simulations. AIC selects M1 in 305 out of 500 simulations, BIC selects M1 in 324 simulations.
S1 Vs S3 (R1 – set of respondents, M1 – correct model, M3 – incorrect nested model). The LRT selects the correct model M1 in 500 out of the 500 simulations. AIC and BIC likewise select M1 in all the 500 simulations.
S4 Vs S5 (R2 – set of respondents, M2 – correct model, M1 – incorrect nested model). The LRT selects the correct model M2 in 433 out of the 500 simulations. AIC and BIC select M2 in 483 simulations.
S4 Vs S6 (R2 – set of respondents, M2 – correct model, M3 – incorrect nested model). The LRT selects the correct model M2 in 490 out of the 500 simulations. AIC and BIC select the correct model in all the simulations.
S7 Vs S8 (R3 – set of respondents, M3 – correct model, M1 – also correct but a saturated model). The LRT selects the model M3 in 241 out of 500 simulations. AIC selects M3 in 225 out of 500 simulations; BIC selects M3 in 420 simulations.
S7 Vs S9 (R3 – set of respondents, M3 – correct model, M2 – also correct but a saturated model). The LRT selects the M3 model in 241 out of 500 simulations. AIC selects M3 in 361 out of 500 simulations, BIC selects M3 in 450 simulations.
Note that when R3 is the set of respondents and M3 is the correct model, M1 and M2 also produce correct estimates of the response probabilities, although with additional estimated parameters. Thus, the fact that the LRT test and the AIC select the M3 model in about half of the simulations is not surprising. The use of the BIC criterion performs better in these cases.
The results are summarized in Table 1.
Percentages Out of 500 Simulations in which each of the Three Selection Procedures Selected the Correct Model, for Different Combinations of Correct (Rows) and Working (Columns) Response Probability Models.
Percentages Out of 500 Simulations in which each of the Three Selection Procedures Selected the Correct Model, for Different Combinations of Correct (Rows) and Working (Columns) Response Probability Models.
Finally, we consider the case where a working model is incorrect but might be a good approximation of the correct model: let R1 be the set of respondents such that M1 is the correct working model. Let M4 be the following working model:
Sverchkov[13] suggested testing whether the response is NMAR or MAR by testing the significance of the corresponding estimated coefficients in the saturated response model. We applied this idea by testing the significance of the estimated coefficients
We considered two samples of respondents, R3 and R1. For R3, we found that when testing the working response model
For the respondents’ sample R1, the response model
In this article we investigate the use of the likelihood function of the observed respondents’ data for selecting an appropriate response model under possible NMAR nonresponse. For estimating the hypothesized model, we applied the missing information principle. Despite of what seems to be a rather complex estimation process, we find in our simulation study that the AIC and BIC information criteria and the LRT test, when applicable, perform well for model selection. Clearly, the use of other likelihood-based tests and selection criteria should be investigated as well.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The authors received no financial support for the research, authorship and/or publication of this article.
