Abstract
This article examines the problem of response error in survey earnings data. Comparing workers’ earnings reports in the U.S. Census Bureau’s Survey of Income and Program Participation (SIPP) to their detailed W-2 earnings records from the Social Security Administration, we employ ordinary least squares (OLS) and quantile regression models to assess the effects of earnings determinants and demographic variables on measurement errors in 2004 SIPP earnings in terms of bias and variance. Results show that measurement errors in earnings are not classical, but mean-reverting. The directions of bias for subpopulations are not constant, but varying across levels of earnings. Highly educated workers more correctly report their earnings than less educated workers at higher earnings levels, but they tend to overreport at lower earnings levels. Black workers with high earnings underreport to a greater degree than comparable whites, while black workers with low earnings overreport to a greater degree. Some subpopulations exhibit higher variances of measurement errors than others. Blacks, Hispanics, high school dropouts, part-year employed workers, and occupation “switchers” tend to misreport—both over- and underreport—their earnings rather than unilaterally in one direction. The implications of our findings are discussed.
Introduction
Economic information reported in household surveys provides essential data for a wide range of policy issues and academic research. A major advantage of survey data is the breadth of information they collect from individuals at the micro level and their widespread availability after they have been collected. A potential disadvantage is measurement error due to inaccurate respondent recall of economic information, especially of personal income. While large surveys such as the U.S. Census Bureau’s Current Population Survey (CPS), the Panel Survey of Income Dynamics (PSID), and Survey of Income and Program Participation (SIPP) provide high-quality, widely used earnings data, they nevertheless contain varying degrees of measurement error (Bollinger 1998; Gottschalk and Huynh 2010; Hurd, Juster, and Smith 2003; Pedace and Bates 2000; U.S. Census Bureau 2006; Alwin 2007). Discrepancies between respondents’ reported earnings and their true earnings can bias the coefficient estimates and/or lead to less precise standard error estimates (Angrist and Krueger 1999).
To date, much of the empirical literature has found that low earners tend to overreport their wage and salary income while high earners tend to underreport these sources of income. That is, measurement errors in earnings are mean-reverting. Despite the presence of mean-reverting errors, studies often find only modest to no correlations between measurement errors and socioeconomic and demographic variables (Bound and Krueger 1991; Bound, Brown, and Mathiowetz 2001; Bricker and Engelhardt 2008). What previous studies have revealed are the directions of bias in measurement errors.
However, measurement errors consist of two components: bias and variance (Groves 1989; Alwin 2007:5-6). Since previous analyses typically assume homoscedasticity of variance conditional on different population characteristics by ordinary least squares (OLS) models, the second component of measurement errors (i.e., variance) is largely ignored, with the notable exception of recent work by Gottschalk and Huynh (2010) and Kim and Tamborini (2012). As a result, the effects of socioeconomic and demographic variables on measurement error at different points of the distribution are, by and large, unknown. If variances are not constant across population characteristics, the conditional mean from OLS regression models could fail to capture the differentiated response error distributions by subpopulations and, thus, may fail to rightly estimate the conditional variances for these subpopulations (Hao and Naiman 2007:26).
With regard to the bias component, an issue not adequately addressed in previous studies is the association between socioeconomic and demographic variables and measurement errors after controlling for levels of earnings. Given the likely presence of mean-reverting errors, it is important to test the association of socioeconomic and demographic variables independent of levels of earnings. For example, if older workers tend to underreport their earnings, it should be tested whether it is simply because they are concentrated at the higher end of earnings distribution (where underreporting is likely), or because they tend to underreport across the entire earnings distribution. To our knowledge, none of the existing research addresses this issue.
This study is designed to begin to fill these gaps. Our primary goal is to assess properties of measurement error in survey earnings; that is, to examine the discrepancy between what respondents report as their earnings in surveys with their “true” earnings. We pay particular attention to (1) the conditional bias of socioeconomic and demographic variables on measurement errors net of levels of earnings and (2) the heterogeneous effect of earnings determinants and demographic variables on the entire distribution of response error. The former question is about the bias component of measurement errors and the latter question is about the variance component. To this end, we employ both OLS and quantile regression techniques.
We use a unique and restricted-use data set that links workers in the SIPP 2004 calendar year file with their W-2 earnings records from the Social Security Administration (SSA). Using a recently matched data set, our analysis provides more up-to-date estimates on measurement error in earnings than previous studies, which rely on older data sets, such as the 1992 SIPP matched data (Pedace and Bates 2000), the 1970s matched CPS (Bound and Krueger 1991; Bollinger 1998), or the 1990s matched PSID (Duncan and Hill 1985; Pischke 1995; Hyslop and Imbens 2001). Furthermore, the richness of the survey data allow us to consider the effects of several independent variables heretofore not examined in the literature, such as the respondent’s occupation, industry, and employment stability over a calendar year.
Our results reveal the heterogeneous effects of covariates on measurement errors in terms of bias and variance. The findings also reaffirm the extent of mean-reverting error in survey earnings. Consistent with the previous research, we do not find strong associations between response error and many demographic characteristics at the first moment. However, significant effects are observed when the sample is divided by the level of administrative earnings (thus controlling for true earnings). Furthermore, we uncover significant relationships between measurement error and demographic characteristics at higher orders, which implies that the extent of variance of measurement errors is not homoscedastic, but varying across covariates. This suggests that greater attention should be paid to potential heterogeneous effects of independent variables on response error. This heterogeneity does not necessarily yield biased estimates for the mean in OLS models; however, it may cause bias in estimating inequality and poverty levels for subpopulations.
Put together, our intention is not to develop formal mathematical theory about earnings measurement error, but rather to examine its properties in a large nationally representative survey and to determine the impact of a broad range of socioeconomic and demographic characteristics on such error. The next section sketches the background literature, focusing on response error in wage and salary income in survey data. The subsequent sections describe our data and methods and report the results. The final section summarizes our findings and draws conclusions.
Background and Significance
Measurement error in survey data has been discussed for decades, among sociologists, economists, and survey methodologists. Topics include, among others, the implication of random measurement errors for causal inferences (Blalock 1965; Althauser and Rubin 1971); the impact of social desirability bias in self-reported religious attendance (Presser and Stinson 1998; Regnerus and Uecker 2007); measurement error in poverty rates (McGarry 1995) and self-reported health (Baker, Stabile, and Deri 2004); the relationship between measurement error and race (Bielby, Hauser, and Featherman 1977; Wolfle 1985; Wolfle and Rorbertshaw 1983) and gender (Corcoran 1980); and measurement error in self-reporting occupations for intergenerational mobility research (Bielby et al. 1977; Corcoran 1980; Wolfle and Robertshaw 1983).
Although the sociological literature has focused on many of the variables mentioned above, the error properties in survey-reported earnings have been analyzed primarily by labor economists and survey methodologists rather than by sociologists (e.g., Bound and Krueger 1991; Duncan and Hill 1985; Morgenstern 1963; Withey 1954). However, given the widespread use of survey earnings for empirical studies in sociology and for policy formulation, measurement error in earnings should not be overlooked by sociologists.
Sources of Measurement Errors and Study Methods
Inaccuracies in surveys can result from various factors. Alwin (2007:15-33) summarizes the main sources of measurement error in general, and Moore, Stinson, and Welniak (2000) provide a useful overview of income measurement error. Inaccurate information may result from survey respondents’ inability to recall correctly. Confusion may arise over income concepts and definitions. A likely source of error is take-home versus gross earnings. Even though respondents are typically asked to report gross earnings, they may not include pretax deferred wages, such as employee contributions to 401(k) retirement plans (Coder and Scoon-Rogers 1996). Response error can also result from proxy respondents, such as a spouse, who may be more prone to provide inaccurate economic information (Huynh, Rupp, and Sears 2002; Pedace and Bates 2000). Survey design factors may also encourage response error. Nonresponse, that is, when respondents refuse to respond to a single income question or entire surveys, represents one factor (Schräpler 2004).
Prior analyses of the accuracy of economic information reported in surveys have adopted various approaches. Validation studies often compare income estimates across different surveys (Coder and Scoon-Rogers 1996). Another technique evaluates aggregate estimates in surveys against independent “benchmarks,” such the National Income and Products Accounts (NIPA) produced by the U.S. Bureau of Economic Analysis (BEA) or Internal Revenue Service (IRS) tax return data (Moore et al. 2000; Roemer 2000). While benchmark sources are useful in that they contain data on aggregate amounts of personal income and are relatively easily accessed by researchers, they tend to have different sampling frames and definitions than their survey counterparts. Another drawback is that research questions focused on the micro level (e.g., association of measurement errors with population characteristics) cannot be addressed.
A different approach, comparable to this study, evaluates survey data linked to administrative data of similar kind (Calderwood and Lessof 2009). A number of empirical studies have followed this method by exploiting U.S. Social Security Administration earnings records matched to national survey data, namely the CPS or SIPP (Bound and Krueger 1991; Gottschalk and Huynh 2005; Huynh et al. 2002; Pedace and Bates 2000). A primary advantage of these data is that they permit a direct comparison of the same individual’s survey-reported and administrative-reported earnings. A drawback is accessibility, which is not widespread due strict disclosure and confidentiality rules for administrative data (Abowd and Woodcock 2001; Davies and Fisher 2009). As a result, a limited number of studies have utilized survey data linked to administrative data.
The Mean-Reverting Error
In most cases, researchers do not observe the true earnings, yi
, for an individual i; instead, they utilize survey-reported earnings, yi
SR. These data, yi
SR may have measurement error, ui
SR as follows:
Thus, measurement error is defined as the discrepancy between the true earnings, yi
and survey-reported earnings, yi
SR:
Several of the previous studies utilizing matched administrative data show that earnings are, on average, higher in administrative records than in surveys such as the SIPP and CPS (i.e.,
One important line of analysis has been whether measurement errors are negatively correlated with true earnings. If the distribution of response errors is classical, the coefficient estimate of survey-reported earnings regressed on response errors will be zero. As long as the distribution of the measurement error is uncorrelated with the true earning, yi
, and has a mean zero, survey-reported earnings, yi
SR is an unbiased estimate of true earnings, yi
on average. However, if covariance between true earnings and measurement errors, cov(yi
, ui
SR), is not zero, survey earnings are no longer unbiased. For this reason, previous studies on the measurement error tested whether the measurement errors are linearly correlated with the true earnings using an equation as follows:
Bound and Krueger (1991) find that respondents tend to overreport their earnings at the low-end of earnings distribution, but as earnings increase, respondents are more likely to underreport earnings. That is, δ in equation (3) is significantly negative, and measurement errors are not classical, but mean-reverting. Bound et al. (1994) provide supporting evidence of mean-reverting errors using a matched PSID data set. Bricker and Engelhardt (2008) find mean-reverting errors among 50 year or older respondents using the Health and Retirement Study (HRS) linked to W-2 earnings records. Studies with different data sets provide additional evidence supporting Bound and Krueger’s (1991) findings (e.g., Akee 2011).
Covariates of Measurement Error in Earnings
Given the presence of nonclassical measurement errors, it is important to examine the association between usual earnings determinants and population characteristics and measurement error in survey earnings (Bound and Krueger 1991; Pedace and Bates 2000). If usual earnings determinants are associated with nonclassical measurement errors, the regression coefficients of those covariates are biased. To address this concern, previous studies estimated equation (4):
where
Several studies find weak to no association between measurement error in earnings and education (Bound and Krueger 1991; Bricker and Engelhardt 2008; Pedace and Bates 2000). Race also seems not to have a significant association with measurement error (Pedace and Bates 2000). Bound and Krueger (1991) find that women have a slight tendency to underreport earnings, but observe no significant association by age and education. Bollinger (1998) concludes that response error in earnings cannot be treated as additive white noise because of its relationship with gender and true earnings. Gottschalk and Huynh (2005), using the 1996 SIPP matched to SSA’s Detailed Earnings Record (DER), find modest underreporting bias for males under age 60, but do not find strong empirical support for the argument that older workers misreport their earnings more than younger workers. The total explanatory power of population characteristics put together—measured by R 2—is surprisingly low.
What previous studies have largely ignored is the effects of socioeconomic and demographic variables after controlling for levels of earnings. Given that many demographic variables are strongly associated with earnings and that earnings and measurement errors are negatively correlated, very weak to no associations between demographic variables and measurement errors are somewhat surprising. Equation (4) examines the mean effects of covariates on measurement errors, but does control for levels of earnings. This is an important gap in literature. For instance, consider a scenario where black workers are more likely to underreport their earnings than white workers with a similar level of earnings. OLS could fail to capture this bias because the tendency of underreporting for black workers and the tendency of overreporting for low-income earners (where black workers as a group are concentrated) cancels each other out, resulting in a zero coefficient of blacks in OLS.
Another issue is variance. Unbiased estimates on average do not necessarily indicate that estimates of means by subgroups are unbiased nor that estimates of variance and measures at higher order by subgroups are unbiased. By using equation (4), previous research mostly addresses the former concern while paying little attention to the latter. The mean effects of usual earnings determinants and demographic variables may not be sufficient enough to understand their impact on measurement errors, particularly when the accuracy of inequality estimates or poverty levels by subpopulations are a main concern.
From equation (1), we can infer that the difference of the variances between true earnings and survey earnings is a combination of the covariance of true earnings and the measurement error (i.e., 2cov(yi
, ui
SR)) and the variance of measurement errors (i.e., var(ui
SR)) as shown in equation (5):
The presence of mean-reverting errors indicate cov(yi , ui SR) is negative, but does not inform anything about var(ui SR). If some subpopulations are more likely misreport, it implies that var(ui SR) is larger for those populations than other groups, which will lead to the overestimation of the level of earnings inequality for those subgroups in survey data.
If, for instance, workers without a college degree tend to overreport at the low end of earnings distribution while, at the same time, underreport at the high end of the earnings distribution, the mean estimate of measurement errors for less educated workers may converge to zero. Nonetheless, this does not necessarily signify that education has no effect on response errors, but rather, that education is associated with both excessive over- and underreports. To this extent, the estimated variance for less educated workers may have downward bias. More problematic, the estimated effects of education at other points of earnings distribution (i.e., not at mean or median) could be biased.
Pedace and Bates (2000) provide empirical support for the notion that studying the mean effects of response errors may not be sufficient. Using a SIPP 1992 file matched to SSA’s Summary Earnings Records, the authors employ logit models to assess the probability of reporting 25 percent below or above true earnings. Findings show that individuals aged 50 to 64, males, and high school graduates are more likely to underreport, while those residing in the south and working in military-related occupation are more likely to overreport. Significant effects of occupation and industry were also found. For example, workers in technical/sales had a higher likelihood to underreport their earnings. The magnitude of misreporting, Pedace and Bates (2000) find, is largest among the lowest earnings group. The authors speculate that significant overreporting at the bottom tenth of earnings distribution could be caused by the tendency of low earners to work in unstable labor markets or in service sectors receiving tips.
Put together, the existing literature suffers from two critical limitations. First, empirical studies have failed to examine adequately the association between covariates and measurement errors after controlling for the level of earnings. Second, most studies have attempted to assess response error in earnings using OLS regression techniques. Consequently, research has tended to overlook the variance component of measurement errors; that is, the possibility of heterogeneity of the effects of standard earnings determinants and population characteristics on measurement error. The current study extends work in this area beyond these limitations.
Data and Methods
Data are drawn from a 2004 SIPP 12-month longitudinal file matched to the SSA’s earnings records. The SIPP is a nationally representative household panel survey of the civilian noninstitutionalized U.S. population administered by the Census Bureau. Based on agreements between SSA and the Census Bureau, SIPP respondents are matched to their administrative earnings using a sophisticated matching procedure. 1 The matched data combine the accuracy of SSA earnings records with the broad scope of information provided in household surveys (Davies and Fisher 2009). McNabb et al. (2009) discuss the matching process in greater detail, and Olsen and Hudson (2009) and Tamborini and Iams (2011, appendix) provide more information on SSA’s earnings records.
We use the SIPP for survey earnings and assume administrative earnings as a gold standard (i.e., true earnings). Our analysis uses SIPP-reported wage and salary income, rather than the income as a whole, which may encompass returns from assets, such as interest income, and transfer income from government programs. SIPP earnings include gross wages, salaries, tips, bonuses, overtime pay, severance pay, and commissions from all jobs before deductions (U.S. Census Bureau 2006). To estimate annual earnings for 2004, we draw monthly wage and salary income amounts reported in the longitudinal core files for waves 1 through 4 of the 2004 SIPP panel. At the time of each interview (every 4 months), respondents are asked to report their monthly wage/salary income for the previous 4 months; annual earnings are derived by summing these amounts over the calendar year.
The current study uses respondents’ administrative earnings contained in SSA’s Detailed Earnings Records (DER) file to assess the accuracy of respondents’ SIPP earnings. The DER uses information derived from SSA’s Master Earnings File, a primary repository of earnings data for the U.S. population sent to SSA by employers. Annual earnings are derived from employer-reported W-2 form(s) for all jobs covering a year and reflect full wage compensation under federal income tax, consisting of uncapped wages, salaries, and other compensation such as bonuses, commissions, and tips. Our administrative annual earnings also include respondents’ deferred wages, but do not add in self-employment income. In the SIPP, earnings-related questions also emphasize “gross pay before deduction.” As a result, earnings in both sources are conceptually similar in this study.
Even though a primary advantage of administrative data are its relative accuracy and completeness (Haines and Greenberg 2005), there are several noteworthy limitations (Abowd and Stinson 2011; Kapteyn and Ypma 2007). First, since administrative earnings rely on IRS W-2 records, they may omit under-the-table earnings. However, the fact that our analysis shows that respondents’ administrative earnings exceed, on average, their survey-reported earnings is evidence of the overall superiority of administrative earnings in this regard. Nevertheless, there remains the possibility that survey-reported earnings could be more accurate for some subgroups of workers. All results reported in this study (as well as other studies using matched data sets) should be interpreted with this limitation in mind.
Second, not all SIPP respondents can be matched to their tax records. As shown in Table 1, the match rate between our sample of workers in the SIPP and their administrative earnings records is high at 87.5 percent. While there are some minor differences in match rates by demographic groups, the matched- and full SIPP samples are very similar in terms of age, income, current marital status, educational level, and race/ethnicity. Most differences do not exceed 1 percentage point.
Descriptive Statistics of the Study Sample.
Source. Authors' calculations using SSA administrative earnings records matched to the 2004 SIPP (calendar year).
Note. a We used the matched sample to create administrative earnings quintiles, thus the match rates by these subgroups are 100% by definition.
Third, although we treat workers' DER earnings as benchmark earnings for the purpose of the analysis, we recognize that administrative data may contain other imperfections not already mentioned. Using Swedish data, Kapteyn and Ypma (2007) and Meijer, Rohwedder, and Wansbeek (2012) provide some evidence that findings in the measurement error literature such as mean-reversion are sensitive to incorrect matching of survey respondents with administrative files. Nevertheless, the initial matching algorithm used by the Census to link SIPP respondents with SSA data is conservative in determining a successful match. Moreover, SSA validates that survey respondents are matched to the correct administrative record (McNabb et al., 2009). As an additional precaution, we contrasted person-level date of birth and gender across SIPP and administrative data, excluding respondents with non-matching values in our final analysis sample.
Our final base sample contains a number of restrictions. First, the sample is limited to respondents aged 18 to 60 assigned a SIPP 2004 calendar weight and matched to SSA’s DER records. Second, respondents must have positive 2004 annual SIPP and administrative earnings. To avoid definitional discrepancies in self-employment income between the data sources, 2 the base sample excludes respondents who report self-employment income in any month during 2004. To focus more clearly on response error, the sample excludes SIPP respondents with imputations for wage and salary income in any month of 2004 (about 20 percent of file). It also excludes respondents in active duty armed forces. Given these boundaries, our final sample yields 23,396 unweighted individuals.
Analytical Strategy
Our analysis defines measurement error as the discrepancy between survey-reported earnings (SR earnings) and actual earnings, represented in this case, by the administrative records (AD earnings). This difference (i.e., SR earnings − AD earnings) is assumed to reflect measurement error in labor earnings. All earnings in our analyses are log transformed. The log earnings gap between SR earnings and AD earnings is the main dependent variable for multivariate analyses.
Following previous research, we start with equation (3) in which the dependent variable is the measurement errors, ui , and the only independent variable is the administrative earnings, yi . With this analysis, we test whether the mean-reverting errors are still evident in a recent matched earnings data set. Next, we estimate equation (4) to test the association between response errors and covariates. To test whether the effects of the covariates vary by levels of earnings, we divide our sample into five earnings quintiles according to their AD earnings and estimate equation (4) separately for each group. Thus, the first group is respondents of the bottom 20 percent of AD earnings; the second group is the second lowest 20 percent; the third group is the middle 20 percent; the fourth group is the second highest 20 percent; and the fifth group is the top 20 percent.
What researchers test with equations (3) and (4) are the mean effects, which are the bias component of measurement errors. However, if a certain subpopulation both over- and underreport, its mean effect converges to zero. If this is the case, we could wrongly conclude that such variables are irrelevant to response errors. To account for this concern, our analysis employs quantile regression models as follows:
where
What equation (7) does is to estimate γ at different quantiles by changing the weights (τ and 1 − τ) on the positive and negative residuals. For example, at the median (τ = .5), positive and negative residuals are given equal weight so that the sum of absolute deviations is minimized (Koenker and Bassett 1978:38). The estimated coefficients, γ provide the estimates of the net measurement error of covariates x at τth quantile.
Quantile regression allows us to estimate the effects of earnings determinants and other population characteristics on measurement error at various points of quantiles net of all other control variables. A distinct advantage is that while OLS computes the effects at the mean level, quantile models estimate them at each quantile specified controlling for all the other variables (Koenker and Hallock 2001). The estimated coefficients at the median will be very close to those of the OLS if the dependent variable is approximately normally distributed. Furthermore, if there is no heterogeneity in the effects of earnings determinants and demographic variables on measurement error at different points of the error distribution, the estimated coefficients across different quantile points should be invariant (Hao and Naiman 2007:23-32). However, if the effects are heterogeneous, then the direction, magnitude, and their statistical significance will vary across quantiles. In this study, we use decile points of the response errors for dependent variables.
If the demographic variables are not associated with the measurement errors, their effects should not be statistically different from zero and vary across quantiles. However, even when the mean effect is statistically zero, if the estimated coefficient of a demographic variable (or a subgroup) is significantly positive at a higher quantiles (say 90th percentile) while it is significantly negative at the lower quantiles (say 10th percentile) and the difference is significant, then that subgroup’s survey-reported earnings have wider measurement errors in terms of variance. In the opposite case, when the estimated coefficient of a demographic variable (or a subgroup) is significantly negative at a higher quantiles while it is significantly positive at the lower quantiles, then the subgroup’s survey-reported earnings have less measurement errors in terms of variance.
In this study, we test the following independent variables.
Deferred compensation
Response error can occur if respondents become confused about their deferred earnings, most commonly employee contributions to a 401(k). As a result, it is important to control for whether respondents have deferred earnings in the year we examine earnings. As noted earlier, in the SIPP, respondents are asked to report their “gross pay before deductions in this month.” While this question is meant to prompt respondents to include their deferred wages in their wage/salary reports, the extent to which they include it is not known. The AD earnings do not suffer from this potential confusion, since they contain separate fields with this information. 3 Looking at respondents 2004 deferred compensation amounts, we create a dummy variable indicating the presence of deferred earnings (1 = yes, positive deferred earnings, 0 = no deferred earnings).
Education
We recode respondents’ educational level into four dummy variables: high school dropout, high school graduate, some college, bachelor’s degree, and graduate degree. The reference group is high school graduate.
Age
Misreporting could be related to age. For example, younger workers are more likely to be in less stable jobs or seasonal employment, which may in turn be more conducive to misreporting earnings (Gottschalk and Huynh 2005). Note that age (and other time invariant variables) comes from the December file of the 2004 calendar SIPP.
Sex
Existing research suggests different patterns of misreporting earnings by sex (Bollinger 1998; Pedace and Bates 2000). To test whether sex significantly influences measurement error in 2004 SIPP earnings, we include a binary variable for female. In addition, we estimate measurement error in gender-specific models.
Race and ethnicity
Race is a control variable coded into dummy variables: non-Hispanic Black, Hispanic, Asian, and other. Other includes Native American, Native Hawaiian, or Pacific Islander. Non-Hispanic white is the reference group.
Foreign-born
Measurement error may be more prevalent among respondents that are recent immigrants to the United States, in part because they may have less stable labor market conditions or different kinds of employment. The SIPP provides information on whether a respondent was born in the United States, which was used to construct a binary variable (1 = not born in United States, 0 = born in United States).
Employment status
Changes in employment status over the year could influence propensities to misreport wage and salary income. Using all 12 months of 2004 SIPP data, a dummy variable denotes whether or not the survey respondent changed employment status over the year (1 = unemployment at some time, laid off, or just not working for at least one month in 2004, 0 = with a job entire year, all months and weeks in 2004).
Occupation switcher
It seems plausible that a change in occupation during the year may associate with a change in earnings, which in turn, may increase a survey respondent’s likelihood to misreport their wage and salary income. We ascertain whether the respondent held the same occupation at the end of the year relative to the beginning of the year (1 = different occupation in December relative to January 2004, 0 = otherwise). The ability to track respondents’ time-varying occupational status over the year highlights an advantage of using a longitudinal survey file.
Industry switcher
Similar to occupation switcher, our model includes a dummy variable measuring whether a matched respondent was working in the same industry at the end of the year compared to the beginning (1 = different industry in December 2004 relative to January 2004, 0 = otherwise).
Occupation
Misreporting may be more prevalent among workers in certain occupations (Pedace and Bates 2000). For each job, SIPP collects data on occupation, industry, and work activities and duties. Based on this information, the Census provides the respondent’s detailed occupation codes, with which we construct occupational dummy variables for 23 categories based on the Standard Occupational Classification (See Online Appendix Table 1 [which can be found at http://smr.sagepub.com/supplemental/] for list).
Industry
Similar to occupation, respondents working in particular industries may be more prone to misreport their wage and salary income. Based on respondents’ detailed industry code, we construct 20 industry dummy variables based on the North American Industry Classification System (See Online Appendix Table 1 [which can be found at http://smr.sagepub.com/supplemental/] for list).
Results
Descriptive Statistics
Table 1 presents the descriptive statistics for our sample. Employment stability over the calendar year is important because workers with less may have greater proclivity to misreport their earnings than those with more. Three variables measure employment stability: not full-year employed, changed industry, and changed occupation. In our sample, 12.1 percent of workers were not full-year employed; industry switchers constituted 16.9 percent and occupation switchers were 21.8 percent. As expected, more unstable labor market conditions are more common at the low end of earnings distribution than at higher ends. All three variables also indicate that female workers’ labor market situation is somewhat less stable than male workers’ situation.
Workers with deferred earnings represent around 40 percent of the sample. Having deferred earnings, most commonly a 401(k), is generally an indicator of high-quality jobs. Consistent with the finding of employment stability, the proportion of female workers who received deferred compensation in 2004 is lower than male workers.
Table 2 displays respondents’ annual SIPP and administrative earnings in 2004. Within our entire sample, the mean value is $36,677 for SIPP and $38,967 for administrative earnings. Thus, on average, respondents’ SIPP earnings were $2,286 lower than their administrative earnings. The discrepancy at the median is smaller (SIPP is $1,361 lower). SIPP log earnings are 4.8 percent lower than administrative earnings. These results are consistent with previous findings, which show survey earnings to be modestly lower than administrative earnings, but largely not biased in estimating true earnings (Akee 2011; Bollinger 1998; Bound and Krueger 1991; Pedace and Bates 2000).
Summary Statistics for 2004 Annual Earnings, SIPP and Administrative Data, by Gender.
Source. Authors' calculations using SSA administrative earnings records matched to the 2004 SIPP (calendar year).
a SR= Survey-reported (SIPP) earnings; AD= Administrative earnings of SIPP respondents; Difference= SR – AD.
By gender, differences between survey and administrative earnings appear larger for male workers than for female workers. However, this magnitude is driven by higher earnings among men. Looking at log earnings, both male and female workers demonstrate similar magnitudes of underreporting. This descriptive result is contradictory to Bound and Krueger (1991) and Bollinger (1998), which found that women have a higher tendency of underreporting than men using older matched data sets. However, it is rather consistent with more recent studies such Bricker and Engelhardt (2008), which report no significant difference between older male and female workers in the HRS.
Figure 1 reveals the shape of the distribution of measurement error (SR earnings – AD earnings) using log-transformed and raw (no-log-transformed) constructs. Both graphs show an approximately normal distribution. Consequently, the difference between the conditional mean estimated by OLS and the conditional median estimated by quantile regression is expected to be small. Following previous studies, and because log-transformed earnings provides easy interpretation (i.e., percentage differences), we use log-transformed earnings in the remainder of this study.

Distribution of measurement errors (A) raw earnings; (B) log-transformed earnings.
Bias Component of Measurement Error: OLS Regression Analysis
OLS regression is used to investigate the influence of earnings and an array of population characteristics on measurement error in SIPP earnings. Model 1 regresses the difference between respondents’ survey and administrative log earnings on their administrative log earnings. Congruent with previous studies, administrative earnings (i.e., AD Earnings) is significantly negative in predicting the measurement error where AD Earnings is the only explanatory variable. 4
Model 2 assesses the influence of various earnings determinants and demographic characteristics. This model includes controls for industry and occupation switcher dummies as well as part-year employment, each of which measures an aspect of labor market instability over the year. It also employs 22 occupation and 17 industry dummies, variables not usually controlled for in prior studies. In spite of all these additional controls, the adjusted R 2 of model 2 is just .018, which is much smaller than the R 2 of model 1 (= .046).
The finding that many of the usual earnings determinants and demographic variables are not significant in model 2 is consistent with past research. Only deferred compensation, age, immigrant, and Hispanic dummies are statistically significant. It is worth emphasizing that deferred compensation is negatively correlated with measurement errors when holding all other variables constant, including occupation and industry. The underreporting at the higher distribution of earnings seems to be partially accounted for by the presence of deferred wages.
Model 3 conducts separate estimations on the AD earnings quintile subsamples. To begin with, we see that the mean-reverting errors are again evident. The intercept of the bottom 20 percent is highly positive (i.e., overreporting), while the intercept of the top 20 percent is significantly negative (i.e., underreporting).
Model 3 also presents a number of new findings that document heterogeneity in the covariates by earnings level. For example, in model 2, the effect of female was approximately zero; however in model 3, we see that female workers are less likely to overreport at the bottom 20 percent of AD earnings distribution, while they are more likely underreport at the second highest 20 percent of the AD earnings distribution. These contradictory trends cancel each other out, resulting in insignificant effects of female in model 2.
High earners tend to underreport their earnings, but as levels of education rises, workers are more likely to correctly report their earnings (i.e., less likely underreport) than high school graduates. When their earnings are low (i.e., bottom 20 percent), highly educated workers seem to more substantially overreport than low educated workers (the coefficient of Grad Degree is .135, and the p value of the coefficient is .06). In short, with regard to education, there seems to be a tendency of misreporting toward the conditional means. When the mean earnings for a group is high, workers in that group tend to overreport if they earn low (thus far from the group mean), and when the mean earnings for a group is low, workers in that group tend to underreport if they earn high (thus far from the group mean).
A similar tendency is found by race/ethnicity. To recall, in model 2, black does not seem to be associated with measurement errors. By contrast, in model 3, black workers underreport their earnings to a greater degree than white workers when their earnings are high, but they overreport to a greater degree when their earnings are low. The pattern of misreporting for Hispanic workers is not different from white workers at the low end of earnings distribution, but at the high end, Hispanic workers more excessively underreport than whites, as did Black workers.
Additionally, workers who experienced labor market instability over the year tended to underreport regardless of their earnings level. Pedace and Bates (2000) suggest that overreporting earnings in surveys may be caused, at least in part, by those who work in unstable labor markets or by workers in service sectors for whom tips are a substantial portion of income. Contrary to this expectation, the three variables that may indicate unstable labor market attachment—part-year employment, industry switcher, and occupation switcher—do not show any significant effects in model 2, nor do they show any significantly positive effects in model 3, suggesting that labor market instability may not cause overreporting at the low end of earnings distribution.
Overall, model 3’s results show significant associations between measurement errors and many of the covariates, which vary across different levels of earnings. These results suggest that contrary to the previous findings, many demographic variables are associated with measurement errors, but the effects are conditional on the level of earnings.
Another noteworthy point is that R 2’s in all models of Table 3 are fairly small. This indicates that less than 10 percent of measurement errors are systematic components, and more than 90 percent are random components. Past research focused on the former but mostly neglected the latter. By employing a set of quantile regressions, we begin to pay attention to the latter component in the following section.
OLS Regression Results of Difference between SIPP and Administrative Log Annual Earnings, 2004.
Source: Authors' calculations using SSA administrative earnings records matched to the 2004 SIPP (calendar year).
Note. The numbers within parenthesis are standard errors. a Quintiles of model 3 are based on administrative earnings.
* < .05. ** < .01. *** < .001 (two-tailed test).
Variance Component of Measurement Error: Quantile Regression Analysis
The OLS analysis, while insightful, can obscure the heterogeneity in the influence of earnings determinants and demographic variables on measurement errors. We cannot rule out the possibility that the conditional residual variances of response error in earnings are not constant for all values of the covariates (the constant conditional residual variance is an assumption of the OLS). To address this concern, we estimate the net effects of regressors on the difference between log annual survey and administrative earnings using a quantile regression technique.
Table 4 presents the results for two models, which are run at several percentiles representing the entire distribution of measurement error. 5 The .50 level reflects the conditional median function. If response errors are normally distributed and the homoscedasticity assumption of the OLS holds, the coefficient estimated in model 2 of Table 4 will be identical to model 2 in Table 3. The .90 level captures the conditional 90th percentile function in computing the effects of independent variables on response errors. As the OLS estimates the expected mean after controlling for other variables, the quantile regression for the 90th percentile computes the expected measurement error for the 90th percentile net of other covariates. Insofar as the homoscedasticity of variance and other assumptions of OLS hold, the estimated coefficient for the 90th percentile should not differ from the 10th percentile.
Quantile Regression Results of Difference Between SIPP and Administrative Log Annual Earnings, 2004.
Source: Authors' calculations using SSA administrative earnings records matched to the 2004 SIPP (calendar year).
Notes. The numbers within parenthesis are standard errors.
* < .05. ** < .01. *** < .001 (two-tailed test).
† indicates that the estimated coefficient is significantly different comparing to the coefficient of quantile .10 at alpha = .05
§ indicates that the estimated coefficient is significantly different comparing to the coefficient of quantile .50 at alpha = .05.
The asterisk marks (*) in Table 4 indicate the significance of the estimated coefficients at various Type I error levels. The dagger mark (†) signals whether the coefficient estimated for a specific quantile is significantly different from the coefficient at quantile .10; the section sign (§) indicates whether the coefficient is significantly different from the coefficient at median (at α = .05). For example, in model 2, the coefficient for less than high school (LTHS) at the 90th percentile is .081, which is significant at α = .01 (indicated by **). As marked by the dagger sign, the estimates for LTHS at the 10th percentile are significantly different from that for LTHS at the 90th percentile.
To show the distribution of measurement error, we first ran an intercept-only model (model 1). At the 10th percentile, the measurement error is −.379; at the median, it is −.036; and at the 90th percentile, it is .241. Therefore, the 90th percentile reflects the extreme overreporting while the 10th percentile reflects the extreme underreporting. Congruent with the OLS results, the intercept for the median is slightly negative.
Model 2 examines the effects of earnings determinants and demographic variables on measurement error at different percentiles. It has the identical specification as model 2 in Table 3 except that we use quantile regression. Figure 2 helps summarize the model’s results at nine decile points. If the effects of the noted independent variables are not associated with the measurement errors, the lines in Figure 1 should be flat along the zero point of y-axis. If the effects are significant but invariant across quantiles (i.e., equivalence of the coefficients), then the line will be flat above or below the zero point of y-axis. By contrast, an upward slope line indicates higher variance of measurement errors than reference groups, while a downward slope line indicates lower variance. In other words, an upward slope of a relevant group shows greater over- as well as underreporting relative to the reference group, while a downward slope indicates relatively smaller amounts of under- and/or overreporting.

Effects of demographic variables on measurement errors.
Overall, model 2’s results reveal several new findings. To begin with, none of the variables tested in our models show consistent under- or overreporting across deciles. This is consistent with the result of Table 3 that covariates are not associated with bias. Looking at deferred earnings, the quantile regression offers a sharper and more easily interpretable picture of the effect across the distribution of error than the OLS results. At the 10th percentile, the effect of deferred compensation is insignificant. As we move toward higher percentiles, its effect becomes more negative and significant (see Figure 1A). This pattern reveals that the main effect of having deferred compensation is a decline in the tendency to overreport at higher earnings levels, rather than a strong tendency to underreport.
Also notable, although education has no power in explaining measurement error in model 2 of Table 3, the quantile regression reveals several significant relationships. The most striking is the extent of overreporting among workers with LTHS education. Results (OLS and quantile regression at median) show that workers with LTHS education do not overreport, on average, compared with other workers, but when they do so, they are more likely to overreport at a higher magnitude. Some college and bachelor degree are weakly associated with more correct reporting than high school graduates. However, we are reluctant to link higher education directly with more correct SR earnings because the coefficients for graduate degree are not significant.
Quantile regression also provides interesting results with respect to race/ethnicity. Blacks are more likely to extremely over- and underreport relative to Whites. In the OLS analysis, these opposing tendencies cancel each other out. Hispanics also show an upward slope; however, they are not more likely to overreport at higher deciles. The upward slope for Hispanics is driven by underreporting at lower deciles.
The results also shed light on the heterogeneous influence of labor market factors. None of the labor market instability dummies were significant in model 2 using OLS analysis. However, according to our quantile regression, workers with part-year employment or industry switcher over the calendar year are more likely to misreport, either over- or underreport, relative to workers with full-year employment and workers who not change industries (see Figure 1D). 6 In sum, labor market instability seems to be associated with measurement errors in survey earnings, but not necessarily overreporting.
With regard to the occupation and industry dummies, 7 the coefficient for Accommodation/Food industry is consistent with Pedace and Bates’ (2000) speculation. That is, workers in this sector tend to overreport excessively. However, our occupation dummies give a more complex picture. For example, workers in food preparation/service occupations, personal service occupations, and sales-related occupations do not necessarily overreport, but instead are more likely to misreport their SIPP earnings, either over or under.
A final consideration is whether the effects of AD Earnings and other explanatory variables are different for male workers and female workers. To investigate the potential for varying effects between male and female workers, quantile regression models were estimated separately by gender. Overall, the gender-separated models show similar results to our analysis with male and female combined (results are not shown).
Conclusions
This study made use of a rich data set that matches respondents in the 2004 SIPP calendar year file to their W-2 tax records to investigate response error in survey-reported earnings. Our findings extend previous work on measurement error in several ways. To begin with, the data set utilized in this study has its own merits and permitted us to bring more recent estimates to bear on this topic, to assess measurement error in annual earnings derived from monthly self-reports recorded every 4 months rather than annual recall, and to use uncapped administrative earnings as the basis of comparison. The data set also allowed us to test the influence of several variables on measurement error heretofore not examined, such as the respondent’s employment status over the year and deferred compensation.
The contributions of this study, however, are not limited to the data set. Our results confirm that respondents’ annual SIPP earnings in 2004 are only modestly different than their administrative earnings for the same year. Congruent with the previous studies, the measurement error appears to be systematically associated with administrative earnings in the form of a mean-reverting error.
It is often speculated that the mean-reverting errors are driven by overreporting at the low end of earnings distribution (Bollinger 1998), and workers with unstable labor market experiences, who also tend to have low earnings, are mainly responsible for this overreporting (Pedace and Bates 2000). Our findings do not provide strong evidence for this speculation. At the low end of earnings distribution, part-year employed workers, industry switchers, occupation switchers, and immigrants, subgroups that are more likely to participate in informal labor markets, are not more likely overreport than other workers. The effects of occupation cast further doubt on the speculation that informal labor market activity is a main cause of mean-reverting errors. Workers in occupations with substantial tips or other compensation amenable to under-the-table earnings (thus, not recorded in their tax records), we find, are more likely to misreport their earnings in the SIPP, rather than solely overreport. Nonetheless, we are hesitant to conclude that informal labor markets are irrelevant to mean-reverting errors, because of the possibility that our variables do not fully capture informal labor market activities. The pattern of misreporting for workers participating in informal labor markets is warranted in future research.
A key finding is that many covariates were significantly associated with response errors once we controlled for levels of earnings. Thus, the moderate-to-no correlations between measurement errors and socioeconomic and demographic variables reported in previous research appear to be a result disguised by the interaction between covariates and earnings. For example, industry switchers tend to underreport their earnings when they earn high, but because most of industry switchers earn low earnings (where overreporting is dominant), the coefficient of industry switcher in OLS becomes zero. For another example, male workers tend to excessively overreport their earnings at the low end of earnings distribution, while female workers do less so when they earn low. Excessive overreporting for men is not observed in model 2 of Table 3, because men are concentrated among higher earners.
Such findings have potentially important implications for mean-reverting error, namely the possibility of varying mean-reverting errors across subpopulations. For example, our results show that black workers tend to overreport to a greater degree when their earning are low, while they underreport to a greater degree when their earnings are high. This suggests that the mean-reverting errors for black workers could be steeper than whites. In other words, δ is not a constant, but heterogeneous by subpopulations. Indeed, very recent research using a similar data set has confirmed this possibility (Kim and Tamborini 2012).
Another contribution is that the effects of socioeconomic and demographic characteristics on measurement error in survey earnings are not limited to the mean effects (i.e., bias), but evident at the higher moments (i.e., variance). In this way, quantile regression provides important insights into the heterogeneity of covariates on measurement errors, of which are obscured by OLS analysis alone. Heterogeneous effects are evident in various aspects: (1) LTHS educated workers tend to extremely overreport their earnings at the upper end of quantiles relative to other workers; (2) black and Hispanic workers are more likely misreport than corresponding white workers; and (3) immigrant workers show a tendency of extreme underreporting at the lower percentiles compared to native workers.
The finding that δ could vary across subpopulations and that the variance of measurement errors is heterogeneous across covariates has potentially important implications for the accuracy of inequality estimates using survey earnings. Recall that the extent of inflation for earnings inequality estimates is affected by the size of cov(yi , ui SR) which is determined by the size of δ and the size of var(ui SR), as shown in equation (5). Larger mean-reverting error (e.g., blacks) suggests larger negative cov(yi , ui SR), which in turn leads to underestimation of inequality. However, the negativity of mean-revering error can be offset by the inflation of variance estimates due to the positive ratio of the variance of measurement error to the variance of true earnings. The substantial heterogeneity of coefficients across quantiles for some demographic variables suggests that the sample variances for some subpopulation (e.g., blacks) can be more positively biased than other groups (e.g., whites). The heterogeneity of mean-reverting errors across demographic characteristics is beyond the confine of this article, but in a separate article, we find that the effects of measurement errors are indeed contingent on demographic characteristics (Kim and Tamborini 2012).
A worthwhile focus for future empirical work is the effects of measurement error in estimating conditional variance across different demographic groups. The empirical findings related to our quantile regression should be complemented by more theoretical and formal modeling of variance and higher moments in measurement error in survey data on earnings. Further research on the effects of response error on poverty and inequality measures, particularly for specific subpopulations, such as those without high school education, Black, or different age groups such as retirees, would be useful. A study using longitudinal data over more than one calendar year is also necessary. Underlying mechanisms of misreporting also need to be studied. A key open question is whether survey-reported earnings may be more accurate than administrative earnings for low earners, who may have more substantial under-the-table earnings. Ultimately, we agree with Alwin (2007:1-3) that “measurement issues are among the most critical in scientific research” and “there is hardly any justification for ignoring survey measurement errors” in assessing the likely biases in analyses.
Footnotes
Authors’ Note
The authors share equal responsibility for this work. The views expressed in this article are those of the authors and do not represent the views of the Social Security Administration. The administrative data used in this article are restricted use; all users must receive approval of the Social Security Administration and the Census Bureau.
Acknowledgment
We are grateful for helpful suggestions offered by Jim Sears, David Weaver, Hilary Waldron, Joyce Nicholas, the editor and reviewers.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
