Abstract
Self-reports of financial information in surveys, such as wealth, income, and assets, are particularly prone to inaccuracy. We sought to improve the quality of financial information captured in a survey conducted by phone and in person by encouraging respondents to check records when reporting on income and assets. We investigated whether suggestive prompts influenced unit response, compliance with the request to check records, precision of estimates, and accuracy. We conducted a split sample experiment in the Community Advantage Panel Survey in which half of telephone respondents and half of in-person household interview respondents were encouraged to check the records. We found a modest positive effect of prompts on compliance but no effect on unit response, precision, or accuracy.
Introduction
Self-reports of financial information in surveys, such as income and assets, are particularly prone to inaccuracy. Without access to their financial records, respondents must estimate values for difficult-to-recall financial data. If respondents provide inaccurate values, the validity of research based on those estimates may be unreliable. In their overview of measurement error in survey reporting of income, Moore et al. (1999) identify factors that occur during the understanding, retrieval, and response stages of survey response that contribute to the difficulty of reporting income. During the retrieval stage specifically, respondents may have difficulty reporting income because they cannot recall the information being requested and use faulty estimation strategies. We anticipate that the retrieval mechanism, and ultimately data quality, may be improved by respondents accessing financial records in reporting income and assets.
Asking respondents to refer to records during an interview has been suggested as an effective strategy before (e.g., Couper et al. 2013; Moon and Laurie 2010). Rather than assuming that a respondent will recall an exact piece of financial information, the interviewer may ask respondents to have copies of their bank statements on hand for the interview for easy reference and retrieval of the correct information. Moore et al. (1999) support the notion that recall and retrieval problems can be mitigated by bypassing memory retrieval through the use of records. Moon and Laurie (2010) similarly contend that better quality data can be collected with little risk to respondent cooperation by encouraging respondents to refer to records in an interview.
Couper et al. (2013) reviewed the (mostly unpublished) research on this topic and results of studies designed to gauge the effectiveness of asking respondents to use records in an interview. They find wide ranges in the reported levels of records use between studies but a similarity in that most respondents do not access records, even when prompted to do so. The authors also conducted two experiments in the Health and Retirement Study designed to gauge the impact of prompting respondents to use records in a web survey. They found that encouraging respondents to consult records did increase records checking significantly (from 46% to 55%), but this was not sufficient to change estimate precision. The group asked to check records had a lower response rate than those not asked to check records (77% and 80%, respectively).
While Couper et al. (2013) took an important step in experimentally attempting to increase record use in web surveys, we are not aware of studies conducted to increase record use for telephone or in-person interviews. We wanted to see if we could improve the data quality in phone and in-person modes by encouraging respondents to check records when reporting on income and assets. Our experiment was implemented in the Community Advantage Panel Survey (CAPS), an annual survey of low- to moderate-income (LMI) homeowners and renters conducted since 2003. The survey aims to capture the experiences of households and assess the pros and cons of homeownership for LMI households in the United States. While not a general population survey, CAPS respondents are largely representative of the LMI population with respect to income and race/ethnicity (Riley et al. 2009).
A primary goal of CAPS is to measure wealth differences between homeowners and renters. Evidence suggests that homeowners tend to accumulate more wealth than comparable renters (e.g., Belsky and Prakken 2004; Joint Center for Housing Studies of Harvard University 2013), possibly due to the forced savings aspect of monthly mortgage payments or because of behavioral changes that accompany the transition from renting to owning. Wealth-and-assets data are typically very noisy because their values change frequently and reporting these data requires recalling a number of different sources. In addition, wealth measures are often analyzed in an aggregated form, which potentially magnifies the amount of noise. These issues are relevant to other surveys such as the Survey of Consumer Finances (Pence 2006). Thus, improving the precision and accuracy with which this type of information is collected is important for ensuring that valid conclusions are drawn from the data. In 2012, the CAPS design featured a multimode administration, with some cases conducted via computer-assisted personal interviewing (CAPI) and others via computer-assisted telephone interviewing (CATI), and an experimental intervention that asked a random half of respondents to check financial records while answering survey questions.
The CAPI and CATI versions of the survey differed only in that the CAPI version included more detailed questions about wealth. For CATI, consolidated wealth questions captured much of the core wealth information, such as bank account balances or retirement account balances, in an aggregated manner.
Half the respondents in each mode were randomly assigned to receive a prompt suggesting they consult financial records during the survey (recent copies of bank statements, mortgage statements, school and car loan statements, retirement accounts, and insurance policies). For CATI respondents, a question near the end of the survey asked: “At any time during these questions have you consulted financial records to determine your answers to these questions?” Interviewers were also asked if respondents referred to records. For CAPI respondents, interviewers answered a follow-up question indicating that the respondent checked their records during the interview.
We tested the following hypotheses related to effects on unit response, compliance with the request to check records, and the precision and accuracy of reporting: Those asked to prepare records for the interview will not respond at the same rate as those not asked to prepare records. Among those who respond to the survey, those encouraged to check records will do so at a significantly higher rate than those not encouraged. Those asked to check records will display fewer behaviors that might indicate suboptimal data quality (i.e., rounding). Asking respondents to check records will result in some significantly different survey estimates compared with those not asked to check records, suggesting a potential improvement in accuracy by prompting respondents to check records. Due to interviewer presence, CAPI respondents will be more likely than CATI respondents to check records when prompted and more likely to provide more accurate responses to financial questions.
Regardless of the outcome of these tests, we were interested to see whether the above effects were similar between CAPI and CATI modes and whether a single method of prompting respondents may be similarly effective for each mode.
One could hypothesize that wealth-and-assets data quality could be higher either by phone or in person. In-person interviews may make it more likely for respondents to double-check financial records, given the presence of the interviewer. On the other hand, greater interviewer involvement has the potential to introduce increased variance in survey measures as a result of differences in the ways in which different interviewers conduct interviews. Such interviewer effects, which may cause respondents to give answers that are more socially desirable, could be larger in person than over the phone because of the proximity of the interviewer. In this context, social desirability could operate to increase the rate at which respondents check records when asked to do so if they view record checking as a form of social compliance. Alternatively, it could reduce record checking if respondents feel that checking records would cause them to reveal less socially desirable information. Such effects are expected to be largest for “sensitive” questions, such as those pertaining to finances.
Methods
Cases were generally assigned to CAPI or CATI based on mode assigned in prior years. The original mode assignment in 2005 for homeowners was based on whether the sample member was matched with a similar renter sample member. In general, owner sample members located in rural areas were not matched with renter sample members, so those households assigned to CAPI will primarily reflect those located in metro areas. This decision was made to contain survey costs, given the additional expense associated with CAPI relative to CATI. In addition, a small number of sample members who were originally assigned to CAPI did not complete the CAPI interview. These cases were later administered a refusal conversion interview via CATI and are assigned to the CATI group in this study. Overall, the goal was to keep the survey mode constant over time for those years in which the detailed wealth-and-assets questions were administered.
As indicated in Table 1, at baseline, the median age and income of 2012 CAPS eligibles were 31 years and about US$28,000; more than half of the sample are female; about 60% are white; nearly 75% had completed only a high school degree; slightly less than half of the sample was married; 20% were widowed or divorced; nearly 85% were employed; more than 65% of the sample were located in the South; and about 70% were homeowners. The characteristics of 2012 eligibles differ somewhat by survey mode assignment. In particular, survey participants assigned to CAPI were slightly older, were more likely to be renting, had a somewhat lower income, were more likely to be female, were slightly more likely to be black or Hispanic than white, were slightly less likely to have completed high school, were considerably less likely to be employed, and were more likely to be located in the South.
Baseline Characteristics of 2012 Survey Eligibles, Overall and by Mode Assignment.
Note: CAPI = computer-assisted personal interviewing; CATI = computer-assisted telephone interviewing; HS = high school.
There were 3,284 cases in the sample. Interviews were completed by twelve CATI and by forty-three CAPI interviewers. A total of 949 CAPI cases were randomly assigned (using SAS version 9.3 software) to the control condition and another 950 to the experimental condition. We tested the experimental and control groups for significant differences across race, gender, marital status, employment status, and level of education and found none.
Among CATI cases, 692 were assigned to the control condition and 693 to the experimental condition. Experimental cases received a lead letter informing participants that for some complex financial questions, records might be helpful. The letter specifically asks respondents to collect “copies of bank statements, mortgage statements, school and car loan statements, retirement accounts, and insurance policies.” The interview began with language that read: “You may recall that in the letter we sent we suggested that financial records may help answer some questions on the interview. Please feel free to refer to your financial records at any time during the interview.” Experimental cases, prior to receiving a set of items on their wealth and assets (where many complex financial questions appeared), also received a prompt that read: “The following set of questions ask for financial information that is difficult to remember. This section may be easier if you have your records ready.” Prompts to encourage the checking of records were purposely brief. We did not want to convey the perception that records checking was a requirement of the interview. We simply sought a way to gently encourage this action. Control cases were not given these prompts in the lead letter or in the interview.
Unit Response and Compliance
To assess the effect of the records checking prompt on nonresponse, we compared response rates between control and experimental cases overall and by mode. Response rates were calculated using American Association for Public Opinion Research (AAPOR) response rate 1 (AAPOR 2011), which divides the number of complete interviews by the number of interviews (complete plus partial) plus the number of noninterviews plus all the cases of unknown eligibility. We also compared the rate at which respondents checked records during the interview between control and experimental cases overall and by mode to determine whether the records checking prompts were effective in increasing the use of records during the interview.
Precision
To assess the effect of the experiment on precision of reported financial values, we measured the amount of rounding that occurred between control and experimental cases. This analysis assumes that the intervention was at least somewhat effective in increasing reference to records during the interview among experimental cases. To confirm that records actually result in less rounding, we also performed a cursory examination of differences based on whether the respondent checked records or did not, regardless of experimental condition. While informative, the latter comparison introduces the potential for spurious associations since we could not experimentally control who did and did not actually check the records.
In other analyses on this topic (e.g., Couper et al. 2013), the metric for measuring rounding examined the presence of terminal zeros in the financial values reported (e.g., 10, 100, and 1,000). Given the nature of the CAPS data, which included several different ranges of values from the single digits to the hundreds of thousands and the desire not to presuppose which values might be overreported without access to records, we sought a more rigorous test. For several key financial measures, we compared control and experimental cases using the Wilcoxon–Mann–Whitney (WMW) test, which orders reported values by descending frequency, plots the cumulative distribution, compares the area under each curve, and produces a Z score indicating whether the samples differ significantly (Hand 1997). We illustrate the WMW method in the Supplemental Online Materials.
Accuracy
To assess whether prompting respondents may have resulted in more accurate reports, we tested for difference in mean values of several financial measures between control and experimental conditions. The supposition is that differences may be attributable to more accurate reporting by those in the experimental condition since they were more likely to have checked records during the survey. We did not test whether self-reported financial data are actually valid (e.g., through our own check against financial records).
Our analysis focuses on several important measures of financial assets, including several where we expected that access to records may greatly increase the precision and accuracy of reporting and others where we expected a smaller impact. For example, some measures like monthly rent amount may be easy for most respondents to recall and would not benefit much from checking records. Others, such as amount of savings put away in the last year, may be too complex and involve too many sources to report accurately, even with access to records. Table 2 presents the measures we included in the analysis, the types of records we expected may be referenced in their reporting, and whether we expected records to be beneficial in the accuracy of reporting.
Selected CAPS Financial Measures, Records that Pertain, and Expected Improvement in Reporting.
Note: CAPS = Community Advantage Panel Survey.
Results
Table 3 presents the response rates to the CAPS survey by experimental condition. The response rate for control cases (not encouraged to check records) was 90% compared to 89% for the experimental condition. This difference was not statistically significant. There was also no difference between response rates among CAPI and CATI cases. The CAPI response rate was 93% for control and 92% for experimental cases. The response rate among CATI cases was 86% for both control and experimental cases.
Response Rate and Rate of Record Checking by Experimental Condition by Mode.
Note: CATI = computer-assisted telephone interviewing; CAPI = computer-assisted personal interviewing.
*Two proportion z test difference by treatment significant at p < .05
With regard to the effectiveness of the intervention asking experimental cases to check records, the treatment did bring about a significant increase in records checking overall and by mode. Overall, 216 of 1,641 (13%) respondents in the control condition checked records and 334 of 1,643 (20%) respondents in the experimental condition checked records, a difference significant at p < .05. Among CAPI cases, 147 of 949 (16%) control cases checked records, whereas 185 of 950 (20%) experimental cases checked records. The difference was more pronounced among CATI cases, with 67 of 692 (10%) control cases checking records and 150 of 693 (22%) experimental cases checking records. While these differences are statistically significant, the treatment resulted in only a minority (about one in five) in the experimental condition checking records. We examined control and experimental response rates by interviewer and did not find significant differences in the rates of compliance.
To assess which survey respondents were most likely to check records, we fitted fit a multivariate logistic regression model including demographic information and an indicator for whether the respondent was prompted to check records. The results for this analysis are presented in Table 4.
Logistic Regression Models Predicting Record Checking, Overall and by Mode.
Note: CATI = computer-assisted telephone interviewing; CAPI = computer-assisted personal interviewing.
***p < .01.
**p < .05.
*p < .10.
For the overall sample, male and minority respondents were significantly less likely to consult records during the interview, while college graduates and respondents who were prompted to check records were significantly more likely to do so. Male respondents were about 70% as likely as female respondents to consult records, while minority respondents were 40–60% as likely as white respondents to do so. Respondents who had completed at least a college degree at baseline were about 60% more likely than those with less education to consult records. Similarly, being prompted to check records increased the likelihood of doing so by about 50%.
The analysis also suggests that the relationship of demographics and prompting differs by mode. In particular, the relationships of age, gender, and race to the likelihood of checking records appear to be strongest for CAPI respondents, with younger, male, and minority respondents being significantly less likely to check records. The only highly significant predictors of record checking for CATI respondents are gender, educational attainment, and being prompted to check records. Being prompted to check records increased the likelihood of doing so by 40% for the CAPI respondents but by more than 100% for the CATI respondents. These results suggest that, controlling for demographics, prompting respondents to check records may be more effective when the survey is conducted over the phone. They also suggest that the likelihood that respondents will check records during in-person interviews may be influenced by social dynamics that are less salient during phone interviews.
To further assess how various demographic characteristics of the respondents may mediate the effect of prompting on the likelihood of record checking, we also estimated several additional specifications that include interaction terms. These interactions generally did not exhibit significant effects and are therefore excluded here. However, we do find that educational attainment appears to mediate responsiveness to prompting during CAPI interviews, with college graduates being significantly more responsive than respondents without a college degree in this context.
Regarding the effect of the treatment on precision, we examined whether the act of checking records was associated with less rounding. As mentioned previously, we did not have complete control over who did and did not check records and so there may be competing explanations as to why those who check records would also supply more precise estimates. Nevertheless, we did find evidence that the act of checking records itself was associated with lower levels of rounding on several measures. One example is household income. For this comparison, we compared a random sample of those who did not check records equal in size to the group who did check records to mitigate the effect of uneven sample sizes on the WMW statistic. Each group included 463 cases. The left-hand graph in Figure 1 presents the distribution of responses among those who did (dotted line) and did not (solid line) check records. As can be seen, the peaks in the distribution are generally higher among those who did not check records at values such as US$25,000, US$50,000, and US$100,000, suggesting more rounding in that sample compared to those who checked records. On the right-hand side of Figure 1, the cumulative distributions show the effect of these peaks with a greater percentage of cases among those who did not check records accounted for by fewer unique reported household income values. The WMW Z score of −1.99 is significant at p = .0461.

Community Advantage Panel Survey household income and rank order cumulative distributions by record checking.
While the example in Figure 1 suggests that checking records may be associated with lower levels of rounding, we did not find similar results for the comparison of control and experimental groups. Across all measures and modes, we found no significant differences in the level of precision used in reporting financial values, suggesting that the encouragement to check records was not effective in this experiment for this purpose. Even with an increase in records checking among the experimental group, there was likely not enough of a boost to result in less rounding as only about one in five respondents in the experimental group actually referred to records in the survey. Table 5 presents the WMW Z scores and p values comparing the control and experimental distributions.
Tests for Reduced Rounding by Experimental Condition and Mode.
Note: CATI = computer-assisted telephone interviewing; CAPI = computer-assisted personal interviewing; WMW = Wilcoxon–Mann–Whitney.
With regard to accuracy of survey estimates, we examined whether checking records was associated with higher accuracy, regardless of experimental condition. Since checking records was not experimentally assigned, it would be inappropriate to draw conclusions from these results, but generally we found very little evidence of differences between those who did and did not check records. Regarding the experiment, Table 6 presents the mean values for the selected financial values by experimental condition overall and by mode.
Reported Financial Values in Dollars by Experimental Condition and Mode.
Note: CATI = computer-assisted telephone interviewing; CAPI = computer-assisted personal interviewing.
*Means between those who were and were not encouraged significantly different at p < .05 (two sample t-test).
In terms of accuracy of survey estimates by experimental condition, we found few differences in means for reported values between those who were and were not encouraged to check records. The significant differences we did find (positive equity in primary home and household income) between experimental and control cases were in the CAPI mode. Again, the fact that so few respondents checked records overall makes the interpretation quite difficult. Perhaps records checking in CAPI, or records checking that can be confirmed by interviewers who are present during the interview, is more likely to result in more accurate reporting on financial items.
We did not have a method in place to verify CATI records checking. So perhaps the CAPI mode showed a few differences because records checking actually happened in that mode. The two significant differences we did find could reveal something interesting about respondent behavior and may be worth some investigation in future studies. Respondents not encouraged to check records reported higher household incomes and lower positive equity in their home. Perhaps these respondents were influenced by their knowledge of national trends where home equity was generally decreasing as were real household incomes. This could represent a strong argument for more records checking in survey interviews. However, we emphasize that very few differences were observed in these results by condition.
Discussion
Our analysis sought to address several hypotheses with regard to record checking by survey respondents. We hypothesized that being asked to prepare records for the interview would not lower response rate. With regard to compliance with the request to check records, we hypothesized that those encouraged to check records would do so at a significantly higher rate than those not encouraged. Regarding precision, we hypothesized that those asked to check records would be less likely to display behaviors indicating suboptimal data quality, such as rounding. Regarding accuracy, we hypothesized that survey estimates would differ among those asked to check records, suggesting a potential improvement in accuracy by prompting respondents to check records. Finally, we hypothesized that CAPI respondents would be more likely than CATI respondents to check records when prompted by a present interviewer and would therefore be more likely to provide accurate responses to financial questions.
Using the 2012 CAPS survey, we designed an experiment in which half the CAPI and half the CATI sample were randomly assigned to control and experimental conditions. The control group received no prompts to prepare and reference financial records during the interview. The experimental group was given the suggestion in the advance letter to prepare records and encouraged to access those records during the interview. We found that by encouraging respondents to check records, we were not discouraging their participation in the survey as response rates were the same across control and experimental groups overall and by mode. The treatment was somewhat effective in boosting the use of records overall and by mode, with around one in 10 respondents checking records in the control group and one in five in the experimental group.
The boost in records checking for the experimental group was higher in CATI. We do not have direct evidence to explain the larger boost in records checking for CATI compared with CAPI, but possible explanations include differential effects due to the sample composition by mode and overstatement of the checking of records by CATI respondents (CAPI respondents were not asked if they checked records—this was simply observed by the interviewer). We assumed that the presence of an interviewer would encourage more records checking. But this did not occur. It could be that CATI respondents indicated records checking when in fact there was none. No interviewer was present to confirm records checking, so while our data show more CATI record checking in the experimental group, this could simply be due to CATI respondents reporting what they think the interviewer wants to hear.
There was no discernible effect on the precision or accuracy of financial data comparing respondents who were and were not prompted to check records. Our treatment was suggestive rather than directive. We did not require that respondents prepare and reference records because we worried about increasing burden. While we did observe some mean estimate differences in CAPI, our data do not support the conclusion that records checking in CAPI leads to more accurate estimates for financial survey items. In the CAPI mode, the means for some items were significantly different by experimental condition, whereas in CATI no differences were present. It is possible that no differences were observed in CATI because records were checked less frequently than respondents reported. In other words, CATI respondents may have indicated they checked records when they actually did not. However, we cannot conclude that CAPI was more effective because we do not know if the means in the experimental groups are closer to the actual mean. We only know that the record checking intervention appears to have changed the mean estimate in CAPI. Future studies should consider designs that compare survey estimates to records that can be independently confirmed.
Any bias present in the CAPS data due to measurement error may be somewhat smaller than what would be observed in other surveys. The median annual income of 2012 survey eligibles at baseline was US$27,996, compared with median income of US$51,371 for households nationally (Noss 2013). In addition, the median age of the respondents at baseline was 31 years, indicating these households were situated toward the beginning of the time during which wealth accumulation is expected over the life cycle. Thus, there may be a restriction of range in measured values of wealth and assets for this population that reduces or attenuates bias relative to the general population or relative to older and wealthier segments of the population, such as retirees. We suggest caution in generalizing our results to surveys that have a different target population.
We suggest further research into the specific types of financial items where records could improve retrieval for improved precision and accuracy of reports. Our analysis makes some assumptions about which items would benefit most, and we did not collect data about the specific items where records should have been or were checked. These types of data, in conjunction with a more effective treatment, would allow for more specific conclusions and could be explored in a cognitive or usability laboratory setting before conducting further experiments in main study data collection. For instance, in a cognitive interview setting, interviewers could ask respondents about the ease of the task and confidence in the reported amounts, and record observations of respondents’ success and difficulties in retrieving and reporting the right information.
Our general conclusion mirrors that of Couper et al. (2013) who examined this issue in a web survey setting. For CAPI and CATI surveys, asking respondents to prepare and reference financial records during the interview will not reduce participation, but it may only result in a modest increase in the rate of records checking. Even when respondents check records, it is not clear that the data provided are more precise or accurate. Without a more directive intervention than the one we employed, suggestive prompts to check financial records will do no harm but may also do little good.
Footnotes
Acknowledgments
We thank the Ford Foundation, the UNC Center for Community Capital, and the Community Advantage Program.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: We received financial support from the Ford Foundation for authorship of this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
