Abstract
Call scheduling is a challenge for surveys around the world. Unlike cross-sectional surveys, panel surveys can use information from prior waves to enhance call-scheduling algorithms. Past observational studies showed the benefit of calling panel cases at times that had been successful in the past. This article is the first to experimentally assign panel cases to previously beneficial call windows. The results from a large-scale national survey in Germany show modest efficiency gains measured in number of call attempts needed until first contact but no gains in efficiency to gain cooperation.
Introduction
Call scheduling is a challenge in surveys conducted around the world, and years of research on improving call schedules have not shown the success expected (Wagner 2012). It is generally known that a variation in calling times reduces the number of attempts needed to reach sample units (Cunningham et al. 2003). It is not hard to imagine that the likelihood of agreeing to be interviewed is higher if the call happens at a time convenient for the household. A report by Statistics Canada (2007:40) rightly points out that: calling during the dinner hour may increase the likelihood of reaching people at home, but it may decrease the likelihood of participation because some people find these calls intrusive. If the ideal time to reach people at home is also the same time of day they are the least likely to agree to participate, then consider alternative call times. This strategy could be as simple as beginning calls earlier in the afternoon, perhaps at 3:00 p.m. rather than 4:00 p.m. A good contact rate is insufficient on its own. People must also agree to participate in an interview.
For the general population, a few patterns have emerged repeatedly over the years; for example, afternoons are more efficient than mornings on weekdays (Weeks et al. 1987), whereas weekend mornings are better than weekend evenings. These findings have found their way into advanced call scheduling algorithms (Brick et al. 1996), which often field a case first in the evening, wait a few days before the next call, and ensure that future calls are placed at other times.
The predictions of best times to call can be improved when auxiliary information about a case is known. Durrant et al. (2011), for example, showed across six U.K. household surveys that Saturday evenings are more beneficial when pensioners are present than when they are not, and Wagner (2012) showed for the U.S. National Survey of Family Growth that persons between age 15 and 44 have the highest contact probabilities when called in the late afternoon on Sundays and Monday evenings. Unfortunately, auxiliary information on a case base level is not available in most cross-sectional telephone surveys (Smith 2011). Panel surveys are in the fortunate situation to have ample covariates from prior panel waves. These variables can be used to predict best contact times for the subsequent wave. First attempts to use sample characteristics in call scheduling have been made by Wagner (2012).
Panel surveys also have many more direct measures of best calling times: call history information from prior waves. Lipps (2012) showed in a post hoc analysis of Swiss household panel data that the probability of contact and more importantly cooperation at field contact increases if respondents are contacted during the call window that had been successful in the prior wave. This effect holds even after controlling for respondent characteristics. Thus, in addition to predicting best call windows through covariates, the direct measure of the best window for a given household can be extracted from prior wave paradata.
There is some appeal in using this information directly instead of model-based call windows or entire call algorithms. Data collectors are often “owners” of the paradata but do not have easy access to respondent information. Also, some call schedulers are hard to program and specific programming efforts are needed to employ individualized call scheduling algorithms (Wagner 2012). Thus, a simple heuristic—such as “start fielding a case during the successful prior wave window”—might be easier to implement and might already provide gains in efficiency. In fact, Lipps (2012:14) suggests that “households should be called more often during the same time (window) at which it was first) contacted in the previous wave, especially at first call,” because there “seems to be a tendency that household-specific preferable calling and contact times persist across years.”
This article sets out to empirically test the suggestion put forward by Lipps (2012). Results from a randomized experiment are reported that used prior wave call information in a subsequent wave of data collection. To our knowledge, this kind of information has never been used experimentally for a sole intervention in a computer-assisted telephone interviewing (CATI) panel survey. The only other experimental efforts known to us were conducted within the Swedish Labour Force Survey, where certain call sequences were prioritized for certain subgroups of the population, but the contact strategies in the experimental group were also changed in several other ways, precluding any statements about causal effects of using the same window at the next wave (Lundquist 2011). Thus, this article gives a first answer to the research question “Does information from prior wave call record data improve process efficiency in panel surveys?”
Data
The survey data used in this article come from the German panel study Labour Market and Social Security (PASS). Since 2006, PASS data have been collected annually by the German Institute for Employment Research at the Federal Employment Agency (Trappmann et al. 2010). After wave 3, PASS changed the fieldwork organization (Müller 2011). For this article, we therefore focus on waves 4–6 of PASS data collection to avoid any possible confounders due to the switch in data collection procedures. PASS is a dual-frame mixed-mode (CATI and computer-assisted personal interviewing [CAPI]) survey. We focus solely on the CATI portion of PASS because call times in face-to-face interviews always reflect availabilities and preferences of the interviewers, whereas in centralized telephone studies calls are made and assigned to interviewers around the clock.
In preliminary work leading up to this study, we examined the effect of the prior (wave 4) call window on the probability of interview at first contact in wave 5 (see online Appendix 1 ). To avoid small sample sizes, we grouped cases into three time windows for weekdays (morning 0:01–12:00; afternoon 12:01–17:00; evening 17:01–0:00) and two for weekends (morning vs. afternoon/evening). Based on these data, we generated two variables: one indicating that the first contact in wave 5 was in the same time window as the first contact in wave 4, and another one indicating that the first contact in wave 5 was in the same time window as the interview in wave 4. Using these two variables, the effect on cooperation (at first contact) in wave 5 was examined through a logistic regression model.
Similar to Lipps (2012), we saw a positive effect of calling at a successful window from the last wave on the probability of gaining cooperation at first contact. However, we saw a stronger effect using the successful interview window from the last wave compared to the successful first contact window. Inspired by these results, we used the successful interview window from the last wave for the experimental manipulation of the call-scheduling algorithm. Although the same time windows were used, the experiment did not use the rough categorization into weekday and weekend but the actual days of the week.
Experimental Design
All panel cases in wave 6 who were also respondents in wave 5 were eligible for our experiment. Within strata, 80% of the panel cases were randomly assigned to the treatment group and 20% were randomly assigned to the control group. The unequal assignment was chosen to maximize efficiency in case the treatment worked, while still keeping a sizable control group. The call scheduler was programmed so that calls to the treatment group were first made on the same day of the week and in the same time window at which they were interviewed in wave 5.
Three time windows were specified for each workday, matching the time slots shown earlier. Cases who were interviewed on a Saturday or Sunday in wave 5 were all assigned to the two Saturday call windows because the call center was closed on Sundays in wave 6. If the contact attempt at the prespecified call window was unsuccessful, the next call was scheduled for the next week on the same day during the same window. After three unsuccessful contact attempts, treatment cases were switched to the standard protocol and were called in a similar fashion to the control cases (see description of the algorithm subsequently).
The control group algorithm randomly assigned a starting window to a case (morning, afternoon, and evening). If a case cannot be contacted during that window, he or she moves to the next window (morning to afternoon, afternoon to evening, and evening to morning). Days of the week are ignored in the assignment. The minimum time before a case is called again is 240 minutes, with the exception of (1) busy cases, who were called again within the next 15 minutes; (2) cases with whom an appointment was made; and (3) cases who had a terminating status during the call (ineligible, refusal, etc.). This control group algorithm is commonly used by Infas, the data collection agency currently in charge of the PASS data collection.
Despite the compressed schedule, the treatment group did not have fewer opportunities to be contacted. Looking at the overall contact rates between treatment and control group, we found no significant difference (with an average contact rate of 0.95 in the treatment group and 0.96 in the control group). The average number of attempts among cases ultimately classified as “noncontact” was 19 in the control and 24 in the treatment group, but the difference is not statistically significant.
Analysis
In our analysis, we focus on two outcome variables: (1) a binary variable indicating whether a sampled case was interviewed at the first contact as a measure of immediate cooperation; and (2) the number of calls until first contact. Both increased probabilities of being contacted and fewer contact attempts would imply increased efficiency and thus potentially benefit fieldwork agencies.
To assess the effects of the experimental change in the call-scheduling algorithm, we first look at these outcomes, differentiating simply by assignment status to treatment and control group. In the literature on evaluation of randomized experiments, this difference by assignment status is also known as an intention-to-treat (ITT) effect. Formally, the mean difference in terms of the outcome variable Y of those assigned to the treatment group (Z = 1) and those assigned to the control group (Z = 0)
2
is:
Since in our context “participation” is voluntary among those randomly assigned to receive treatment, there is the issue of compliance with the experimental intervention. That is, what we have manipulated here by concentrating first calls in the same time window as the previous wave’s interview is just an “offer” of treatment. In many experiments—and here, too—it is possible that the treatment assignment is not equal to the actual treatment intended, which would be a successful contact in the designated time window. There are several reasons for this. First, some control group cases might have been called and contacted at the same time window by mere chance. Second, individuals in the treatment group are free to ignore a call or may simply be unavailable at the assigned calling time. That is, of those offered treatment, only a—potentially self-selected—subset is successfully contacted (treated) at the designated time.
That said, the ITT effect is indeed the policy-relevant parameter. It tells us the causal effect of the offer of treatment, building in the fact that many of those made an offer have declined (Angrist and Pischke 2009:161–66). However, the ITT effect is too small relative to the average causal effect on those who were, in fact, treated. Therefore, an adjustment can be made to the ITT effect that gives us a measure of the effect of treatment on the treated: dividing the θITT by the difference in compliance rates between treatment and control groups; with D = 1 if the first contact was in fact in the same time window, and D = 0 otherwise.
In econometric parlance, this is called a local average treatment effect (LATE), which is the average treatment effect on the subset of individuals whose treatment status is actually changed by the treatment offer. It is equivalent to instrumental variable (IV) estimation using the randomly assigned treatment status as an IV for treatment received.
When reporting the results, we assess statistical significance at the 10% level in a one-sided test—testing if the treatment improves efficiency outcomes compared to the control. To assess sensitivities of the results, we analyzed the data in several different ways. We restricted the analysis data set to weekday cases only, assessed the effects by call windows, and changed the coding of the dependent variable such that appointments are also treated as successful outcomes for the first call attempt. The alternative analyses produced the same patterns as the results shown subsequently; however, the fine-grained analysis within windows reduced the sample sizes too much to assess results with any confidence.
Results
Starting with the results of the ITT analysis, Table 1 shows that, overall, the probability of cooperation at first contact is not increased significantly. Restricting the analysis to the first three calls—the range over which the experimental manipulation of the call scheduling was in effect in the treatment group—shows a similar result.
Comparison by Assignment Status to Treatment/Control Group; ITT.
Note: Positive ITT for Pr interview indicate higher probabilities in the treatment group. Negative ITT for number of calls indicate reduction in the number of calls for the treatment group.
*p ≤ .1. One-sided test of improved efficiency in treatment group versus control group.
Turning away from immediate cooperation toward contactability (i.e., the number of call attempts it took to establish the first contact), we again find no effect over the range of call attempts during which the experiment was in effect. We find a significant difference between treatment and control group cases with respect to the number of call attempts when we look at the full sample. Analyses not shown here indicate that there is no difference between treatment and control up to about 10 calls. Therefore, the mean difference of .364 calls until first contact (which would, scaled by the mean, amount to a sizeable 9% reduction) appears to be driven by a few cases with a relatively large number of contact attempts and is most likely unrelated to the experimental intervention. In line with our discussion in the analysis, the ITT effects just described are smaller than the local average causal effects on the treated, mainly due to noncompliance with the treatment. With our six experimental interventions, we could only shift calls into designated time windows, but not contacts.
For example, in the treatment group, only 63% had their first contact in the same time window (complied with the treatment). On the other hand, about 6% of the control group cases happened to be successfully contacted in the same time window just using the usual calling routines. Table 2 shows that even the effects on those cases actually being treated are not significantly different from zero, again, except for the number of calls until first contact in the full sample.
LATE; IV Estimates.
Note: IV = instrumental variables. One-sided test of improved efficiency in treatment group versus control group.
*p ≤ .1.
Summary and Discussion
This article set out to empirically test a suggestion put forward by Lipps in a 2012 issue of Field Methods. A random subset of panel cases was assigned to successful windows from the prior panel wave. Control cases were randomly assigned to call windows. In the experiment reported here, no efficiency gains were found for calling cases at the same day and same time window as in the prior wave.
Although the experimental implementation of the call window assignment did not yield the expected outcome, the observational analyses in PASS supported the notion that calling at the successful window from the previous wave would increase efficiency. This suggests that the significant coefficients in models for cooperation at first contact by Lipps (2012) and in our earlier analyses for PASS might reflect mainly selection effects. That is, individuals who are contacted in that time window are different in terms of unobservable characteristics that are not—or cannot be—controlled for in such analyses.
We used assignment status as an IV for the actual treatment received and were able to identify the causal effect of the same time window on the chosen outcomes and thereby assess the magnitude of selection effects. Our analysis highlights the need to carefully identify the underlying causal mechanisms when designing experimental interventions. Otherwise, alterations to existing calling routines and other fieldwork procedures will remain ineffective.
It would be a valuable contribution to this line of research if other panel surveys included similar experimental manipulations. Future research could also benefit from expanding the set of paradata used from prior waves to include patterns of attempts from several waves. Likewise, frame or registry data could be added to build more powerful models when predicting best call windows (Wagner 2013).
The maximal gain in terms of cost effectiveness for these efforts will likely vary by fieldwork organization. In the absence of precise cost data, one could estimate a lower bound of the gains achieved here by taking the hourly rate of a telephone interviewer (e.g., US$30) and assumed 1.5 minutes per call for a failed call. With an average number of four call attempts per case and an average reduction of .04 contacts attempts, this would amount to roughly US$1,800 in a 6,000-case survey. Supervisor costs and other costs of maintaining a telephone facility are not counted in this example calculation. Such savings can add up in large-scale surveys and are likely to be larger in CAPI surveys, where contact attempts are much more costly. Implementing prescribed call windows in face-to-face surveys has other challenges, however (Wagner 2013).
Footnotes
Acknowledgment
We thank Mark Trappmann for allowing us to implement the experiment into PASS, and Infas for the implementation. We also thank Russ Bernard, Stephanie Eckman, Andrew Mercer, Alexandra Birg, and three anonymous reviewers for their critical comments and their suggestions to this article, and Gabriele Durrant for her input to preliminary analyses during her research visit at the Institute for Employment Research in 2011.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
