Abstract
This special issue is devoted to discussion of probability-based survey panels that collect data either solely or partly through online questionnaires. Panels of this kind have been around for a long time, though they have been few in number, but recent years have seen several new panels start up in Europe. This has led to renewed interest in the methodology of such panels and also to deeper questioning of the role of these panels. On one hand, the probability-based panels are to some extent competing against cheaper non-probability access panels. On the other hand, the probability-based panels are increasingly being seen as possible alternatives to more expensive probability-based survey methods. In both cases, clients and data users want to better understand the relative advantages and disadvantages of the probability-based panels.
Evolution of Probability-Based Online Panels
The idea of having respondents participate regularly in computer-assisted self-administered interviews goes back to the mid-1980 (Saris, 1998). In 1986, the probability-based Dutch Telepanel encompassing 1,000 households was established, each equipped with a stationary computer, modem, and software to communicate by telephone with a central modem pool. Once every weekend, a new survey was downloaded via a temporary telephone-based dial-in connection and upon completion, the data were uploaded to the central modem pool accordingly. Although this technological procedure appears utterly outdated, the methodological solutions established to reduce survey errors and biases (Groves, 1989) are still impressive. Specifically, reducing coverage error was achieved by selecting a random sample of the Dutch population, which was then equipped with technology to participate in regular surveys. The sample size was chosen in a way to keep sampling errors within acceptable ranges. Measurement error was controlled by standardizing the visual appearance of survey instruments, and by randomizing item orders whenever deemed appropriate and feasible. Finally, nonresponse bias was reduced by implementing sophisticated panel maintenance measures. Respondents were not paid for their participation, but the provision of a computer could be considered as a serious benefit of participation—certainly in those days—reducing nonresponse bias as well.
In 2000, the Telepanel, by then moved from the University of Amsterdam to Tilburg University and renamed to the CentERpanel, switched from a telephone-based dial-in connection to communication via the Internet. Since then, cultural and technological change associated with Internet adoption have made online panel-based research a well-accepted strategy for conducting surveys in the social sciences (Callegaro et al., 2014), allowing the fielding of both cross-sectional and longitudinal studies. This development has led to the emergence of a multitude of online panel vendors, with non-probability access panels, where people select themselves into a panel, dominating the landscape. In contrast, probability-based panels, where all members of a population of interest have a known, nonzero probability of receiving an invitation to join, are still few, perhaps because of the expense of setting up the panel, which usually involves face-to-face contact and provision of equipment and Internet connection. Based on the blueprint laid down by Saris, probability-based online access panels were also emerging in the United States, of which the longest lasting are the GfK KnowledgePanel, which began recruiting in 1999, 1 and the Gallup Panel, which started in 2004. 2 These two panels took different decisions about how to achieve coverage of the part of the population that were not already online. The GfK KnowledgePanel decided to provide off-line sample households with equipment (since 2009, this has consisted of a windows-based laptop) and an Internet connection, while Gallup decided to use postal survey methods with these households, thus creating a mixed-mode panel. The first of these two solutions has a much higher initial set-up cost, while the second has higher data collection costs for each survey carried out, and results in slower data collection and mixed-mode data.
Building on the success of the CentERpanel, a new (and larger) probability-based online panel was set up in the Netherlands in 2007, that is, the Longitudinal Internet Studies for the Social sciences (LISS) panel.
3
More recently, in 2012 and 2013, three other panels have commenced, two in Germany and one in France. Two of these panels have chosen to provide off-line sample members with equipment and connections, one is a mixed-mode panel using a combination of postal and online methods, and the fourth has provided all sample members with standard equipment. The article by Blom et al. (
Survey Errors in Online Panels
The main sources of statistical error in the data collected by probability-based online panels are no different from any other survey, namely, coverage, sampling, nonresponse, and measurement (Groves, 1989). However, the issues that determine the nature of each of these error types are quite distinct, as are the methods and procedures that can be used to attempt to reduce the errors.
With respect to coverage, the key challenge is to find an effective and cost-efficient way to include the off-line population. As the size and nature of the off-line population is changing over time and differs between countries (Mohorko, De Leeuw, & Hox, 2013), any decision about the most appropriate cost-error trade-off may not necessarily be generalizable. The approaches outlined in the previous section, of either providing equipment to sample members or using a mixed-mode design, are designed to address this issue and thereby avoid the coverage error that would be incurred in the off-line population were to be excluded from the panel. However, a fundamental question, addressed by Eckman (2016, in this issue), is whether the efforts made to include the off-line population can be justified in terms of error reduction.
A key issue that affects both sampling and nonresponse errors is the absence of high-quality general population sampling frames that include both names and e-mail addresses (Lynn, 2013). In consequence, it is usually necessary to select samples from frames that include one or more of name, address, and phone number—but rarely all three. The frame information tends to drive the mode of initial approach. If this needs to be face-to-face, then the sample is likely to be geographically clustered for cost-effectiveness reasons, thereby ruling out one of the potential error advantages of online surveys—reduced sampling variance through the use of an unclustered sample. The constraints imposed by available national sampling frames therefore loom large in any discussion of costs of errors for probability-based online panels. On the other hand, sampling bias can usually be assumed to be negligible in probability-based online panels, in contrast to non-probability access panels. Statisticians have proposed a number of methods to overcome this disadvantage of non-probability panels, including sample matching, which is evaluated by Bethlehem in his article (2016, in this issue). Most of these methods amount to some form of either quota sampling (imposing a structure on the unweighted sample) or population weighting (imposing a structure on the weighted sample), and it is as yet unclear whether they can be relied on to deliver similar statistical properties to a true probability sample.
Whether or not the initial approach to sample members is interviewer-administered may have an important impact on response rates and nonresponse bias at the recruitment stage, as may other design considerations including whether off-line sample members are being asked to accept and use new equipment and what kind of incentive is provided for participation. Furthermore, once recruited, obtaining high levels of participation in each survey is an ongoing challenge. However, the probability basis of these panels at least provides a context in which it should be possible to assess nonresponse bias in fairly robust ways. Currently, attempts to compare outcomes between surveys using different methods are, however, hampered by the lack of a standardized approach to recording and summarizing outcomes (DiSogra & Callegaro, 2016, in this issue). Furthermore, even when panel members participate in surveys they may not answer all questions. This brings with it a risk of item nonresponse bias. Methods for minimizing item nonresponse in online surveys are evaluated in the article by De Leeuw, Hox, & Boevé (2016, in this issue).
With respect to measurement error, surveys conducted on probability-based online panels are much like other self-completion surveys. Thus, to obtain valid and reliable measurement, researchers would ideally design questions to be engaging, understandable, easy-to-answer, nonsensitive, and nonthreatening (De Leeuw, Hox, & Huisman, 2003). However, there are a couple of additional concerns. One is that respondents to online panels may become “conditioned” by repeated participation in similar surveys and consequently that their responses may no longer reflect the broader population that the sample is designed to represent. This concern is addressed by Struminskaya (
Future Research Agenda
As probability-based online panels are still few in number and relatively new, methodology is still evolving. Different panels have made different decisions regarding important aspects of design and implementation and it is unclear which methods are best. Several methodological challenges still exist. For a large-scale shift to occur away from probability-based methods using other modes, particularly interviewer-administered modes, to online panels, we believe that it will be necessary to demonstrate that the error structures associated with online panels are broadly similar to, or better than, those that are achieved in other modes. On the other hand, for survey clients to choose probability-based online panels rather than their cheaper non-probability access panel cousins, they need to be persuaded that errors are likely to be substantially smaller. For both of these reasons, a body of research evidence on the relative merits of different methodological features is required.
We hope that this special issue makes a useful contribution to that evidence. The call for papers for this issue focused on all aspects of survey error, including coverage, sampling, nonresponse, and measurement, the aim being to present recent research on the nature of survey errors in probability-based online and mixed-mode panels, and on the effectiveness of methods and procedures to reduce those errors. We think we have achieved that aim, and the article authors are to be congratulated on collectively providing a very useful set of insights. But of course, this collection does not provide all the answers to the questions that the research community will have about probability-based online and mixed-mode panels. We hope that the articles in this special issue will stimulate further discussion of the issues and new research to address remaining gaps in the knowledge base.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
