Abstract
Evaluating the quality of the data is a key preoccupation for researchers to be confident in their results. When web surveys are used, it seems even more crucial since the researchers have less control on the data collection process. However, they also have the possibility to collect some paradata that may help evaluating the quality. Using this paradata, it was noticed that some respondents of web panels are spending much less time than expected to complete the surveys. This creates worries about the quality of the data obtained. Nevertheless, not much is known about the link between response times (RTs) and quality. Therefore, the goal of this study is to look at the link between the RTs of respondents in an online survey and other more usual indicators of quality used in the literature: properly following an instructional manipulation check, coherence and precision of answers, absence of straight-lining, and so on. Besides, we are also interested in the link of RT and the quality indicators with respondents’ auto-evaluation of the efforts they did to answer the survey. Using a structural equation modeling approach that allows separating the structural and the measurement models and controlling for potential spurious effects, we find a significant relationship between RT and quality in the three countries studied. We also find a significant, but lower, relationship between RT and auto-evaluation. However, we did not find a significant link between auto-evaluation and quality.
Introduction
Web surveys offer both new challenges and new opportunities. On one hand, the absence of interviewers makes the control of the situation and of what the respondents are doing more difficult. Respondents can answer for another person. They can do several tasks at the same time. They can complete the survey in any place, at any time, with or without others persons being present. They can randomly answer the questions without even reading them. There is nobody to see it and stop it.
On the other hand, web surveys offer rich possibilities to collect paradata, meaning data generated automatically by the data collection process that can be used as additional information to describe or evaluate this process (Couper, 2000). Even if nobody is there to observe what the respondents are doing, it is possible to record many kinds of information that can help the researchers in identifying the respondents’ behaviors during the survey completion. For instance, it is possible to track the movements of respondents’ eyes, in order to see whether they looked at all the information on the page. It is also possible to record the movements of the mouse, which can not only indicate hesitation or direct answering but also record all the clicks the respondents did, which allows seeing whether respondents change their answer or go back in the survey. Moreover, it is possible to detect whether the respondents went out of the survey to open another Internet window, indicating that the respondents did some other tasks on Internet.
In this study, we want to focus on one specific kind of paradata: the response times (RTs) or response latencies. RTs are defined as the times between the moment when the page starts loading and the moment when respondents click on the “Next” button.
Even if there is a lot of potential paradata available, a lot of them are difficult to analyze in a systematic way that could be used to learn more about the response process. RT seems to be one of the most accessible ones.
Moreover, in practice, it was noticed that some respondents of web panels are spending much less time than expected completing the surveys. This fast responding, also referred to as speeding, creates worries about the quality of the data collected through these web panels. By answering surveys regularly, panelists get practice, so we can expect them to be quicker than nonpanelists. However, part of them are spending such little time completing the surveys that it is nearly impossible that they have read the questions, even more that they have thought carefully about the answers before selecting them. Consequently, some survey institutes commonly apply as a quality rule the exclusion from the final data set of the panelists who did the survey in less than x% of the expected time of completion.
One limit is that different respondents need different times to properly answer the questions. Some may be quicker than others in reading, making up their mind, and finding the adequate answers. Yan and Tourangeau (2008) find that age and education, as well as Internet skills, affect RT.
It is clear that when RTs are extremely short, it is simply impossible that the respondents read the questions before answering. However, when RTs are quick but not so extreme, we do not know whether this is because the respondents are highly skilled and therefore they are able to answer quickly while going through all the necessary cognitive steps (the four steps described by Tourangeau, Rips, & Rasinski, 2000) or whether the respondents are just rushing through the questionnaire without reading and thinking properly.
Even if the idea that speeding is an indicator of low quality is quite common, there is little research about how RTs and quality relate to each other. Malhotra (2008) looks at the link between completion time and primacy effects (i.e., bias toward selecting earlier response choices). He finds that for low educated respondents, shorter RTs are associated with higher levels of primacy effects. Zhang (2013) considers the link between RTs and straight-lining, an extreme kind of nondifferentiation. Straight-lining consists in selecting the same answer for all the items in a battery. The author finds that the respondents who speed more also tend to straight-line on more grid questions, suggesting that the tendency to speed is indeed related to the quality of answers.
Following this line of research, we want to go further in the investigation of the link between RTs and quality, by considering more indicators of quality and also adding the auto-evaluation of the respondents about the efforts they made to answer the survey.
More precisely we want to investigate: How strong is the link between RTs and quality? How strong is the link between quality and auto-evaluation of efforts done? How strong is the link between auto-evaluation of efforts and RTs?
We hope the answers to these three questions will give us some indication to start answering a very general and practical question: Can we use paradata to improve the overall quality of online surveys and how? More specifically, we focus on one kind of paradata, RTs, so we can wonder whether we can use RT as a proxy for quality. Said differently, can we use RT to decide which respondents we should exclude from our analyses? Besides paradata, can we use auto-evaluation of the respondents for the same goals?
The next section gives some details about the data used to study these questions. Then, the structural and measurement models are presented and combined to propose a complete structural equation model (SEM). After that, we explain how the analyses were done and how we corrected the initial model in order to get to the results that are finally presented and discussed.
The Data: A Survey From Netquest in Spain, Mexico, and Colombia
We used data from the web panel Netquest (www.netquest.com), which is accredited with the ISO 26362 quality standard, specific for online access panels. Netquest uses a database of users of many websites that agreed to receive e-mails from one of these websites. From this database, they invite people with the profile they need to participate in their panel. For each survey completed, panelists get points that they can exchange for gifts. The number of points is proportional to the expected length of the survey but not the actual completion time of each respondent. Therefore, speeding can look advantageous to the panelists since it allows them to get the same amount of points in less time. If respondents’ participation is mainly drawn by the incentives, “the purely rational approach is to satisfice” (Malhotra, 2008, p. 915), meaning to “minimize effort in responding to surveys and simply provide the appearance of compliance (Krosnick 1991; Krosnick & Alwin 1987).”
The survey selected has been proposed to panelists in Spain, Mexico, and Colombia between May 14, 2013, and June 18, 2013. One advantage is that it includes questions that can be affected by different kinds of satisficing behaviors. Another is that it includes an auto-evaluation of the efforts made and of the difficulty of the questions.
The survey was expected to take around 25 to 30 min to be completed (around 125 questions) and was about various topics: mainly drinks and food consumption, brands of cars, and media use.
Quotas for age and gender were used to get samples’ distributions similar on these variables to the population distributions. In each country, around 1,000 respondents completed the survey. However, an experimental design was used such that each sample was randomly split-up into three groups. For this study, we focus only on the first group of each country, in order to have similar questionnaires for which we can compute the same quality indicators. At the end, we have 345 respondents in Spain, 305 in Mexico, and 336 in Colombia.
The Structural Model
Our main interest is to investigate what are the links
1
between the RTs, the quality of the answers, and the auto-evaluation of the efforts made. Our main hypotheses are:
In order to estimate these relationships properly, we need to introduce as control variables the ones that we expect will create spurious relationships between our main variables of interest. These variables are, first, age and education. Indeed, we expect them to affect RTs (e.g., following the results of Yan & Tourangeau, 2008) and to affect the quality of answers (e.g. following the results of Alwin & Krosnick, 1991). Then, we also include an auto-evaluation of the difficulty of the questions as control variable: If the questions are easier, the quality should be better and the RTs shorter.
The Measurement Model
RTs, quality, and auto-evaluation are theoretical concepts that are not directly observed. The indicators used to operationalize them are presented in this section. Because there are errors, the indicators are never perfect measures of the concepts of interest, but by using several indicators for each latent concept, we can correct for measurement errors (Saris & Gallhofer, 2007). This is the approach followed in this article.
Indicators of RT
RTs are paradata. As explained for instance by Yan and Tourangeau (2008, p. 53), they “can be collected or from the server or from the respondents’ computers. Server-side response times show the elapsed time from the moment the server delivers a survey question to a respondent’s computer to the moment when it receives an answer from the respondent. By comparison, client-side response times include the elapsed time from when a survey question is fully displayed on respondent’s computer to when an answer is sent.” In our study, we use server-side RTs (the only ones we could get access to), which therefore include the downloading time. By consequence, if some respondents use quicker Internet connections than others, we may observe differences in RT that do not reflect really the differences in the time spent to process the question, decide, and select an answer.
Besides the downloading time, another issue of the measure of RTs is that these times correspond to the time spent on a page. However, if a respondent let the survey on and start doing other activities (e.g., speak with another person, go to the kitchen to check the diner, go to the bathroom, etc.), the time spent on the page will not be a good measure of the time spent to answer the question. It will be much longer. Detecting the multitasking behaviors is quite difficult. Nevertheless, in the case of very long RT for one page, we can be pretty sure that respondents have interrupted the survey process to do another task and came back later.
Therefore, in order to compute the RT of each respondent, for each page, we substitute the times of the 1% respondents with the highest time 2 (considered as the ones that clearly were multitasking) by the average time spent by the other 99% to answer to the questions on that same page. We substitute by the average time and not the maximum time of the other 99% because we believe that the very long times do not indicate extremely slow respondents but respondents who interrupted the survey. 3 Then, there is no reason to expect these respondents to spend a long time on the page once they come back to it.
Still, the measures of response times cannot be expected to be free of errors. For instance, if respondents are interrupting the completion of the survey for a short time, this may not be detected and taken into account. Therefore, in order to correct for measurement errors in our final model, instead of using only one indicator of RTs, we use three. Each one corresponds to the average RT for a set of questions from the surveys. Together, the three sets of questions constitute the complete survey. This is a crucial step to get correct estimates of the relationship between RTs and other variables. So our three indicators of RT (corrected for very long times) are, for i = 1, 2, 3:
In Spain, the mean RTs for the three sets of questions are, respectively, 0.23, 0.17, and 0.18 min. In Mexico, there are, respectively, 0.26, 0.20 and 0.20 min and in Colombia, 0.29, 0.22, and 0.23 min. Part of the differences could come from differences in average connections’ speed across countries.
We have to note that, in general, there is only one question per page, but when questions are part of a battery, there are several questions on the same page. Therefore, when there are batteries of questions, we can expect the RT to be higher.
Indicators of Quality
In order to measure the quality of the responses, we use different traditional indicators and some a bit more specific, but all are directly derived from the answers to the questionnaire (“direct” data and not paradata anymore). We will now go through each of them.
Instructional manipulation check
First, we use an instructional manipulation check (IMC), which “consists of a question embedded within the experimental materials (…) that asks participants (…) to provide a confirmation that they have read the instruction” and is supposed to measure “whether or not participants are reading the instructions and thus provides an indirect measure of satisficing” (Oppenheimer, Meyvis, & Davidenko, 2009, p. 867).
In our study, the IMC was included within a grid about media where respondents had to select the five most important options of a list of 18, and this in three different cases. The IMC consisted of one additional row in this grid asking the respondents, whether they were reading, to mark all three buttons on this row, besides the five most important options. This additional row was placed in the middle of the substantive alternatives, always on the same position.
The variable summarizing the results is called “passIMC.” It is a dummy variable that takes the value 1 if the respondents correctly achieve the manipulation asked (i.e., checked the three boxes) and the value 0 otherwise.
In our sample, only 21.67% of the respondents did it correctly: 30.0% in Spain, 21.9% in Mexico, and 14.1% in Colombia. There are clear differences across countries but overall, these are very low percentages that can make us worry about the quality of the data. But we should notice that this IMC was part of a really complex and badly designed grid. This grid was actually used in this survey because it was such a badly designed one, so it was part of an experiment aiming to improve it. Therefore, even respondents who are not by nature “bad” can be tempted in this context to take shortcuts. It is also possible that some respondents did not understand exactly what they had to do in a context of such a complex grid.
Moreover, as Berinsky, Margolis, and Sances (2013, p. 2) mentioned the IMC “are not immune to measurement errors.” Their suggestion is to use multiple measures. Here, we do not use multiple IMCs, but we indeed use multiple measures for quality in order to correct for measurement errors.
Number of selected items in the media grid
The same grid was used to compute a second quality indicator, which is the selected number of items in the three questions of this grid. As mentioned earlier, respondents should choose the five most important options in each of the three questions, apart from the boxes in the row of the IMC.
In such situations, web survey allows using an automatic check to be sure that respondents comply with the requirement of choosing exactly 5 items in each question. However, the automatic check was not applied here in order to see how many respondents will not at first be able to carry out the task according to the instructions.
The information is summarized in a variable called “select5.” For each respondent, it scores from 0 (did not select five options in any of the three questions) to 3 (selected five options in the three questions). Table 1 gives the distribution of this variable in the different countries. Table 1 shows some variations across countries, but overall the percentage of respondents who really respected the instructions is only between 32% and 37%.
Proportions of Respondents Who Selected Five Options as Requested.
Nondifferentiation
The next indicator of (bad) quality is nondifferentiation between items. It usually occurs in grids with many items. In its extreme form, it means respondents select the same answer category for all items in a given battery. This is referred to as straight-lining. This is one of the most used indicators for satisficing in surveys (Green & Krosnick, 2001; Zhang, 2013). However, it is also possible that respondents select the same category in almost all items (all except one or all except two). Therefore, some authors (Couper, Tourangeau, Conrad, & Zhang, 2013; Krosnick & Alwin, 1988) consider not only pure or full straight-lining but also near straight-lining. Besides, they propose to look at the variance in each respondent’s answer to the list of items in a grid. A larger variance indicates a lower nondifferentiation level.
The Netquest survey included two grids of 16 items about the frequency of consumption of several drinks and one grid of 12 items about opinions about brands of cars.
For each grid, we look whether the respondents selected the same answer category for all items. If they do, then they are considered as pure straight-liners for this grid. Combining the scores for the three grids, we create the variable “nbstraight,” which counts in how many of the grids the respondents are pure straight-liners.
Table 2 shows again some differences across countries, with Colombia performing better. Overall, pure straight-lining is not present as the previous undesirable behaviors. However, the questions for the two grids about drink consumptions were not very demanding since they were about central behaviors. Also, we only counted the pure straight-liners. We can also consider the variance in each respondent’s answers to the list of items in each grid as suggested by Krosnick and Alwin (1988). Then, to summarize the information, we take for each respondent the sum over the three grids of the variances in answers. Table 2 gives the average over all respondents of this sum of the variances. We can see that the higher level of differentiation is found for Spain now, followed by Colombia, Mexico coming last. However, it is more difficult to interpret the variances in answers. If pure straight-lining clearly indicates satisficing, a low variance may happen because of the drinking habits of a respondent. A respondent drinking only water (every day) but nothing else will have a very low variance even if he or she is answering the questions very carefully. Since our goal is to measure really the bad quality respondents, we therefore use straight-lining in the main analyses. Nevertheless, as a robustness check, the analyses are repeated using the variances. The main results are not affected (see Appendix A).
Proportions of Respondents Who Are Pure Straight-Liners in 0–3 Grids and Average of the Sum Over the Three Grids of the Respondents’ Variance of Answers.
Incoherence of responses between repeated or opposite items
Then, we consider as an indicator of quality the incoherence across responses for repeated questions or opposite items. Indeed, the two grids about drink consumptions mentioned before are similar except that the scale is reversed. One is asked at the beginning of the questionnaire, and the other at the end. The drink consumption of respondents could not have changed in between the two sets of questions, so the differences can be interpreted as incoherence in answers. Moreover, we have three pairs of opposite items (with one item worded positively and one negatively), such that if respondents agree with one, they should disagree with the other one if they are coherent in their answers. For example, “[name] is a trustworthy brand” and “[name] is a brand in which I have no trust.”
For each of these questions, we compute how incoherent the responses are in the following way: If the respondents selected the same category twice, it is not incoherent at all, so they got a score of 0 for the corresponding question. If the second answer is one category next to the expected answer according to the first question, there is a small incoherence and they get a score of .01 for this question. If the second answer is n categories next to the expected answer according to the first question, there is an incoherence of level n and they get a score of .01 × n for this question. Finally, we sum up all these scores in the variable “incoherence” (from 0 to 1.08).
The percentage of respondents who answered in a perfectly coherent way (i.e., incoherence = 0) is only 3.6% overall, with big differences again between countries: in Spain, it is much higher (9.9%) than in Mexico and Colombia (both 0.3%).
The percentages of respondents whose answers varied in average no more than one category (i.e., incoherence ≤ .19) are 86.6% overall, 91.9% in Spain, 87.2% in Mexico, and 80.6% in Colombia. Spain is still performing better, but now Mexico is doing better than Colombia.
Precision of answers in open narrative questions
In the case of open narrative questions, other indicators of quality are available. First, the precision of the answers can be estimated via the number of characters the respondents wrote. Respondents who do not want to make efforts will indeed tend to write less.
In our study, two open narrative questions are present. The first one asks respondents what they would do to improve the survey and the second one asks them to indicate all the topics they can remember from the survey.
The number of characters written in both questions is summed up to create the variable “charac.” All countries together, the number of characters written varies from 2 to 621, with an average of 120. In Spain, it goes from 2 to 613, with an average of 111. In Mexico, it goes from 6 to 621, with an average of 125. Finally, in Colombia, it goes from 8 to 502, with an average of 125.
Nonsense in open questions
In open narrative questions, in order to reduce the efforts, besides writing short answers, respondents can also write “easy” answers: just a few letters or signs that have no sense, some unrelated text that does not answer the question or simply “don’t know.”
We create a variable “nonsense” that counts how many of these undesirable answers we get (0, 1, or 2 since we have two narrative questions). Table 3 gives the percentages.
Proportions of Respondents Who Wrote a Nonsense 0, 1, or 2 Times in the Open Questions.
Table 3 shows that the percentages of nonsense are relatively low. Overall, almost 94% of the respondents did not write any. Still, a small percentage did and when they did, it indicates a bad quality of the answers.
Indicators of the Auto-Evaluation of the Efforts Done
The previous indicators of quality are based on the idea that if respondents are doing the necessary efforts, the quality will increase, whereas if they satisfice, the quality will decrease. Therefore, the efforts done seem to be a key variable. But who knows better the efforts they did than the respondents themselves?
In our study, the following question was asked: “how much efforts did you put in answering this survey?” on a scale from 0 = minimum effort possible to 10 = maximum effort possible.
Even in an online survey, this question is susceptible to social desirability bias, that is, overreporting of socially desirable behaviors (here having performed the maximum efforts) and underreporting of the undesirable ones. The mean for this variable is 6.6 in Spain and Colombia and 6.8 in Mexico. This is quite high, maybe because some respondents do not need to make effort to answer properly since answering questions is an easy task for them. But it can also be because of social desirability bias. Still, 30.14% of the respondents reported a level of efforts lower than 5 in Spain, 33.01% in Mexico, and 34.02% in Colombia. Are these respondents also the ones with low quality according to our different indicators?
This is what we want to find out. If the efforts reported are highly correlated with the quality measured by the indicators specified in the previous section, then the easiest way of assessing response quality would be to ask directly to the respondents.
In order to correct for measurement errors, since we have only one available indicator for this latent concept, we need to get an estimate of the quality of the question from an outside source (Saris & Gallhofer, 2007). We use the program SQP 2.0 (Saris et al., 2011) to predict the quality of this question (available for free at http://www.sqp.nl/) and get a value of .538. We use the same value in all three countries since they share the same language.
Indicators for the Variables of Control
For age, we use the answers to a direct question asking the age of the respondents. For education, different answer categories are used in the three countries. But Netquest also provides a harmonized variable ordered from no or low education to high education (six categories). We use this harmonized variable.
For the last control variable, “easy,” we use the question, “How do you feel about the questions of this survey?” on a scale from 0 = extremely difficult to answer to 10 = extremely easy to answer.
Since we have only one indicator for each, we need estimates of the quality of these questions in order to correct for measurement errors. For the variable “easy,” using the program SQP 2.0, we obtain a quality estimate of .613.
Alwin, 2007, p157, Table 7.4 “Estimates of reliability of self-reported measures of facts by topic of question”. We use the estimates for age and education provided in this table, which are respectively .997 and .884 for education. These estimates have limits since these are not based on data from the countries of interest and since the education variable is not measured in the same way in our study than in the ones of Alwin. However, this is the best we can get for these variables. Besides, for age, the size of the errors is so small that in any case it will not change the results. But for education, it is more realistic to use Alwin’s estimates than to assume that there are no measurement errors.
General SEM
A Preliminary Analysis
Before considering the general SEM, we briefly consider some bivariate descriptive statistics. Figure 1 gives, for each country, the plots of the different indicators of quality as a function of the average RTs (box plots for the indicators of quality with few values and scatterplots for the others).

Plots of the different indicators of quality in function of response times (RTs).
We can see that high levels of incoherence and straight-lining are mainly found for respondents with short RTs in all countries. In Spain, respondents with more nonsense and those who never selected the five items in the grid also have shorter RTs. To some extent, this seems to hold in the other countries too. Very quick respondents write fewer characters, but slow respondents also write quite few so no clear pattern appears for this indicator.
However, little can be concluded from the descriptive statistics. Using only separately the indicators of quality is not enough. All contain errors so we should combine them in order to get a better overview of the quality of answers of the respondents and to be able to correct for measurement errors. As mentioned earlier, RTs are also not measured without errors so we should also use several indicators there. Finally, we should include control variables in order to correctly estimate the correlations between our main variables of interest. This is what we do in the next sections.
Initial SEM
A path diagram of the complete initial model can be found in Figure 2. It is a combination of the structural model and the measurement model discussed previously.

The complete structural equation model (SEM).
For the sake of simplicity, the control variables are treated as observed ones in the model. In order to correct these three variables for measurement errors, we use the reduction in variance technique, putting the quality on the diagonal of the correlation matrix. For the two latent concepts quality and RT, the correction is done by using several indicators. For the auto-evaluation of the efforts, the loading between the latent variable and the observed one is fixed to the quality coefficient, which is the square root of the quality prediction from SQP mentioned earlier, that is, √.538 = .734. The error variance is fixed for this indicator to 1 − .538 = .462.
Analyses and Corrections of the Model
The maximum likelihood estimation of LISREL (Jöreskog & Sörbom, 1991) for multiple group analyses is used to get the results. Each country constitutes a different group. We first specify all parameters invariant across countries. Then, the ones that are misspecified are allowed to be free.
The testing is done using both global fit measures (chi-square, root mean square error of approximation [RMSEA], and comparative fit index [CFI]) and local fit measures (looking at the Expected Parameter Changes, Modification Indices, and Power) using the program JRule (Van der Veld, Saris, & Satorra, 2009) based on the procedure developed by Saris, Satorra, and Van der Veld (2009).
The initial model is corrected step by step by following the suggestions. First, we had to introduce correlations between some error terms (in one or more countries): between “select5” and “passIMC,” between “nonsense” and “charac,” between “incoherence” and “nbstraight,” between “passIMC” and RT1 and between “charac” and RT2. The first three can be expected because pairs of indicators are based on the same questions, so it makes sense that they correlate more. The last two also makes sense because the IMC is only included in the set of questions used to compute RT1 and both open questions are part of the set of questions used to compute RT2.
Second, cross-country differences are suggested for some correlations between control variables and spurious effects. It seems acceptable to think that indeed education and age, for instance, correlate differently in the different countries. It seems also reasonable that these control variables can affect differently RTs or quality. For example, similar levels of education may in practice correspond to different levels of knowledge in the different countries, which can explain different sizes of the effects. Therefore, we allowed cross-countries variations in the size of the spurious effects and correlations across control variables when JRule indicated it.
Results
By correcting the model as just indicated, we get an acceptable fit, according to both the global fit measures, χ2(220) = 302.37; RMSEA = .032; CFI = .96, and the local fit indicators (JRule did not suggest any big misspecification anymore). The final LISREL input is provided in Appendix B. Table 4 gives the completely standardized estimates for the general SEM as well as the unstandardized ones and the corresponding t values. 4
Standardized and Unstandardized Estimates in the Three Countries.
Note. CFI = comparative fit index; comp. stand. = complete standardized; unstand. = unstandardized; df = degrees of freedom; RMSEA, root mean square error of approximation; RT = response times. When the estimates are similar for all three countries, they are presented only once. There is a star next to the estimates when the t-value indicates that the coefficient is significantly different from 0 when a significance level of 5% is used.
Table 4 contains a lot of information. But we focus on the main results. First, about the measurement part of the model, we see that the estimates are similar across countries. Even if they are all significantly different from zero, none of the indicators of quality are very good since all loadings are quite low. This may be because here the error terms include not only the measurement errors but also the unique components.
Nevertheless, if we would have to select fewer indicators of quality, the straight-lining seem to be the strongest one (standardized loadings between .47 and .55), followed by “incoherence” (standardized loadings between .41 and .48). On the contrary, the IMC, which is used in some companies as the unique measure of quality (respondents that fail the IMC are sometimes immediatly excluded from the final sample), has loadings of only −.24 to −.20. Using the IMC to exclude “bad respondents” seems therefore not very appropriate. However, we should mention that this may be due to the specific IMC used (in a very complex grid, such that most respondents failed it). Another choice of IMC may have led to a higher loading. But Berinsky et al. (In Press) also found that using a single IMC is not a very effective way of separating bad and good respondents. They suggest to use more than one IMC.
Now, about the structural part of the model, we can notice that the three main relationships of interest are similar across countries. The standardized correlation between quality and RTs varies from −.22 to −.27, depending on the country. It is significantly different from zero. The one of quality and auto-evaluation of the amount of efforts done is around −.05 (not significant) and the one of auto-evaluation and RT is around .15 (significant).
What is changing is the size and sometimes even the signs of the spurious effects. Nevertheless, most of the effects of the control variables are not significantly different from zero. There are less spurious effects than we expected. The nonsignificance can be linked to the relatively small sample size (around 330 in each country). But overall the results suggest that the spurious effects created by “easy” and education are relatively limited. For age, more significant effects are found.
Discussion
Our main interest was to investigate the links between RTs, quality, and auto-evaluation. By developing a complete SEM with a measurement and a structural part, we were able to estimate the size of the different relationships, controlling for possible spurious effects and correcting for measurement errors. Based on the results, we can answer our three introductory questions.
First, is there a link between RT and quality? Yes, there is. How strong is this link? The size is around −.25. This means that our first hypothesis is confirmed: a worse quality of answers is directly related with shorter RT, that is, with more speeding.
Second, what about the link between quality and auto-evaluation? The analyses suggest that there is no significant relationship between them. A worse quality does not relate directly with less reported efforts. We do not find support for our second hypothesis. This can be due at least partly to social desirability bias, which pushed even bad respondents not to report the little efforts they did.
Third, how strong is the link between auto-evaluation and RT? Table 4 gives a standardized effect of around .15, statistically significant at the 5% level. This supports our third hypothesis: more reported efforts go together with longer RT. It may be because more reported efforts imply more real efforts (even if the relationship is not perfect) and more efforts lead to the respondents taking more time for answering. It can also be because if respondents take more time, they will give better quality answer. However, the size of the effect is small.
All in all, which practical advices can we derive from these results? Can we use paradata, and in particular RT, to improve the overall quality of online surveys? Can we use auto-evaluation of the respondents for the same goal?
The relationships with auto-evaluation are really small or even not significant. Whatever the reasons, it does not seem to be possible to use the auto-evaluation of efforts as a proxy of quality and therefore also not to decide which respondents did a “serious job” and which not based on this variable. More research would be needed to check whether by dealing with some of the limits our analyses are facing it is possible to get stronger relationships, but based on Table 4, we have to conclude that we should not use the auto-evaluation of the efforts to decide which respondents to exclude.
Concerning RT, our results suggest there is a stronger relationship. Nevertheless, the standardized coefficients of around −.25 are not enough to conclude that we can use RT as a proxy of quality. RT can be used, as it is already done in practice by many companies, to exclude some clear speeders. But RT is very imperfectly linked with quality so we should not give them too much importance by themselves. However, they may be used in combination with a few other indicators (for instance, incoherence and straight-lining that appear to be the indicators with the highest loadings) in order to identify a set of respondents who overall give bad quality answers. But further research in that direction would be needed to define exactly how this could be done or which cutoff points could be used.
Footnotes
Appendix A
Estimates Using the Sum of the Variances Instead of Straight-Lining.
| Estimates | Compl. Stand. | Unstand. With t-Value in Parenthesis | |||||
|---|---|---|---|---|---|---|---|
| Spain | Mexico | Colombia | Spain | Mexico | Colombia | ||
| Structural model | Bad with RT | −.21 | −.19 | −.21 | −.07* (−4.05) | ||
| Bad with AE | .00 | .00 | .00 | .00 (−.06) | |||
| AE with RT | .14 | .15 | .14 | .12* (3.19) | |||
| Easy on Bad | −.19 | −.47 | −.18 | −.07* (−3.31) | −.21* (−5.33) | −.07* (−3.31) | |
| Easy on RT | .04 | .02 | .04 | .03 (.89) | .01 (.27) | .03 (.89) | |
| Easy on AE | .00 | .00 | .00 | .00 (−.05) | |||
| Educ on Bad | .06 | .06 | .06 | .02 (1.41) | |||
| Educ on RT | −.14 | −.05 | −.15 | −.12* (−3.63) | −.04 (−.82) | −.12* (−3.63) | |
| Educ on AE | .04 | .04 | −.11 | .04 (.83) | .04 (.83) | −.11 (−1.43) | |
| Age on Bad | −.23 | −.29 | −.32 | −.09* (−2.99) | −.13* (−4.82) | −.13* (−4.82) | |
| Age on RT | .27 | .27 | .27 | .22* (7.72) | .22* (7.72) | .22* (7.72) | |
| Age on AE | −.01 | −.01 | −.20 | −.01 (−.15) | −.01 (−.15) | −.20* (−2.61) | |
| Measurement model | Bad by Inco | .39 | .44 | .40 | 1 | ||
| Bad by Nonsens | .11 | .13 | .11 | .29* (2.53) | |||
| Bad by Charac | −.19 | −.21 | −.19 | −.48* (−3.73) | |||
| Bad by SumVar | −.62 | −.67 | −.63 | −1.55* (−6.30) | |||
| Bad by Select5 | −.16 | −.18 | −.17 | −.41* (−3.31) | |||
| Bad by PassIMC | −.21 | −.24 | −.22 | −.54* (−3.87) | |||
| RT by RT1 | .81 | .80 | 1 | ||||
| RT by RT2 | .79 | .78 | .95* (25.48) | ||||
| RT by RT3 | .82 | .81 | 1.00* (25.62) | ||||
| Global Fit | χ2(df) = 354.17 (220); RMSEA = .043; CFI =.94 | ||||||
Note. CFI = comparative fit index; comp. stand. = complete standardized; unstand. = unstandardized; df = degrees of freedom; RMSEA, root mean square error of approximation; RT = response times. We should notice first that the fit of the model was much better when straight-lining was used than here. This is another argument in favor of using straight-lining. Here, the fit may not be good enough to really interpret the results. If we are willing to accept the fit and consider the estimates, then we see that the main results are not changing: there is a significant relationship between response time (RT) and bad, as well as AE and RT, but no relation between bad and AE. The main difference is for the effect of age on bad: a bigger effect is found when using the variances.
Appendix B
Authors’ Note
We are very grateful to Netquest for providing us with the necessary data for this article and to Willem Saris and two anonymous reviewers for their very useful comments on previous draft of this article.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
