Abstract
The EVS survey comprised in France two subsamples, one built randomly, the other according to the method of “enhanced quotas”. This article first presents the mechanism implemented by emphasizing the difficulties to obtain random interviews. Then it compares the two samples, first for variables entering the quota, and for other background variables, finally for dependent variables. The quota sample proves better than the random sample. Too educated, too rural, too old, the latter is also more politicized than the quota sample, which contradicts the usual critics of this methodology. The article finally shows that after weightings, the two parts of the sample have results sufficiently close to constitute an entity regarded as representative of the French population.
Proponents of statistical theory often categorically refuse samples by quotas because they do not allow good representativeness 1 . Indeed, that procedure does not designate the person to be questioned based on random choice; and surveyors keep a certain freedom of choice for respondents. In fact respondents ultimately have very different chances of being selected since, in the absence of an individual corresponding to a sought quota in the first household visited by the surveyor, he or she can try the household next door. Only random samples can truly know the entire population from a small portion of that population because they do not allow substitution of selected individuals by a random procedure. We can calculate non-contact and refusal rates, and a margin of error, all of which are reassuring in the eyes of a statistician.
The weight of statisticians in the organization of international social science surveys leads to forbid (in principle) the quota sample method (this is the case for the European Social Survey (ESS), the International Social Survey Programme (ISSP) and the European Values Survey (EVS). Yet, in many countries, including France, surveys with random samples are very difficult and expensive to carry out in exchange for benefits of data quality that are not always realized in practice. This is what we intend to show with an experiment carried out on the French part of the EVS 2008 survey 2 .
The Intended Survey Structure
The international coordination of the EVS required in 2008 that the sample for each country be at least 1,500 individuals selected randomly. Such a number is enough to make a comparison between European countries, but not enough to make detailed comparisons of national data and to draw conclusions on the evolution of different sub-groups of the population. Given the very significant costs of a random survey, and also given previous French EVS experience with quota samples, and finally given our desire to methodologically compare responses from random and quotas samples, it was decided to cut the sample in two, and carry out half of the interviews on the basis of random selection, while the others would be carried out with the use of quotas in a manner strengthened compared to the more commonly practiced pollster quota methods. Thus the entire survey had to have 3,000 respondents, with the two sets of 1,500 responses 3 .
Believing that a good survey sample is based largely on a broad geographic dispersion, it was decided to limit the number of interviews per area to six and therefore randomly select 250 municipalities (communes) 4 or districts (arrondissements) for Paris, Lyon and Marseille for each sub-sample, according to a matrix by region and town size, in proportion to its share in the metropolitan French population aged 18 or over; this is, 500 geographical collection zones in total.
The Random Sample
We must remember that we cannot, in the strict sense, carry out in France, as in many other countries, good random sample surveys for at least three reasons: – Exhaustive lists of individuals or households in France are very rare and are not available for public research. Only INSEE can use for its own surveys the lists of individual census forms
5
. – Finding the selected person on the list is not easy, despite a high number of visits at the home, and it is easier to reach people living in a household of several members than persons living alone and often absent from home. – In a democratic country, the answer to an opinion survey is not compulsory. Those selected can always refuse to answer, despite the efforts of the interviewer to encourage participation. And we know that the social situation of the interviewees and the theme of the survey can introduce differential interview acceptance
6
.
Due to the lack of an exhaustive population register in France, the method most often applied when you want to make a random survey (administered at home) is a kind of substitute for the random sample by using the technique of random route. Based on a random selection of geographical points of departure, the interviewer follows a path imposed to arrive at the front door of a household. He therefore does not choose who he wants within the limits of the quota, but must interview an individual “randomly” determined by the route taken. One of the problems of the method (used in particular for the Eurobarometers) is the difficulty to describe a route suitable for both rural and urban areas, which also makes it difficult to control the work of investigators: the indications for routes are not always easy to follow and by repeating the route, a controller does not necessarily end up at the same door. In addition, investigators may be tempted to use a neighbor as a substitute when a normally designated person refuses to answer. To overcome this difficulty at least partially, the method used in France for the European Social Survey (ESS) – a method developed by the Institute ISL – is to send a first investigator to the geographically selected spot with the responsibility to noting addresses on mailboxes, according to the “random route” principle. A few days or weeks later, another investigator will contact the households previously listed. It was this rather expensive but more rigorous methodology that was also used for the EVS survey.
Based on experience with other surveys (including the ESS), we planned to construct a list of 3,750 names 7 , 15 by selected area (commune), but in the first instance, two addresses were kept “in reserve”; that is, 500 emergency addresses. The ratio of available addresses (3,750) to expected interviews (1,500) was therefore 2.5, corresponding to a minimum success rate of 40 percent. But the goal was to do better (60 percent), without hoping to achieve the utopian goal of 70 percent required by the European coordination 8 . For this, five home visits, at different times of the day and week, were planed (at least one on a weekend and one in the evening).
The interviewers then sought in every household whose address was previously selected, the response of a person, randomly chosen (according to Kish method, after having made upon arrival a list of the household’s permanent inhabitants).
The Quota Sample
The second half of the interviews was obtained according to a so-called “enhanced quota” method with a quota crossed between sex and age (18-29 years, 30-44 years, 45-59 years, 60 years and over, a total of eight categories), occupation of head of household (current or last exercised) in six categories (farmers or artisans or merchants, professionals and senior executives, middle management, employees, workers, inactive residuals), level of education in five groups (no diploma, basic certificate, high school diploma (baccalauréat), technical and vocational diploma, higher education diploma). The introduction of this latter criterion to establish a reduced model of the population significantly increases the constraints of the surveyor but provides a much better representativeness of the samples, as we will be able to check it thereafter.
Carrying out the Survey in the Field
As the survey progressed in the field, it appeared that carrying out the random part was slow because of the difficulty in obtaining the required number of interviews and because of the number of visits required before giving up on an address. In fact, it was necessary to add the supplementary addresses around mid-July and sometimes to include an additional address to reach the anticipated 1,500 interviews. Table 1 shows the results: 3,993 addresses were used. A contact was only established with 74 percent of addresses, some units turned out to be vacant and others are only occupied very intermittently. Contact with the individual randomly selected was established for only 49 percent of the households. Because of the refusal of certain selected individuals, the success rate in relation to the list for the representative population was ultimately only 37.6 percent, which is low. For targeted individuals, the success rate is of course much better (76.8% accepted), and can even be considered good compared to what we know of the growing refusal attitudes toward surveys 9 . The crux of the problem actually lies in the difficulty of establishing contact with certain randomly selected households and individuals.
Success Rates for the Random Sample
Note also that the demographic monitoring of the sample (Table 2) showed the appearance in early July a significant lack of young people aged 18 to 29 years (due solely to the random sample). The shortage of young people was very problematic for us. We wanted in fact to have a sample of at least 600 people aged 18-29, due to commitments to a partner interested in the values of French youths. So we decided – after much hesitation – to include an additional sample of 70 young people aged 18-29 by quotas to correct the lack of young people in the random part of the sample.
Samples (Random and Quotas) by Sex and Age (Vertical %)
Comparison of the Two Samples - Infrastructural Aspects
On the Demographic Characteristics Constituent of the Quotas
The final sample by sex and age (Table 2 again) confirms what we saw coming up as things progressed: the random sample was too feminine and too old. The quota sample was instead very close to the desired percentages.
The lack of young people in the random part of the sample is primarily due to two structural factors: – Samples of households taken as a basis for representation of individuals over-represent individuals living in smaller households (one or two persons), whose average age is higher. – Samples consist only of ordinary households: all young people living in dormitories and hostels for young workers are largely missed by the survey because they are very seldom present in the parental home, even on weekends.
Moreover, youth mobility and frequency of their outings makes it more difficult to contact them (which is just as true in the random sample as by quotas).
The deficit in men is largely due to their higher rate of employment and their lower presence at home.
Consider now the sample obtained from the perspective of socio-professional groups (Table 3). For part by quotas, it refers to the professional group, past or present, of the head of household, who was declared by the respondent at the beginning of interview. In the random part, we consider the current or past socio-professional group of the individual, as measured in the demographic portion of the sample. Both sets present some differences with reality, but relatively limited: the most important difference concerns self-employed people, highly under-represented in the two sub-samples. These categories working a huge number of hours, don’t reply easily to surveys, and the great length of the questionnaire is likely to be very prohibitive for these population groups. There is also a significant excess of intermediate occupations in the random sample.
The Final Sample (by quotas and Random) by Socio-professional Group
* Current or previous occupation of head of household, source 1999 Census (RGP, 1999), updated by the 2005 employment survey (INSEE).
** Current or past occupation of the individual, according to the 2008 employment survey (INSEE).
In terms of educational level, the quota sample was based on the level of the highest degree obtained, to avoid a tendency to overstatement. Moreover, the demographic portion of the questionnaire measures (for the whole sample) the highest level of education attained (Table 4). For these two indicators, there is available INSEE data for judging the quality of the sample. The quota part shows a lack of lower degrees and a surplus of higher degrees. It is with the indicator for level of studies that the differences are strongest, especially for the random sample, with a surplus of 7.8 points for university studies compared to the desired sample structure. The differences are much smaller for the quota sample, although those at the primary level are also quite under-represented (in favor of the intermediate levels). We feel it is, regardless of the method used, difficult to obtain survey acceptance from people of low cultural and social levels, just as it is difficult to obtain acceptance from the professional groups most involved in their employment and perhaps from the less “relational” and sociable.
The Final Sample (by Quotas and Random) in Terms of Diplomas
* Statistics INSEE based on the 1999 census of the population, with adjustment in late 2007.
On Other Social-cultural Dimensions
If the quota sample appears better on the tables presented so far, it can be argued that this is normal since the quotas are precisely designed to provide a good representation for the controlled criteria. It is in fact on other variables that one needs to compare the two samples to assess their relative quality. This is what Table 5 does on a set of socio-cultural variables, retaining those for which the differences are significant.
The Most Important Socio-cultural differences for the Random and Quota Samples (Unweighted)
We first observe a surplus of rural inhabitants (8.2 points) and a deficit of people in large cities (6.2 points) in the random sample. This is confirmed by the region of residence (not shown in Table 5): the Paris region (as defined as ZEAT) represents only 13.3 percent of the random sample but 18.9 percent of the quota sample. If the regional distribution has been met in quota sample, it wasn’t for the random sample. Obtaining the desired number of interviews was only possible by offsetting the lack of the Paris and Mediterranean regions with surpluses in the South West, South East, the Paris Basin West and the West. It seems in fact very difficult to fulfill (within a reasonable time, in this case 4 months) the schedule of interviews in major cities where people are not often present at home. In rural or semi-urban zones, random sampling has a much better performance. However the random sample has a percentage of active full-time employees substantially greater than the quota sample (5.7 points). Active full-time employees are very difficult to reach in cities, which probably also affects the level of household income. This is the variable for which there is the most important difference: 10.1 points.
Just as it is more difficult to obtain respondents of lower educational level in the random sample, then it is also difficult to represent there the population of foreign nationals even if these people are a bit underestimated in the both samples (5.7% of the French population, according to INSEE 10 ).
Some differences appear regarding family situations, particularly for the never-married or civil partnership persons but who live as couples: they are more present in the quota sample, probably because of its more urban character (6.7 point difference).
A noticeable difference (7.8 points) exists in associative membership: according to the random sample, the French are noticeably more associative than according to the quota sample. However, the differences in religious affiliations are weak. It may be noted that the two samples represent the Muslim minority 11 , but minimizing it.
Comparison of the Two Samples - Superstructural Aspects
For All the Opinions and Values Measured in the Survey
If, as we have seen, some socio-cultural differences between the two samples are sensitive, they appear to have only limited effects on opinions and values. A systematic comparison of all survey variables, sorted according to the type of sample, shows that in the vast majority of cases the differences are minimal. Table 6 shows again the most significant differences (for which the degree of significance, based on Cramer’s V, is still very weak or non-existent). It presents the results both with and without weighting.
The Most Important Differences of Opinions and Values in the Random and Quota Samples
Individuals of the quota sample appear a bit more critical on the dimensions considered (more dissatisfied with their jobs, less confident in parliament and government, more critical of the democratic process…), they are also more materialistic, more attached to permissiveness and freedom of behavior and individual choice, more believers to hell and effectiveness of good-luck charms. These attitudes can probably be explained quite well by the younger and more urban character of the quota sample.
We also observe that the introduction of weights change only very little responses by quotas, but significantly more responses from the random sample due to their larger difference compared with the structure of the French population.
On the Degree of Social and Political Participation
If criticism of statistical theorists against quota samples is justified, especially if this methodology makes it easy to accept the refusal of answer by a potential respondent and replace him or her by the neighbor provided that the latter is included in the quotas (while in a random sample, the investigator would be forced to insist to obtain the response of the randomly selected individual), the respondents of quota sample should reveal more participation in social and political life, and be more eager to express themselves, more knowledgeable and sophisticated in social and political fields 12 .
A first result (already presented in Table 5) tends to discredit this supposition because people in the random sample are more often members of an association than in the quota sample (gap of 7.8 unweighted points), where individuals are expected to be more involved in social and political life. For all other tested variables of participation and competence, the differences are often small (Table 7), but they are also most often contrary to this supposition.
Social and Political Participation in the Random and Quota Samples
In the rough, individuals of the random sample are slightly more sociable than in the quota sample: they have a little more trust in others, they consider themselves a little more altruistic; they are less worried about European integration, and also less worried about their security. One the other hand, they are a bit more selective concerning the categories of undesirable neighbors (only variable that tends to go in the sense of the hypothesis). Concerning the relations with politics, still in the rough, the random sample is more politicized, less abstainer, more concerned with the development of democracy, all of which contradict the hypothesis of less politicization or participatory in the random sample as compared to the quota sample 13 .
Note here again that taking into account the weightings, the results for random sub-sample change pretty much with the differences often small but reversed with respect to the raw data.
Extreme-right voting intentions deserve special comment. We know that all surveys understate this electorate for two reasons: these individuals accept less than others to respond to surveys, and they sometimes hide their intention to vote for a party often considered harmful. The second explanation is valid for both the random and the quota method. Nonetheless, according to specialists critical of the quota methodology, the first reason should make itself felt more on quota samples where the insistence of the interviewer to obtain answers is considered lower. The quota sample should therefore have fewer supporters of the extreme right than a random sample. Again, this supposition is contradicted. The representation of the extreme right is more biased in the random sample than in the quota sample.
Comparison of the Two Samples - The Non-responses
According to the same theoretical hypothesis, a quota sample, since it is deemed to select more sophisticated and participatory individuals, should lead to non-response percentages lower than in a random sample. We have also tested that hypothesis. It must first be emphasized that non-response had begun to decline with the 1999 survey and it collapsed in 2008. This downward trend is also seen in other studies and may have several explanations: – One accepts less to reply to surveys, but as a result, those who agree are more competent individuals and therefore would be more able to reply to all questions; – The progress of educational level would lead to a stronger sense of competence for responding in matters of opinion; – Interviewers have become more professional and better for encouraging respondents to continue with the interview.
But due to the large decline in non-response in 2008 Values survey, we cannot exclude a priori a specific effect by this survey, which could be linked to the choice of a new survey institute to carry out the survey in the field in 2008 (one other than those that did the work in 1999 and 1990), reputed institute for its seriousness 14 . One can imagine that such an institute investigators know more leave some respondents the time to understand the issues (and to slowly repeat once if necessary) instead of delivering the questionnaire mechanically, in a standardized way and too fast. As a result, the respondents would have a better response rate. If this explanation is valid, which cannot be proven 15 , it applies to both the random and quotas sample.
Considering all non-response to the whole questionnaire shows that non-response to many questions have become negligible (less than 1 percent of respondents). When we consider the cases where non-response is at least this percentage, one can generally observed that the differences between the two samples are often very small and are not always in the same direction. We can consider that non-response for two-thirds of the variables are slightly higher in the random sample, whereas for one third of the variables, it is higher for the quota sample 16 . The result is therefore rather uncertain with respect to the hypothesis. It appears that the two samples are very close from this point of view and there is no systematic tendency, however slight. Table 8 shows the differences for the variables where non-response is the most frequent. The column for the differences clearly shows how small they are (only four cases with more than one percent).
Maximum Rate of Non-response, Comparing Parts of the Sample
Non-response may actually have several causes. On the left-right position scale, it was 21 percent in 1990 and 17 percent in 1999, which indicates both an inability to find a place on such a scale (ignorance of the political universe) and a rejection of this framework. This is what explains why non-response remains high. Non-response to the question about voting or abstention if a national election were held tomorrow also likely reflects both a challenge to publicly acknowledge one might not be a good citizen, a shortage of political competence and closeness with a party, and finally suspicion of the entire political class. The importance of non-response in trusting NATO and the UN is primarily due to insufficient knowledge of these organizations. Similarly, the inability to say if the earth may or may not be able to support an increase in population is due to the highly technical nature of the question: strictly speaking, only economists and demographers should be able to attempt to offer a cautious response. But obviously many agree to answer according to their ideological outline and it is obviously what legitimizes the question. In religious matters, the fairly high level of non-response for a significant number of questions is probably due to secularization. A number of people believe that the religious world does not concern them. But non-response to dichotomous questions about beliefs can also have to do with the clear-cut nature of answers (yes or no), at a time when many have uncertain beliefs.
Conclusions
The structural comparison of the samples showed fairly significant differences in selected sub-populations, at least for some dimensions. The random sample is a bit too old, lacks some self-employed, is much too rural, not well distributed in the major regions, has too many persons with jobs, too many diplomas, and too many people with high incomes. Of course, the quota sample is not structurally perfect, but it is less offset from the structure of the population and not solely for the variables controlled for in the sampling process.
The formidable problem encountered during the implementation of the questionnaire was to succeed in limiting the problems of random sampling, while the realization of the quota-sampling plan was without surprises, and gave a rather better sample cheaply. The assumption that the quota sample would not sufficiently select categories of difficult-to-reach people proved to be false, at least in this survey, where important means were available to practice “enhanced quota” sampling 17 . It is rather the opposite that has taken place in our case: with more diplomas, the random sample is more sophisticated and participatory, missing – even more than the quota sample – members of disadvantaged minorities.
Fortunately, these weaknesses do not seem to have significant effects on the responses in the different areas of values, which are at the center of this study. This is basically not very surprising since the relationship between values and infrastructure variables are often only of low or medium intensity. Overall proximity of responses in both samples permits us to conclude that it is perfectly legitimate – especially using weighting which often makes some raw differences disappear – to consider the sample as a single unit, for which the quality of data is equivalent.
The introduction of the method of random sampling, for part of the survey, could not explain the changes of 2008 data compared to previous waves. These changes are linked to social change at work in France as in other European countries. In contrast, maintaining a partial quota sample introduces no particular problem with the comparison with the data of countries using the random method.
The quality of a sample is less dependent on the choice of the random method (which would necessarily be, in principle, always the best) than on the strength of the planned survey, whether it be random or quota samples, the number and the choice of criteria used in a quota procedure, the quality of work of the network of interviewers, the establishment of strict rules of interviewer practice, and the verification of these practices, not to mention the importance of actual content of the questionnaire 18 .
In the future, the coordination of international surveys would benefit from more attention to national habits in sample survey research. The random sample is certainly a good practice when you can have a very good population list. The stopgap “random route” - in lieu of the real random method - is largely an illusion. Quota sampling, when performed on a sufficient number of criteria and with professionals who master the possible flaws at the level of the interviewers, proves to be a very good method of administering questionnaires, and at a cost that is reasonable.
Footnotes
Acknowledges
I thank all those who were willing to share with me their comments on an earlier version of this text, especially the heads of GFK-ISL (Bernard Mandin, Hervé Bastide, Claire Blanchard), Dominique Joye (professor in Lausanne and international survey specialist), Nicolas Sauger (FNSP Researcher, then responsible for the ESS survey in France), Yannick Lemel (INSEE), Michel Lejeune (emeritus professor of statistics at the UPMF in Grenoble), Frédéric Gonthier (lecturer in political science at IEP in Grenoble), and Benoit Riandey (INED). I also would like thank a lot Karl Van Meter who realized the first translation of my text.
