Abstract
Self-reports in linguistic study, which were central to the dialect surveys of the twentieth century, have, by and large, been relegated to the sidelines by more advanced sociolinguistic techniques in recent years. This article probes into the validity of written self-report surveys in relation to the fieldwork method for Vancouver, British Columbia. Confirming Chambers’s general findings of equivalence, it produces insights into the preferred length of written questionnaires and offers recommendations as to question type. The present article also compares the written questionnaire results to acoustically analyzed recorded data for yod-dropping and the low-back vowels before /r/, identifying linguistic items that correlate well with results from self-reports and those that fail to produce reliable results because of ongoing linguistic change or reindexicalization in the case of yod-dropping. Overall, written self-report surveys are found to be highly reliable data gathering tools if certain factors are kept in mind.
Keywords
Introduction: The Written Questionnaire and the Vernacular
Not long after the completion of Georg Wenker’s (1881) pioneering postal questionnaire in late-nineteenth-century Germany, dialectologists in general and later sociolinguists in particular became skeptical toward the written questionnaire. Because of Wenker’s methodological oddities and the success of Jules Gilliéron’s (1902–10) field-worker method with the Atlas Linguistique de France, the written questionnaire has never again reclaimed the status of a fully legitimate data gathering tool. Historically, if dialectologists used postal questionnaires at all, they have been hesitantly doing so. Chambers (1998a:222) summarizes the situation aptly:
In nearly all cases, its inclusion [the written questionnaire’s] has been announced almost apologetically. There is often a sense of opportunism associated with its use, as if it might have been preferable to avoid it if it were possible.
As their major characteristic, written linguistic questionnaires are filled out directly by the informant, without an intermediary, and are therefore self-reporting in nature. They can take the form of postal questionnaires (sent out by mail), email questionnaires, web-based surveys, or distributed paper questionnaires, among others. Twenty years since its rediscovery in the context of Dialect Topography (Chambers 1994), written questionnaires remain relatively underused in the bigger picture. This is surprising, as the advent of Internet-based surveys and inexpensive tools, such as surveymoney.com, have further increased the written questionnaire’s unrivaled time-effectiveness when compared to any other data gathering methods. A critical look at the written questionnaire’s advantages and drawbacks for linguistic purposes therefore seems to be timely and beneficial.
The primary goal of the present article is to test the written questionnaire’s validity by comparing the data it produced with data gathered with the field-worker method on one hand and with acoustic analyses from sociolinguistic interviews on the other. Chambers (1998a) compared postal data (written questionnaire) and fieldwork-based data for lexical variables from the Linguistic Atlas of the Upper Midwest and found solid evidence for the robustness of the written questionnaire. The present study aims to test the reliability of the written questionnaire beyond the lexical level. All three data sets were collected from English speakers in Vancouver, British Columbia, Canada. The written and fieldwork surveys were carried out in the spring of 2008 and consist of 423 postal-type and 187 fieldwork questionnaires for thirty-two linguistic variables. They include phonemic, morphological, syntactic, lexical, and usage variables. For the acoustic analysis, ten sociolinguistic interviews were recorded in the spring of 2009 and analyzed for two variables, the low-back vowel merger and yod-dropping, in a number of contexts.
The present article seeks to address data-related and methodological questions. Of the thirty-two variables, twenty-one are a resurvey of Chambers’s (2004) Dialect Topography of Canada survey, while eleven questions are new in the Dialect Topography framework. Going methodologically beyond the original test in Chambers (1998a), I compare results for yod-dropping and the low-back vowels from an acoustic analysis with the self-reported data. The results are further used to provide an assessment of the chances and challenges of written questionnaires and their role in sociolinguistic and dialectological studies.
The comparisons suggest further support for the written questionnaire as a data gathering tool and show, more clearly than before, problem areas for written questionnaires and possible workaround strategies. Some researchers have shown that self-reporting might be the best option for low-frequency items (Pratt 1983), and Bailey, Wikle, and Tillery (1997:57) are very clear in their assessment of low-frequency variables that “self-reports might be more valid and reliable measures of linguistic behavior than linguists have supposed.” Chambers (1998a:234) reports that, when compared with fieldwork data, “the postal questionnaire cannot be shown to be different.”
In what follows, I first briefly place the current surveys in the context of studies on Vancouver English, before discussing the survey designs, data collection methods, and sample sizes. Second, I compare the Vancouver data from the written questionnaire (WQ) and the fieldwork questionnaire (FQ) and classify deviations qualitatively, before introducing a quantitative analysis. Third, I put the results of the sociolinguistic interviews in relation to the participants’ self-assessments. Finally, the article presents recommendations for self-reporting in general and WQs in particular.
Data and Data Collection Methods
The city of Vancouver is a fairly young city. It was incorporated in 1886 and has been settled by English speakers for little more than 130 years, which allows present-day studies to reach back, via apparent time, to the second and third generations of English-speaking settlers. First linguistic studies in the area were carried out in the 1950s and provide a useful real-time correlate. The basic formation of a new dialect in a newly settled area, as laid out in new-dialect formation theory (Trudgill 2004; Dollinger 2008), happens in the first three generations since the initial settlement and makes present-day Vancouver data relevant. These reasons, paired with the city’s relative understudied status, make Vancouver a prime location for this methodological comparison.
Since we are dealing with real-time and apparent-time components, a word on the methodology of merging these approaches seems in order. In this study, the time depth is measured by the difference between the oldest speakers’ ages at the time of data collection, minus twenty years (an age of twenty is taken to represent the end of a speaker’s formative linguistic years). The earliest linguistic studies on Vancouver, starting in the 1950s, can today be used as historical signposts to reconstruct long-term linguistic change. Gregg (1957), the earliest study, is a description of the phonemic inventory of his Vancouver university students from the early 1950s and reports important impressions, as we will see below, from a real-time perspective. Scargill and Warkentyne (1972) and Scargill (1974) use the results of a national postal survey of school children and their parents, which provide an apparent-time window to just about the time when Gregg interviewed his Vancouver students. Chambers and Hardwick’s (1986) apparent-time study reaches back to the 1960s, while Gregg’s (2004) study on Vancouver English, for which data were gathered in the late 1970s, pushes the diachronic window back to the 1940s. This temporal benchmark is closely reached by Chambers’s 2004 survey and the 2008 WQ and FQ from Vancouver. Beyond the 1940s other means, such as written evidence, and, where available, early audio recordings, would need to be used (Dollinger 2010b).
Data collection for the WQ was carried out following the principles of the Dialect Topography (DT) of Canada framework (Chambers 1994). At present, DT of Canada covers seven Canadian regions and offers comparative baseline data from four adjacent U.S. border regions. 1 Vancouver, British Columbia, and its urban environs (metro Vancouver) were surveyed in 2004, under the direction of Tony Pi. The DT questionnaire is intended to produce instantiations of the vernacular as much as possible. Following Chambers’s approach, the linguistic part of the WQ was opened with the following instructions:
Please read through these questions one by one and mark the option that is most natural to you. We are interested in the language you use when you are among friends, and not in what somebody else might consider “better.” Usually, gut reactions work best here. If you use a word that we do not list, please make sure to write it into the margins next to the particular question. Please Do Not Go Back And Change Any Answers (we are interested in your gut reactions)!
The WQ and the FQ
The written survey uses a self-report questionnaire that is equivalent to the DT of Canada questionnaire in all but two respects: first, it is significantly shorter, and, second, it was not sent by mail. Rather, students in my English dialectology course (ENGL 323A, fall 2008) approached potential informants in public places and asked friends and relatives whether they would be willing to spend ten to fifteen minutes filling out a questionnaire on Canadian English. Once the participants agreed in principle, they were informed about the survey procedures (informed consent), then handed the questionnaire and left alone. In case participants had any questions, they were told to fill out the form “as best as they could on their own” and that interpreting the instructions was part of the survey.
Fieldwork followed a similar approach. An adapted FQ was used by the students as a detailed guideline to interact with the interviewees. The questions from the WQ were transposed and the instructions were paraphrased to be read out to the interviewee. The FQ was used as a guide to ask questions but was much more specific than the traditional linguistic atlas field-workers’ “work sheets” (Kurath et al. 1939:147-158). When compared to traditional fieldwork, the results of the FQ are biased toward more homogeneity by a highly structured interview routine.
Table 1 illustrates the basic difference between the WQ and the FQ. The WQ questions are on the left, the FQ versions on the right. As can be seen in question 18, in some cases the written instructions were taken over unchanged into the FQ, with the difference that the informant had to choose from options that were offered verbally. Field-workers were instructed to show the word in writing if required, but they did not offer the written form unasked. For phonemic questions, as shown in question 10, the informants were shown the word on paper and asked how they pronounce it, which the field-worker would then note down. The participants did not see the respelling (ad-ver-TAHYZ-ment or ad-VUR-tis-ment), which was offered as guidance to the field-worker. If an informant produced an option not shown, it was noted as well. While in the FQ the likelihood for other, unlisted answers is smaller than in the WQ (where informants were left alone), the documentation method inevitably steered the field-worker toward a binary documentation of phonemic incidence.
Comparison of Written Questionnaire and Fieldwork Questionnaire
The generational spread of the data is quite satisfying: the youngest informant in WQ and FQ is 14 years old, the oldest informant is 101 years old. Many long-term residents are included, such as one informant who spent her entire formative years (ages 8–18) in the metro Vancouver area, where she has lived for 84 years. The survey data offer an apparent-time depth of more than 80 years, which accounts for about two-thirds of the city’s settlement period. With these data, we can tap into the third generation of settlers in Vancouver, some sixty years after the original settlement, which invites comparisons with Trudgill’s findings from New Zealand. 2
The first two lines in Table 2 show the absolute numbers of WQs completed by participants in each age cohort for the original Vancouver DT survey (DT 2004) and the 2008 written survey (WQ 2008).
Sample Sizes for DT in 2004, WQ 2008 Surveys, and FQ 2008: N per Age Cohort, Based on WQ, q1 – Different
DT = Dialect Topography; WQ = written questionnaire; FQ = fieldwork questionnaire.
The absolute figures are solid in both DT and WQ, especially for the two youngest age cohorts. Only the octogenerians and older participants appear to be less optimally represented. Although DT 2004 has slightly better coverage across all age cohorts, both DT and WQ represent large databases that are hitherto unparalleled in the area. 3
The last line in Table 2 (FQ 2008) shows the number of informants for the FQ. The figures are generally lower than the WQ, while the twenty-year-olds stand out with three times as many interviews as the second-largest age group. The age cohorts of the septuagenarians and up have, again, markedly fewer participants.
A juxtaposition of WQ and FQ questions is provided in Appendix A, which also introduces the variables. As is shown there, the questions are virtually identical with the ones asked in DT 2004 but have been occasionally adapted to the local situation. The “new” variables complete the set (Appendix A, second table). Six of them, to the best of my knowledge, have not been used in Vancouver, while five are new in the DT context. The new variables include pronunciation variables (advertisement, with the principle variants /ædvзtəzmənt/ or /ædvзtaizmənt/, presentation /prϵzϵnteiʃn/ or / prizϵnteiʃn/), and lexical variables: the word for coffee cream—creamo, a local variant, or cream; half-and-half, the word for whole milk—homo milk or homo as Canadian variants; and the word, hydro, meaning a marijuana joint from hydroponically grown plants. Hydro is first documented in the Bank of Canadian English (id# 2208) in 1998 (Dollinger, Brinton, & Fee 2006). In the area of word formation one variable, the incoming form wait time, as opposed to traditional waiting time, was polled. Previous variables are the pronunciation of groceries (Gregg 2004), as /groυʃəri:s/ or /groυsəri:s/, the pronunciation of sorry as /sɔri/ or /sɒri/, with its midclosed vowel before /r/ in Canada, and two local lexical items that are on the cusp of extinction, slough “ditch” and skookum, from Chinook Jargon, meaning “great, big.”
Comparison I: WQ and FQ
The written and fieldwork data sets are compared to each other in two ways: First, an overall assessment is given by visually comparing the apparent-time charts, which suggests a rough classification into five types of fit between the surveys. Second, statistical tests will be run on the variables and the entire data set. To facilitate the first type of assessment, a numeric threshold will be introduced for deviations between the surveys, which will then be applied to assessing effects of the length of the questionnaire, the age cohorts, and different sample sizes.
Classifying the Match Qualitatively
A straightforward yet not entirely precise way of assessing the WQ and FQ is to match the patterns visually when graphed. Five general patterns evolve, with are shown in Table 3, with the number of variables for each category and the order of the question in the questionnaire (e.g., q22 for the twenty-second question of thirty-two).
Five Types of Matches between the WQ and FQ
WQ = written questionnaire; FQ = fieldwork questionnaire.
For reasons of space, I illustrate the method with three cases. Figures 1 through 3 illustrate categories E (typical, pretty good fit), D (good fit, with outliers in the older age cohorts), and A (extremely different fit). Here, as elsewhere, the results from the WQ are shown on the left, the FQ data on the right.

Slough: Percentage of knowledge

The written questionnaire (left) and the fieldwork questionnaire (right) show a good fit, with two outliers in the sixty or older age cohorts (type D)

The written questionnaire (left) and fieldwork questionnaire (right) results are extremely different (type A)
A typical fit can be seen in Figure 1 (type E), which shows the familiarity of speakers with the relic lexical item slough /slu:/ (q20). Slough can either signify a slow-flowing and stagnant body of water, such as the side arm of a river, or a man-made water hole. The FQ tends to underreport the knowledge compared to the WQ, where a more gradual loss of the term, as one would expect, is seen. This I call a “typical, pretty good fit,” because the overall progression is maintained in both data sets, while one can see deviations of up to 20 percent between the WQ and FQ. It is typical, because it accounts for twelve of the thirty-two variables.
The second example, Figure 2, is a “good fit,” with outliers in the older age cohorts. The variable concerns the palatalization of the first-syllable fricative in asphalt (q13), which is either a palato-alevolar fricative /ʃ/ or an alveolar fricative /s/. As the WQ data suggest, there is a slight trend toward /s/, which is not borne out by the second-oldest age group in the FQ or the oldest cohort, eighty and older, in the WQ. As will be shown later, the older age cohorts are an important subgroup for the discussion of potential problems with WQs.
Finally, Figure 3 is a grammatical variable, which summarizes the data for the three most frequent prepositions following the comparative adjective different (q1).
Here, the percentages for different than in the FQ diverge drastically from the WQ data, especially in the youngest and oldest age cohorts; different from is the absolute majority form in FQ, while WQ shows fierce competition with different than. A peak for different than can be seen in FQ in the forty-year-olds, which appears in WQ in the thirty-year-olds and thus one decade earlier. With such differences, it can be classified only as “extremely different,” type A. Types B and C are classified in a similar fashion.
Taking all types together it can be said that, generally and impressionistically, the FQs tend to jump or fluctuate more abruptly between the age cohorts (see, e.g., the categorical scores at the 0 and 100 percent levels in all three figures, which seem to occur much more frequently in FQs). Also, the FQ data seem to produce more troughs and peaks that fall outside a trend, such as with outliers in type C and type B. It seems that, overall, the WQs provide a smoother pattern of changes in progress than the FQs, which cautiously indicates sounder and more reliable data in the WQs.
Classifying Statistically: Paired t-Test and Wilcoxon Test
One way to assess the match between the FQ and WQ statistically is by comparing the means individually for each variable and age cohort on one hand and for all variants combined on the other hand. This means that we compare the average scores in the WQ and FQ of, for example, all ten-year-olds that say st/ju/dent, then all twenty-year-olds, and so forth. We repeat this for each age category and variable score. Then, as an overall test, we can compare all scores in the WQ for all variables with all scores for the FQ.
Statistically, the choice of the instrument depends on a number of factors. In this survey design, each value in WQ has a matched partner in FQ. A paired t-test would be the test of choice for the paired samples of WQ and FQ. This test, however, can be used only if the data are normally distributed. This means that each set of answers for the WQ and FQ first has to be tested for the characteristics of the normal distribution, by way of a Shapiro–Wilk test. 4 If the criterion of normality is met, one can apply the paired t-test. If not, another test, the Wilcoxon signed rank test, can be used. This procedure was carried out for each variable and for the overall sample.
Comparison of the overall match of all variables between the WQ and FQ must rely on the Wilcoxon test because the overall sample is not normally distributed (as shown in Appendix B). Taking all variables together between the WQ and FQ, the p value is .163 and thus not significant (i.e., it is greater than the threshold of .05). 5 This means that, overall, the WQ and FQ results cannot be shown to be significantly different from one another. This supports the reliability of the WQ, as it cannot be proven that the WQ answers are different, that is, “worse” or less reliable, than the FQ answers.
Figure 4, which shows graphically the correspondence between the WQ and FQ answers, serves as a useful guide to compare the variables individually. In the scatterplot, each mean value for each age group and variable is plotted as a dot for FQ and WQ. Most points cluster around the smoother line, which is close to a forty-five-degree line. If there was an ideal match, all points would be exactly on a forty-five-degree diagonal. The figure also reveals outliers, the extreme cases along the x-axis and y-axis. In these cases, a particular age group answered for one variable in one survey categorically with only one variant, where the other survey group did not. If their answers were more similar, they would be found more in the middle along the diagonal line. It must be in the interest of any survey to eliminate these outliers, and this especially concerns the ones farthest away from the diagonal. The individual results help explain this patterning. As Appendix B shows, four variables show significant differences between the WQ and FQ: avenue, lever, presentation, and zed. All four have in common that the three oldest age groups show big differences between the surveys As we will see later, the Length of the Questionnaire section isolates one contributing factor that affects the misfit.

Scatterplot of means per age group and variable for the fieldwork questionnaire and written questionnaire
Deviations and the Less Than 10 Percent Threshold
The basic typology from Table 3 offers a first assessment of the WQ and FQ data, while the Wilcoxon test has shown the overall statistical acceptability and equality of the samples. In Table 3, typical, “pretty good” fits show deviations up to 20 percent between data points between WQ and FQ. This entails that the choice of a threshold needs to be set at less than 10 percent of difference in either direction between the WQ and FQ scores. If WQ is, for instance, at 15 percent, then FQ could result in 24.9 percent or 5.1 percent, still producing the 20 percent range around the WQ value. Deviations that are equal or greater than 10 percent in either direction (≥ 10 percent), count as “big” deviations, smaller ones as “small.” Figure 5 illustrates the rationale, concerning the pronunciation of presentation. For each age group and variable, the deviations between the WQ and FQ were calculated. The percentages of the initial syllable /prϵ/ (as opposed to /pri/) are depicted in the graph. In the teenagers to the fifty-year-olds, the deviation is less than 10 percent per cohort; however, for the sixty- to eighty-plus-year olds, the deviations are equal or greater than 10 percent.

Measuring deviations for presentation (percentage of /prϵ/ shown)
Less than 10 percent deviation between the two types of survey data is a fairly good match; 10 percent or more is where deviations become progressively problematic to the point where little to no correlation between the data sets can be seen. In figure 4 the older age cohorts produce three deviations of ≥ 10 percent, which would be counted as three deviations.
Table 4 shows the deviations that are equal or greater than 10 percent per age cohort and variable, starting with six deviations, and going to down to one. There are additional fields on the left in Table 4. The linguistic level of description is shown next to the question number. Any clustering of lower or higher question numbers would reveal a bias of errors depending on sequence. Next, the field “answer” lists the answer choices of each question: whether they are binary (as in both examples shown in Table 1) or multiple answer/open questions (marked as “m/open”, such as in question 1, different than, where there are three choices offered or question 3, which elicits the target word tap or any of its variant forms, with an open answer line).
Deviations Greater Than or Equal to 10 Percent between the WQ and FQ Surveys
WQ = written questionnaire; FQ = fieldwork questionnaire.
= includes the forms homo milk, homogenized mil and homo
As shown in Figure 5, the results produce big deviations in the sixty, seventy, and eighty-plus age cohorts. All big deviations are entered in Table 4. For instance, question 19 (q19), presentation, is found under category “3 deviations”, which puts the variable presentation, therefore, in the middle range of reliability. Different than is the most erratic variable, as in every category, except the forty- and sixty-year-olds, WQ and FQ differ by 10 percent or more; several variables have only one big deviation.
The juxtaposition of the WQ and FQ is interesting for a number of reasons. It becomes clear that there is no single linguistic level that shows more or fewer deviations; the trend goes across all levels: six big deviations are shown in one syntactic variable, and five deviations are shown in three phonetic, two morphological, and three lexical variables. This means that all four linguistic levels are found in the category of five or six big deviations. Four deviations are found in one syntactic, one lexical, and two morphological variables; three deviations in four phonetic, one syntactic, one morphological, and one lexical; two deviations in three phonetic and one lexical; and one deviation is found in five phonetic and two lexical variables. Every variable has at least one deviation of greater than or equal to 10 percent between the WQ and FQ. The four significantly different variables are all pronunciation variables (Appendix B). However, pronunciation variables are also among the most reliable variables, featuring five times in the category of one deviation above.
Linguistic levels apparently do not explain these deviations. However, the way the answer choices are offered in the questionnaire provides interesting insights. We distinguish between multiple-choice/open-answer and binary-answer options. The syntactic variables all score three or more big deviations: different than yields six deviations, Tom and I yields four, and sick to yields three big deviations. What they have in common is that syntactic variables work with multiple answer choices: three options for different than: than/from/to, three for Tom and I: Tom and I, I and Tom, Me and Tom, and four for sick to (the stomach), with the options, sick to/at/in and I do not use this expression.
On the other end, those variables that have three or fewer deviations are almost all in a binary answer format: asphalt or ashphalt (q13), zee or zed (q17), or avenyoo or anvenoo (q8). There are four vocabulary items in this category, all of which have been polled with multiple-choice/open-answer questions: slough, tap, napkin, and hydro. These four share that they represent extreme cases: the first three are changes that have run their course to (near) completion, and the fourth, hydro, is just taking off at the other end of a lexical s-curve (hydro is so new that it is known by only six people in WQ and one person in FQ).
There are other factors that influence the match between WQ and FQ data. For instance, wait time (as opposed to waiting time) is a binary variable and would thus be expected to produce a fairly good match. However, wait time has had more time to run the cycle of a lexical change, as it is used by about 50 percent of the population and is very much on the steep slope of the change, which spoils the match. As a vigorous change, it produces five deviations. Similarly, different than, also hovering around 50 percent, produces also very deviant results with six big deviations, which suggest the recommendation that vigorous changes are not well suited for self-reports.
These observations suggest that binary answer options tend to produce more similar results between the WQ and FQ, which is intuitively appealing: the fewer choices there are, the more closely the data will match. The DT tradition has long worked with binary answer options, and in this tradition WQs are often conceived as being limited to variables that have two discrete, or largely discrete, variants. Multiple answer choices and open-answer questions are also used, and the results here suggest that with these instruments the WQ produces different results from the FQ. In any case, researchers will have to know the variables well to design an effective and reliable questionnaire. For variables with binary variants the results between the survey formats are most compatible. Needless to say, binary choices can be used only where there is great certainty that no intermediate variants occur. If, for instance, one found comments in the margins about another answer choice, one would need to rethink this particular question.
It is crucial in self-reporting that speakers have conscious access to their linguistic behavior. If variables are known to be on the steep slope of a change in progress, the self-reporting survey method, whether WQ or FQ, will influence the results. What possible safeguards can be taken against this effect of the method? In the following section, we aim to isolate an independent variable that might explain the deviations. Table 5 shows all big deviations ordered by age cohort: seven in the teenagers, six in the twenty-year-olds, and so on to the older age cohorts, the sixty-year-olds, seventy-year-olds, and eighty-plus-year-olds, who produce more big deviations than the younger ones (seventeen, twenty-one, and twenty-four, respectively).
Deviations Greater Than or Equal to 10 Percent per Age Group Juxtaposed with Sample Size
Boldface indicates WQ sample size, normal font FQ sample size per age cohort
What appears to be an age effect from this perspective—perhaps the participants sixty and older tend to be more easily distracted and are thus more prone to offer different answers in the WQ and FQ questions—turns out to be intertwined with another factor: sample size, the sizes of which are shown in the bottom line in Table 5. The bottom cell in the twenties column, for instance, reads: in the twenties age cohort, 153 WQ answers were obtained, 66 FQ answers, and six deviations ≥ 10 percent were counted. If we reorder the columns by sample size (for WQ), an implicational scale unfolds: the larger the sample, the fewer the deviations between WQ and FQ, and this holds true not just for the WQ sample sizes, but, mostly, also for the FQ sample sizes: six deviations in the twenty-year-olds are based on 153 WQs and 66 FQs; seven in the ten-year-olds on 82 WQ and 21 FQ; and twenty-four big deviations in the eighty-year-olds and older on the smallest sample sizes of nine and six, respectively. Age has an effect, but sample size appears to be another crucial factor that influences the closeness of the match between the WQ and FQ.
Because of the involvement of two independent variables, age and sample size, it is not possible at this stage to say which of the two has the greater effect. By looking at the big deviations for the first and second halves of the questionnaire separately, however, some recommendations can be made, which are explored in the next section in the context of questionnaire length. Questionnaire length provides a clue toward untangling the intervening factors age and sample size and thus proves to be an important factor for the design principles of self-report questionnaires.
Length of the Questionnaire
Every survey design needs to work with a questionnaire that is free from bias, and questionnaire length is one key factor to consider. Besides the fifteen personal background questions, the DT of Canada framework (Chambers 1994), polls seventy-six variables that prompt for eighty-one answers and lasts about twenty minutes. In the 2008 Vancouver WQ, only thirty-two questions were asked (and eighteen background questions), making the questionnaire about half the length of the original DT survey. Chambers’s (1994:39) return rate in DT of the Golden Horseshoe was 53 percent, which is excellent for a postal questionnaire. In the WQ, with the method to approach potential informants directly, it was close to a 100 percent. Only very occasionally did an informant not finish the questionnaire.
The completion times for the WQ, however, were somewhat surprising: while the average completion time was around fifteen minutes—the questionnaires could often be finished in less than ten minutes—many older informants, ages sixty and up, took considerably longer; in the case of one eighty-year-old, a break had to be taken, and in other cases up to an hour was spent on the questions.
These different experiences between the “long” DT survey and the “short” WQ survey invite a closer look at the questionnaire. The results for WQ and FQ can now be used to gauge possible length effects. Table 5 shows a correlation with sample size. Can a difference also be seen between the WQ and FQ when the first half of the questionnaire is compared to the second half? Figure 6 offers this information for the linguistic questions (the respondents reach the first linguistic question after eighteen social background questions):

Deviations greater than or equal to 10 percent in the first and second half of questionnaire, by age group
An interesting pattern can be seen. The younger age cohorts, teens to thirties, show more deviations in the first half of the questionnaire with the thirty-year-olds in the lead with two deviations in the second half but six in the first. The forty-year-olds show five big deviations in each half. Then, the fifty- and sixty-year-olds show an increase in deviations in the second half: the longer they work on the questionnaire, the more the WQ diverges from the FQ scores. The seventy- and eighty-plus-year-olds are in a category of their own: they show, regardless of the section, a steady deviation with ten or eleven and thirteen or eleven errors in the first and second halves, respectively. These figures indicate that the two older age cohorts may indeed be more unreliable respondents.
Figure 6 captures two levels of deviations: one level is around five, the other around twelve, as represented by the dotted lines. The sixty-year-olds provide a clue to what is happening: their deviations in the second half skyrocket, 6 from five to twelve—that is, from one level to another. It seems as if the sixty-year-olds are the missing link between the more reliable younger-than-fifties and the least reliable seventies and older; whereas sampling size affects the younger age cohorts, it is the sixty-year-olds, whose sample size is in between the younger and the older ones, who clearly show some fatigue factor in the second half. In terms of sample sizes, the seventy and up cohorts have only ten (and fewer) informants, the fifty-year-olds and younger, around fifty informants and up. The sixty-year-olds are in the middle, with twenty-four informants. Their increasing deviations, when compared to the FQ, points toward fatigue that seems to affect the sixty-year-olds. Since their deviation score is low in the first half, their deviations are not an effect of sample size.
The paired t-tests have shown that the variables avenue, lever, presentation, and zed are significantly different in the WQ and FQ. Inspection of these four variables reveals big deviations in all of them in the age cohorts sixty and up. For three, marked gray above in Table 6, the three oldest age cohorts produce bigger deviations than the five younger age cohorts combined.
Deviations in Percentage between the WQ and FQ for the Four Significantly Different Variables
WQ = written questionnaire; FQ = fieldwork questionnaire.
It is clear that the three oldest age cohorts cause—and in the case of avenue, considerably contribute to—the significant differences in their paired t-tests (the t-test is computed using the standard deviation, which depends on the points’ differences from the mean).
In sum, fatigue may plausibly explain deviation in age groups sixty and up: older people tend to tire more quickly than younger ones. Other factors may contribute to fatigue, such as the difficulty of writing because of pains (e.g., arthritis), which would need to be considered in any WQ. As long as these factors are derived from the physical difference between younger and older speakers, we may summarize them all under “fatigue.” A linguistic reason for the deviations, such as that older speakers tend to avoid stigmatized forms, would present a serious problem of data collection. However, for the variables studied here, especially the four significantly deviating ones, such an interpretation can hardly be brought forward. Zed, certainly, is a prestige variant in Canada, as is yod-ful avenue. The pronunciations of presentation and lever have, apparently, not been commented on. In those cases, there certainly is no discourse of the type one can find for surface-level features such as Canadian spellings with -our, such as colour or favour (Dollinger 2010a).
At this point, the following conclusions can be drawn from the comparison. Most importantly, the WQ and FQ are not significantly different. These study’s data indicate no apparent effect of linguistic levels on the deviations, as they occur across the board. There is an effect of sample size on deviations between the methods. As a recommendation, if possible, WQs should target at twenty-five informants per cell (which corresponds to the sixty-year-olds), better yet fifty (forties and younger). In the seventy- and eighty-plus cohorts only ten and nine informants, respectively, in the WQ were the most problematic cases. The sixty-year-olds show an effect by the length of the survey. One way to alleviate the situation is to produce shorter questionnaires so that older informants (sixty and older) would receive questionnaires with fewer linguistic questions, perhaps twenty-five or fewer. It also might be useful to ask the background questions at the end of the questionnaire to ensure maximum attention and alleviate factors of fatigue in the linguistics section. Conversely, however, fatigue might then interfere with the reliability of the background data. While a reversal of the linguistic and background section is easy for the WQs, in the fieldwork method it would do away with a smoother, more natural transition into the interview process. Naturally, there is a limit to how short the linguistic section could be: with fifteen to eighteen background questions, it seems that the linguistic section should not be just a few questions either, so they should at least match the number of background questions. Given the informants are probably used to filling in background questions, this part will still be shorter than the linguistic questions, to which they are less likely accustomed.
Comparison II: Sociolinguistic Interviews
While the results from the previous section point out some critical issues, they furnish support for the validity of WQs when compared with the field-worker method. There is, however, the chance that both surveys, which are both essentially self-reporting, could err in the same direction. In questionnaires some variables are more prone to show a bias than others. Pronunciation variables are often considered the most problematic when polled with self-reports. To gauge the possible effect of the WQ, two phonemic variables were chosen for acoustic analysis. The variables are yod-dropping, also known as glide deletion and the low-back vowel merger. Yod-dropping was chosen for its status as a well-established sociolinguistic variable in Canada (e.g., Clarke 1993, 2006; Chambers 1998b), which also figures prominently in DT self-report surveys (e.g., Chambers 1994). The low-back vowel merger was chosen since it is a “pan-Canadian development” (Boberg 2008b:136) but has, since Scargill and Warkentyne (1972), not been polled with self-reports, apparently because of validity issues.
By choosing phonemic variables, and therefore variables that are conventionally treated as less conducive toward written surveys, this comparison seems to be biased against the WQ on one hand. On the other hand, Table 4 shows that eight phonemic variables (pho) have two or fewer big deviations, which points toward the equivalence of these two data sets. All contexts for yod—lexical items avenue, coupon, and student—show one big deviation, as does the variable relating to the low-back merger (sorry). Since these variables behave so uniformly, this comparison cuts to the core of the validity of both the WQ and FQ. We have seen from above that WQ are more reliable than FQs. If it can be shown that their results are reliably reflected in the acoustic data, WQs must be considered legitimate research methods.
Five questions from the WQ and FQ surveys are relevant to the interviews, as shown in Table 7 with their question numbers. The differences between WQ and FQ results are not significant (chi-square test, p < .05) and do not come close to statistical significance, with the notable exception of the term student (p = .073).
Percentages of Self-Reporting in the WQ and FQ for Four Variables for the Regionality Index 1–3
WQ = written questionnaire; FQ = fieldwork questionnaire.
It is interesting to note that the FQ overreports yods in avenue (83.9 percent) compared to the WQ (79.5 percent). Yod in this context is a Canadian marker compared to American English variants, and it seems plausible that the presence of a field-worker, doing a survey on Canadian English, triggered more Canadian responses in favor of yod. Coupon, which is no such marker, does not show the same phenomenon, and FQ underreports yod in this context compared to the WQ. Student, however, is overreported in FQ (38.2 percent vs. 25.0 percent in WQ), reaching almost statistical significance values. News, in addition, also overreports the yod-ful variants in FQ (33.9 percent vs. 26.6 percent), with a p value that is the second highest for these five variables (p = .14). The fact that these comparisons come close to significance may be an indicator that these two variables may not be as reliably polled with WQ surveys.
The Acoustic Analysis
All interviewees are native born and raised in metro Vancouver, 7 and, on Chambers and Heisler’s (1999) Regionality Index, they would score very low, that is, would be very local, with a score of 1 (both parents born in Vancouver), 2 (one parent born in Vancouver), and 3 (no parent born in Vancouver), but no higher than 3. Six informants are in the young age cohort (ages fifteen to twenty-six), two in the middle-aged group (forty and forty-four), and two in the older cohort (sixty-one and sixty-two). The interviews lasted for about thirty minutes. After an initial conversation about a topic that was shared between the interviewer (the author) and the interviewee, tokens were elicited from a reading passage, which were then used in the acoustic analysis. Table 8 includes the interviewees’ age and ethnicity.
Interviewees in the Sociolinguistic Interviews
AC = Asian Canadian; CC = Caucasian Canadian.
A few weeks after the interview, the interviewees were given a short questionnaire that included the following questions, intermingled with a number of dummy questions. For the low-back merger we asked,
3. Are the first parts in the words
7. Do the words
13. Are the names
And for yod-dropping, the questions were identical to those in Chambers (1998a:236):
8. Does the ending of AVENUE sound like
12. Does the u in STUDENT sound like the
15. Does the beginning of COUPON sound the same as
For these ten interviewees, a full set of their self-reporting answers is compared with their acoustic measurements. In addition, their scores can be compared to their group aggregate averages (based on Age and Regionality Index), shown in Table 7.
Low-Back Vowels
The merger of the low-back vowels is an ongoing linguistic change in North America, in which Canada has shown advanced features. Originally, the merger was most likely brought to Canada with the United Empire Loyalists in the eighteenth century (Dollinger 2010b; Wetmore 1971), but it has spread across Canada faster than in the United States, creating such homophones as cot and caught, Don and Dawn, Otto and auto. While the merger now covers more than half of North America, Labov, Ash, and Boberg (2006:65) assess that “[o]nly in Canada is the merger well enough established to show no correlation with age” (see also Boberg 2008b:135).
The cot-caught merger has been surveyed in self-reports for some time. In 1971 (Scargill & Warkentyne 1972), the merger showed a wide dissemination in Canada: respondents in their late thirties and early forties reported merged vowels at 85 percent for males, 85 percent for females; for the fourteen-year-olds it was 84 percent for males and 87 percent for females. These figures include Newfoundland, whose percentage at around 70 percent was much lower than mainland Canada, which decreases the national averages by 2–3 percent. Already in the early 1950s, when the adults of the study were completing their formative years, the low-back merger seems to have occurred more or less categorically (around 90 percent) across Canada, especially as these figures include first-generation parents, which are likely to decrease the merger count. There is further support for this interpretation. Gregg (1957:22) reports from “local” Vancouver university students that cot and caught and caller and collar are merged in a low-back vowel with “slight lip-rounding.” In most North American dialects, the merger usually occurs before /r/, merging the sets for historical NORTH (morning, for, war, born) and FORCE (mourning, four, wore, borne; Wells 1982). Where /r/ is in non-tautosyllabic position, many U.S. speakers have /ɑ/, as in sorry and tomorrow, and some also in horrible or orange (cf. Wells 1982:476). This U.S. innovation gives rise to stereotypical Canadian pronunciations. In Canada, sorry does not sound like sari /ɑ/, as the vowel retains its /ɔ/ quality (see Boberg 2008a:153).
Figure 7 shows the acoustic vowel plots for the ten informants for single tokens of sorry and sari. The data are not normalized, which means that comparisons can be made only by judging the distance between the two vowel heights for individuals but not between speakers. The acoustic analysis was carried out in Praat (Boersma & Weenink 2008) using F1 and F2 measurements at the vowel midpoint to measure the low-back vowels. The analyses are based on single tokens from a reading passage.

Vowel plots for caught/cot (non-normalized)
Figure 7 shows all ten interviewees at a glance. Mario, twenty-five, is a good example of a complete merger, while Ella, twenty-six, Carla, sixty-one, and Carl, sixty-two, have auditorily indistinguishable mergers but show some minor measurements differences. The difference in vowel height is detectable by Praat, but not for most hearers. Nancy, 40, is a borderline case, showing a bigger gap, and should be considered “close,” not merged, while Anton, 15, shows a distinct vowel sound for the two contexts.
Table 9 shows the interviewees’ self-reporting and the match with the acoustics. On the whole, apart from Ella, Anton and Nancy, seven out of ten self-report faithfully for cot/caught, with Ella being the only real case of wrong self-assessment. In the case of cot and caught, Ella is the only one who did not self-report her vowels to be merged. Clearly, her plot in Figure 7 shows the contrary. Nancy’s recordings, auditorily, appear to be very similar for both caught and cot. Anton’s data prove to be outliers. Anton, who was very nervous for most of the interview period, may have lost some control over his linguistic performance (he even said so after the interview), or, he may maintain the distinction in production, but not in perception. In any case, his case is a noteworthy exception.
Match Between Self-Reporting and Acoustic Analysis
For Don/Dawn (Figure 8), nine out of ten cases self-report correctly. Ella correctly asserts her linguistic idiosyncrasy by reporting a distinction between the two vowel sounds, which is shown in the vowel plot. Gustave, believing he pronounces the two names differently, is the one case here that is in error. Chad and Carla, in the older age cohorts, are close enough not to produce noticeable auditory differences.

Vowel plots for Don/Dawn (non-normalized)
Figure 9 presents the vowel plots for sorry and sari, the stereotypical Canadian maker, which shows some interesting behavior in the youngest speakers. For the most part, these variables show faithful self-reporting, with nine out of ten that can be considered accurate. Anton, again, is the exception, as he merges the two vowels but self-reports otherwise. Anton is not the only one who merges sorry/sari, with Lola, the nineteen-year-old female, coming close to a merger. She, auditorily, merges the two sounds, which might be triggered by unfamiliarity with the lexical item sari, or might represent a change toward raising in the younger age cohort. Carla, sixty-one, has a case of fronting sari, which may be the result of unfamiliarity with the token (as she mentioned).

Vowel plots for sorry/sari (non-normalized measurements)
Overall, the self-reporting involving the low-back merger is highly reliable. While the auditory data are needed to fully interpret the graphs, Figures 7–9 show that seven out of ten (cot/caught) and twice nine out of ten cases (Don/Dawn, and sorry/sari) are faithfully matched between the WQ and the acoustic analysis.
Yod-Dropping
Glide deletion in palatal glides in such contexts as student, news, tune, or duke has long been commented on in Canadian English (CanE). Until recently, the retention of the glide has been seen as a marker of CanE by both the public and linguists alike. As many commentators have argued, the yod-ful variants are vehicles for Canadians to assert their Canadianness in opposition to variants that are considered American: “When a particular pronunciation is clearly identifiable as American, the majority of Canadians tend to shun it without hesitation” (Orkin 1970:124).
The identification of what is American and what not, however, is another issue. Today, more reliable data are available than ever before. Three large-scale sociolinguistic studies report on glide deletion in three Canadian cities (Woods 1999 in Ottawa, Gregg 2004 in Vancouver, and Clarke 1993 in St. John’s, Newfoundland). In a well-argued article, Clarke (2006:234) notes that “the +glide variant is not the formal target for all groups” in Canada, which indicates that statements such as Orkin’s are either overly simplistic, or have become so since. Chambers (1998b:19) reasons, based on DT data, that glide-ful variants carry overt prestige, using the small percentages of yods in his data as evidence that in Canada “[y]od-dropping is not only common and standard but also unmonitored.” In the Ottawa data, females older than forty lead glide retention, a finding that is matched in St. John’s and Vancouver. Glide loss is led by males and blue-collar workers; that this pattern runs counter to the typical role of women in language change leads Clarke (2006:236) to propose a “change in indexicality,” as “glided and glideless variants have come to symbolize different social values for different segments of the Canadian population.”
A comparison of the self-report data with the acoustic measurements in Table 10 reveals very different scenarios for the three contexts. The duration for the /j/-glide was measured from the beginning of the periodic wave form until the start of F2 drop as the glide transitions into the back vowel /u/. The raw duration was then divided by the duration of the entire syllable to obtain a score that is normalized for different speech rates and allows for comparisons between speakers. This score is called the yod-ratio. The column “match” reports whether a participant’s self-reporting matches the acoustic measurement.
Glide Deletion in Vancouver: Yod-Ratio and Self-Reports
For the word avenue, for which around 80 percent and more reported yod-retention in Table 7, a robust match is shown between self-reporting and observed behavior. All interviewees self-report some degree of yod-retention, with the exception of Carl, 61, Anton, 15, and Lola, 19. Interestingly, Anton and Lola, the youngest in the sample, self-report the glide-less variable, while having fairly long glides, especially Lola (her yod-ratio of 0.348 is the second highest). Since all interviewees show some kind of glide, seven out of ten can be considered as reporting faithfully.
Coupon, which has been reported as a variable of divided usage (Chambers 1998a:243), shows precisely this kind of bimodal distribution in the data. The lexical item has a different history than student or news, which may partly account for its divided usage. From Anton, 15, to Nancy, 40, everyone is categorically glide-less; from Chad, 44, to Mario, 25, increasingly longer glides can be seen. The self-reporting is, however, apparently unaffected by the bimodal distribution and very reliable: eight out of ten self-reports match with the acoustic data. Carla and Kelsey are the two that are off.
Finally, student is a highly interesting case: taking out Carla, who could not decide in her self-reporting questionnaire, 9 only four out of nine self-report reliably. For the assessment of the categories glide-ful/glide-less, a cutoff point had to be introduced, since the phonetic environment of the alveolar stop /t/ invariably triggers some glide element. The threshold for yod-less and yod-ful variants was set at 0.15, where a quality change seems to occur (Table 10, gray shading). This cutoff is confirmed by Carla, who sits right at the transition point, which explains her confusion.
By taking all three contexts together, a most striking fact emerges. Without exception, all informants underreport their use of yods. Where informants self-report incorrectly, they overreport the glide-less variant, while, in reality, they show glide-ful pronunciations, as is highlighted by the gray shadings in Table 10. In all three contexts, in every single case of a nonmatch between self-reporting and the acoustic analyses (ten nonmatches), respondents have self-reported the glide-less variants as their target, while actually showing glide-ful pronunciations. This suggests that glide-less variants carry overt prestige. The result goes against many of the statements in the literature concerning glide deletion, such as Pringle (1985:190), who asserts that when Canadians
want to stress how their English differs in sound from American English, they are particularly likely to settle on these [palatal glides, i.e., yods]. In former years many school English texts which concerned themselves with matters of elocution regularly included advice to say “news” not “nooz.”
Judging from the self-report answers, the glide-less pronunciation, which has been part of CanE since its inception, has now become, for the ten Vancouver informants, the target variant carrying overt prestige, and they no longer adhere to shibboleths of British input, such as yod, that have become associated with anti-Americanism. Table 7 shows differences in the FQ data, which elicit more yods for student, news, and avenue than the WQ, suggesting influence by the field-worker.
Self-reporting for student is only badly matched in the acoustics. One reason might lie in its lexical status. For yod-dropping in news, Clarke (2006) suggests a social recoding for a yod-ful prestige variant in North America. This reindexicalization of yod-ful news is not a Canadian phenomenon, as it is found in American media as well, and has therefore little to do with Britshness in CanE. In U.S. media, news retains the glide to give it more “respectable formality” (Pitts 1986:136, qtd. in Clarke 2006:243).
A similar process seems to account for the case of student and one can conceive of it as an exception to generally glide-less Canadian target forms based on the social importance of the lexical item for some groups. Similar to news in the media, st/ju/dent, as opposed to st/u/dent, may have taken on a specialized indexicality for “learning.” Carl, 61, and Lola, 19, both live on a university campus and are part of the academic community. They correctly self-report yods. Kelsey and Ella are undergraduates, Anton a high school student, and Nancy an art teacher: they are all part of a larger workspace of education and learning. As the data show, they all retain their glides without being aware of it. Chad, 44, in contrast, has no academic connection or pretensions and is, correctly, a self-reported yod-dropper. Carla is a secretarial assistant who staffs a front desk. She is surrounded all day by literature professors and was “confused” as to what form(s) she uses. She is, importantly, precisely on the cutoff value in her measurement. Mario, however, is a law student who correctly self-reports yod-dropping. He is very ambitious and focused on building himself; he is a person who has a keen interest in Canadian politics and is more than likely to run for office one day. The indexicality explanation, while not perfect, accounts for nine out of ten cases, and is thus preferable. The best reason I can give for Mario’s behavior is that he chooses not to code student for “learning” to not alienate parts of a potential voter segment—he has the education, but does not “show” it with a yod-ful pronunciation. This, however, would need to be supported with further evidence.
Discussion
The comparisons of the WQ and FQ, and the WQ and interviews, suggest that WQs have their place in linguistics. Some preference can be given to WQ results when compared to FQ results because of a more gradual progression of changes, which appears to be more typical and more representative of linguistic change in real-world scenarios. This preference is plausible, but will need to be cross-checked against observed linguistic behavior.
In terms of linguistic levels, not one area can be singled out as being more prone to deviations than any other: the list of big deviations (≥ 10 percent) is led by a syntactic variable, followed by morphological (three times), lexical (three times), and phonetic variables (three times). Phonetic variables are the only ones that reach significance levels of difference between the WQ and FQ, but they are also among the most closely matched ones.
The age of the informants should be considered as a potential factor. Evidence was found that the older age cohorts, those in their sixties, seventies, and eighties or older, are more prone to deviations than the younger cohorts. Sample sizes play a role in the reliability of the surveys, and it is recommended that at least twenty-five (as seen in the sixty-year-olds) but ideally fifty participants per cell should be polled in WQ surveys. This may be challenging for some age or social groups, but it seems to control one problem. These sample sizes put WQs in rough proximity to the recommendations for random sampling, 10 as laid out in Schilling-Estes (2007:168): for populations smaller than 1,000, the sample size should be 300, for populations greater than 150,000, which is the case in metro Vancouver and other cities, the sample should be 1,500 and more. If we were to aim for 50 participants for both men and women in eight cohorts, a sample of 800 would be a minimum. Some DT studies have worked with about 1,000 respondents, while others, such as the ones reported here, made do with 400 to 500 respondents. In light of the current findings, this might be suboptimal.
The length of the questionnaire is a point worthy of further attention. In the present study, it was suggested that the reliability would be increased in the older age cohorts (fifty and older) with shorter versions.
WQ versus FQ
Evidence for the usefulness of WQs is increasing. Bailey, Wikle, and Tillery (1997:57) suggest that “self-reports might be more valid and reliable measures of linguistic behavior than linguists have supposed.” In their comparison of four data sets, very similar frequencies of self-elicited variables are found and are taken as evidence for equivalence. For socially stigmatized variables, such as the double modal might could, they suggest that self-reporting is a more reliable method “of linguistic behavior than observations of usage” (Bailey, Wikle, & Tillery 1997:58). Bailey, Wikle, and Tillery, however, exclude phonological variables in their case for self-reported data. Chambers (1998a:244) pushes the agenda one step further, stating that “in alleviating field-worker bias and minimizing the Observer’s Paradox, it must be conceded that it [the WQ] is in some ways more reliable than field-worker interviewing.” By omitting the observer, it is clear that one reduces the Observer’s Paradox. With no observers physically present, it is only the informant’s awareness of answering for someone else that might interfere. The anonymity of data collection in the WQ, where no data could identify an informant personally, is an incentive to report as faithfully as possible. As Bailey and Tillery (1999:395) have shown, field-worker bias can distort survey data. In the case of the Linguistic Atlas of the Gulf States, one interviewer, Barbara Rutledge, conducted almost 18 percent of all interviews, and Bailey and Tillery illustrate that her interview style elicited might could significantly more often by using a more direct elicitation style than other field-workers. Self-reporting techniques are prone to produce better data than trying to observe the double modal might could in the interview situation.
There is more support for self-elicitation in linguistic work. In the area of lexis, Pratt (1983:151) suggests that “in certain kinds of lexical investigation direct questions are not only unavoidable but even highly desirable.” Pratt sees low-frequency phenomena, such as lexical items that are outside of the common core or register, as most appropriate for a particular interview situation. To count as a real occurrence, and not just as the interviewee trying to please the interviewer, he suggests that the interviewee needs to “display his knowledge of it [the word]” (Pratt 1983:153).
In the Canadian context one is reminded of pragmatic marker eh? and its elicitation in self-report questionnaires in recent studies (Gold 2008). Stigmatized in some or in all functions by some speakers, such variables can relatively easily be elicited with WQs that yield a good idea of the distribution of the features in the survey area that can be followed-up with more fine-grained analyses.
Self-Reporting and Observed Linguistic Behavior
The previous section examines phonological variables and offers a difficult test case, in which phonological information gained from self-report written surveys is compared with phonetics data from audio-recorded and spectral-analyzed data. For pronunciation variables, researchers have doubted the usefulness of self-reports. Bailey, Wikle, and Tillery (1997:59) state they “would be hesitant to rely upon self-reports for phonological data.” Their minimal pair data compare well with the reading passage that the sociolinguistic interviewees were given in the present study and so might allow a reassessment of their skepticism. Chambers (1998a) expressly aims to show that variables such as yod-dropping can be elicited, claiming that “it is no more difficult to elicit certain pronunciation variants from postal respondents than to elicit lexical items” (236).
The comparison of self-reporting and acoustic measurements in this study is largely congruent. First, the phonological variable of the low-back vowel matches with self-reported data. Here, it was shown that at least a solid match (seven of ten for cot/caught) between self-report and observed behavior could be found, in the other cases an impressive congruence could be seen (Don/Dawn and sorry/sari, both nine out of ten). While, as Labov (2001:517) argues, mergers generally “run their course without taking on symbolic value of any kind,” the sorry/sari distinction is a highly conscious variable in Canada and is reported very faithfully by the ten respondents as well as in the overall WQ and FQ samples, with 84.0 percent (WQ) and 86.5 percent (FQ) of local Vancouverites reporting “different” vowel sounds in the pair. For sorry/sari, the acoustic analyses of the youngest informants suggests the possibility of a future merger in a new, slightly raised, mid-low-back position, that is, sari, with /ɑ/, pronounced more like Canadian sorry, with /ɔ/, and therefore still different from an American merger in a lower position. A bigger sample, however, would have to verify this effect.
The discussion of yod-dropping shows errors in self-reporting of the ten interviewees by overreporting of the yod-less variant, contrary to the expectation that Canadians would tend toward a yod-ful variant that fits with traditional norms mimicking British RP. This is somewhat in line with Chambers’s (1998b:19) assessment that in CanE yod-dropping is “not only common and standard but also unmonitored.” The data of yod-ful variants in the observed behavior show a consistent error in underreporting yod-ful variants and suggest a change in indexicality (Clarke 2006). It seems clear that the general Canadian norm is now the yod-less variant, with only small percentages for variants that preserve the yods. These exceptions seem to cluster around lexical items that are indexed for some social feature such as “elegance and good breeding” for— n/ju/s (Clarke 2006:242) and, as I suggested, “learnedness” for st/ju/dent. Word frequency plays an effect too. As Phillips (1994:120) has noted, the less frequent a lexical item, the less yod-ful it is perceived. Student, avenue, and coupon, alongside news, certainly reach high lexical frequencies. The matches between self-reporting and acoustic measurements are good for avenue (seven of ten) and coupon (eight of ten), with student not faring as well, with only four of ten. In the latter case a social interpretation is possible, considering the ties of the individuals to academia.
The case of student may be further complicated since linguistic change appears to be in progress. For the same reason that Bailey, Wikle, and Tillery (1997) are doubtful of self-reports for sounds that undergo a number of changes in progress, self-reporting methods must be used with caution for linguistic changes that are socially coded, such as a reindexicalization of student. One must consider this caveat when using WQ methods as a survey tool for an area in which a particular variable has not yet been studied independently.
Overall, self-reporting should be seriously considered as a legitimate data collection method, especially so for dialectological projects. While some phonetic nuances cannot be operationalized in written format, basic phonemic differences can be fruitfully employed. The acoustic data independently confirm this assessment but also show that some caution is needed when dealing with the possibility of vigorous, socially driven changes in progress.
Self-reports in the form of WQs are not likely to go away in the information age; on the contrary, new technologies could help revitalize them. As this study suggests, if applied consciously, there is no reason to throw out this tool. The sheer range of speakers one can reach with an online questionnaire is unparalleled and provides macro-level data that are bound to have an impact on the practices of social dialectology and other disciplines. While other techniques are now increasingly being used in undergraduate courses, such as interview transcription or acoustic phonetics, self-reporting remains a useful tool that is unmatched in its efficiency. If applied consciously, linguistic self-reporting can enrich our knowledge of dialects in a most economical way and enable us to introduce the method in our lower-level classrooms. WQs are one good survey tool among many. Like any tool, self-reports have their strengths and weaknesses. While the results have highlighted some problematic areas, the method’s general validity has been confirmed by producing insights on the CanE low-back vowel merger and yod-dropping.
Footnotes
Appendix A
Statistical Tests
| # | Variable | Normal? | t-value | df | p value | Differences? |
|---|---|---|---|---|---|---|
| 6 | lever | Yes | −2.5882 | 7 | .018 | Significant |
| 8 | avenue | Yes | −2.05 | 7 | .039 | Significant |
| 17 | zed | Yes | −1.903 | 7 | .049 | Significant |
| 19 | presentation | Yes | −1.916 | 7 | .048 | Significant |
| All | overall, all 32 variables | No | Shapiro–Wilk test for Normality: W = 0.9803, p value = 2.829e-05; Wilcoxon test: V = 33739.5, p value = .163 | Not significant |
Normal distribution was checked with the Shapiro–Wilk test (criterion p = .05 or smaller). The table shows the results of the paired t-tests or Wilcoxon tests (where normality was not given) for variables reaching significant differences between the written questionnaire and the fieldwork questionnaire, and the overall comparison.
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
