Abstract
In view of the ubiquitous increase in the use of C-tests, which are almost unanimously believed to measure general language proficiency, this study investigates whether the aspects of language proficiency tapped into by the C-test format are the same when the test is taken by a learner population other than that of foreign language learners. Specifically, we conducted a differential functioning analysis and compared the types of mistakes that 113 foreign language learners of Russian made when completing C-test gaps, with the performance of 89 heritage language learners on the same C-test. The results showed that almost half of the C-test gaps are biased towards either learner group. In addition, the error analysis for a number of the biased items demonstrated that, although heritage language learners seem to have an advantage in reconstructing the meaning of C-test gaps, they fail to translate their recognition skills into producing the right form. Furthermore, the study reveals a possible sensitivity of the C-test construct to the traditionally used dichotomous scoring method. We conclude with a discussion that includes the implications of the results regarding the construct measured by the C-test and the possible consequences for its actual use.
Keywords
The C-test format, a specific type of cloze test, was originally developed by Klein-Braley and Raatz (1984) as a test of global language proficiency. It has become a widely used instrument in various educational and research contexts. Grotjahn’s (2016) C-test bibliography, which provides a comprehensive overview of C-test studies conducted for different languages, includes more than 500 entries in which the C-test format is either the focus of the publication or used as an instrument for answering a specific research question in second language research. The widespread use of the test format is not surprising since a C-test is a rather economical procedure in comparison to other tests. Furthermore, C-test scores have been shown to be capable of ranking examinees according to their proficiency in the tested language. Therefore, C-tests are widely used as screening and placement instruments (e.g. Eckes, 2014; Mozgalina & Ryshina-Pankova, 2015; Norris, 2008). Especially in Germany, C-tests appear to be the most commonly used instruments for assigning students to language classes at university language centers.
Most previous studies on the C-test format have been conducted with foreign language learners (FLLs) and have primarily focused on C-tests in English, Spanish, French, and German, with only two publications (Drackert, 2015; Steurer, 1986) dealing with a Russian C-test. In the context of Russian as a foreign language, however, a large number of students in different learning contexts are not typical foreign language learners who learn Russian exclusively in an instructional setting. Rather, they are heritage language learners (HLLs) who at least partially acquired some knowledge of Russian at home prior to instruction. Without evidence that the C-test works for this target group in a similar manner as for FLLs, caution is advised when generalizing the research results to the HLL population, as well as when using this test format to establish HLLs’ general language proficiency and making decisions based on the test results.
To the best of our knowledge, no studies have explicitly examined the C-test performance of heritage language learners, a group widely present in L2 classes of Russian as well as other languages worldwide (e.g. Spanish in the USA, Turkish and Polish in Germany). This exploratory study attempts to address this gap by investigating what kind of language proficiency or aspects thereof the C-test taps into if taken by Russian HLLs with their unique language proficiency profiles. Answering this question is also potentially relevant for the use of C-tests with other populations of heritage learners.
C-test and its construct
Like classic cloze tests, C-tests are written integrated language gap-tests based on the principle of reduced redundancy (Klein-Braley, 1997), which is operationalized by systematically deleting parts of the words in a number of short authentic texts on different topics. A C-test usually consists of between four and eight texts which normally appear in order of difficulty, which is generally determined according to a scale such as the CEFR or a language curriculum.
In a C-test constructed according to the classic deletion rule, texts are mutilated by deleting the second half of every second word, beginning with the second word of the second sentence until 20–25 deletions per text have been made. In a word with an odd number of letters, the larger part is deleted (Grotjahn, Klein-Braley, & Raatz, 2002). The first and the final sentence of each text are usually left intact in order to provide necessary semantic context.
In most settings, C-tests are scored dichotomously, where one point is awarded for the correct reconstruction of a gap on all levels, that is, semantically, grammatically, and orthographically. This scoring system has proved to be especially practical in online applications of the C-test, where the answers are scored automatically. Figure 1 shows an example of an automatically scored Russian C-test text on Moodle. 1

Example of a C-test in Russian.
In the context of testing German as a second language, Baur and Spettmann (2007) introduced the use of an additional score, which they call the word recognition value (WE-Wert). It is awarded for semantically correct answers and accounts for the receptive skills of C-test candidates in cases where the former do not correlate with their productive skills. Used alongside the traditional binary scoring system, this score allows the diagnosis of writing deficits and identification of the additional support required by the students (p. 100). However, for practical reasons, the recognition score has not been implemented in an online environment.
Each text of a C-test can be viewed as a sample of the language tested because the systematic deletion leads to a representative sample of all the linguistic features of the given language being deleted (Sigott, 2004). Test candidates have to restore the deleted parts of the words in each text using their lexical and grammatical knowledge of the language as well as their text-processing skills. The higher the level of proficiency of a testee in the given language, the more gaps they are able to complete correctly.
In the research literature, the C-test is presented as an appropriate instrument for measuring general or global language proficiency (GLP): “C-tests have proved to be objective, highly reliable and very economical means for measuring global language proficiency” (Grotjahn, 2012, p. 181). These claims are based on numerous attempts undertaken to gain a better understanding of the construct measured by the C-test format (e.g. Eckes & Grotjahn, 2006; Grotjahn & Schiller, 2014; Sigott, 2004; Stemmer, 1991). Some of the research methods used have been, for example, correlational and factor-analytic analyses, logical task analysis, introspective research into the cognitive processes involved in solving a C-test, and the analysis of learner responses.
C-test scores in various research studies were correlated with the scores obtained through a number of standardized tests of different language subskills. Low/moderate to high correlations were found, both for oral/aural and written skills, for example 0.64–0.68 for TestDaF or 0.36–0.88 for TOEFL (see Eckes & Grotjahn (2006) for an overview of the correlational studies). For the researchers, these low/moderate to high correlations with different subskills served as indicators that the C-test measures general language proficiency. Correlational analysis conducted specifically for a C-test of Russian also showed a rather high significant correlation (r = .79) between this test and an Elicited Imitation test that targets aural and oral skills (Drackert, 2015).
Since the studies summarized by Eckes and Grotjahn (2006) and the study by Drackert (2015) were conducted with foreign language learners, comparable levels of L2 proficiency in both the oral and written modes were to be expected based on the instructed language acquisition which the learners had been involved in. Thus, the obtained correlations used for demonstrating that the C-test targets general language proficiency do not provide any evidence that a C-test taken by heritage language learners would target the same construct of GLP.
In order to investigate the construct of general language proficiency measured by a German C-test, Eckes and Grotjahn (2006) conducted Rasch measurement modelling and confirmatory factor analysis and obtained strong evidence that the C-test they investigated represents a unidimensional instrument which measures the same general dimension as the four TestDaF sections (reading, listening, writing, and speaking) together. At the same time, the researchers showed that the GLP measured by the C-test format is divisible into more specific constructs, defining general language proficiency as “an underlying ability comprising both knowledge and skills and manifesting itself in all kinds of language use” (Eckes & Grotjahn, 2006, p. 291).
To provide additional information on the construct of a C-test, Stemmer (1991) used a thinking-aloud method to understand which processes are involved in solving C-test items in French. First, she found that the construct measured by the C-test format depends on the characteristics of individual C-test texts, namely on the number of completion possibilities, the number of cohesive ties in the text, as well as the length of the meaning units in the text (p. 327). Furthermore, her analysis showed that “knowledge about texts or specific content knowledge was much less often activated than knowledge about L2,” thus indicating that the C-test format favors bottom-up processing. Stemmer concluded that if the proposed general language proficiency format includes higher level comprehension skills, then the C-test format does not measure this kind of GLP.
According to results obtained by Grotjahn and Schiller (2014), the use of top-down versus bottom-up processing of C-test texts depends on the level of GLP of the test candidates. They analyzed learner responses to 12 individual gaps of a Spanish C-test and found that learners with a higher total score on the C-test, that is, more proficient learners seemed to “make better use of the context than those with a lower score” (p. 275). They interpreted this observation as evidence that a C-test can be considered a valid measure of general language proficiency, although they did not specify what other aspects the construct of GLP comprises besides utilizing the context.
Sigott (2004) also approached the question of the C-test construct by investigating how much context C-test takers need to process when they reconstruct a C-test gap; in other words, whether they solve an item at the word, phrase, clause, sentence, or text levels. He found that different learners tend to apply different strategies to solving a C-test, and that the same C-test gap can be completed correctly by means of low- or text-level processes depending on the individual learner. Furthermore, contrary to Grotjahn and Schiller (2014), Sigott (2004) observed that the more proficient the learners are, the more use they make of lower-level processes. He concluded that the C-test construct is rather complex and includes orthographic, morphological, lexical and syntactic knowledge, the knowledge of such text properties as cohesion and coherence as well as the ability to process language at all levels from the individual letter to the text (p. 200). At the same time, he emphasized that C-test texts may elicit these aspects of knowledge to a different extent depending on the test takers’ language ability and passage difficulty. This unstable quality of the C-test led Sigott (2004) to conclude that a C-test has a fluid construct.
The existing research that has addressed the question of the construct of a C-test by investigating learner performances as reflected in their test scores, or by examining individual gaps, provides insights into different aspects of general language proficiency in terms of the knowledge and skills needed to complete a C-test. However, few of the studies actually define the construct of GLP as measured by a C-test. Even researchers who did define GLP in terms of knowledge and skills (compare the definition by Eckes and Grotjahn above) did not specify which particular aspects of knowledge and skills are targeted by C-test gaps. At the same time, research has clearly demonstrated that at least two types of knowledge are required to reconstruct an individual gap: lexical knowledge (i.e. the meaning of the word) and grammatical knowledge including orthography (i.e. the form of the word). These have to be activated simultaneously, or, as Hastings (2002) put it, correct responses to C-test items “can be achieved only by applying and integrating several facets of language competence, each of which acts as a check on the others” (p. 54).
To conclude, the existing research on the C-test construct has been conducted with FLLs who learned the language in question through instruction; in other words, in a setting where language learning typically takes place simultaneously in written and oral mode. Thus, previous research disregards the settings in which the language is acquired (primarily in an oral mode) prior to instruction, as is typical of HLLs. However, these settings are bound to produce different proficiency profiles which reflect a different kind of ability, and this, according to Sigott (2004), will influence the construct. In other words, theoretical knowledge about the construct of a C-test is limited to a certain learner population, namely FLLs. However, in reality the test is frequently used with a different learner population, namely HLLs. Yet the results are still interpreted as a measure of the same construct, namely, general language proficiency. This clearly ignores the generalization argument in Kane’s (2006) argument-based approach to validity and emphasizes the need for separate studies on the C-test construct with heritage language learners.
Language proficiency of heritage learners
Benmamoun, Montrul, and Polinsky (2010) defined heritage speakers as “early bilingual speakers of ethnic minority languages who have differing degrees of command of their first or family language, ranging from mere receptive competence in the first language to balanced competence in the two languages’’ (p. 8). Typically, heritage speakers were born outside the parents’ home country or left the home country at an early age. At home, at least one person speaks to them in the heritage language or spoke to them when they were children. However, beginning in kindergarten and continuing through middle and high school, heritage speakers become increasingly proficient in the majority language and use it at the expense of the home language, sometimes even when talking to their parents. In many cases, heritage speakers do not acquire written skills in their home language or do it unsystematically and as a rule to a lesser extent than children who receive school education in their L1 (Benmamoun, Montrul, & Polinsky, 2013; Zyzik, 2016).
Language acquisition of heritage learners results in the following proficiency profiles: first, heritage language learners have (near) native-like phonetics and quite extensive vocabulary in comparison to FLLs (Mehlhorn, 2016); and, second, because the language is not acquired systematically in a written mode, HLLs have difficulty with reading and writing in their home language and show deficiencies on different linguistic levels including morphology and syntax (Benmamoun et al., 2013; Brehmer, 2007; Polinsky, 2006). Very often these deficits are manifested in the realm of academic skills in general and academic vocabulary in particular (Meskill & Anthony, 2008). Furthermore, in languages with irregular grapheme–phoneme correspondences, such as Russian, in which one grapheme can represent different consonants or vowels depending on their phonological environment, HLLs tend to have additional difficulties with correct orthography (Böhmer, 2011). Notwithstanding these typical common features, it is worth mentioning that HLLs are a very heterogeneous group.
Describing the prototype model of a HLL based on research on Spanish, Zyzik (2016) came to the conclusion that the language proficiency of a prototypical HLL is limited to Basic Language Cognition (BLC) and implicit knowledge of the language. Thereby she drew on Hulstijn’s concept of BLC, which relates to “(a) the largely implicit, unconscious knowledge in the domains of phonetics, prosody, phonology, morphology and syntax; (b) the largely explicit, conscious knowledge in the lexical domain (form–meaning mappings), in combination with (c) the automaticity with which these types of knowledge can be processed” (Hulstijn, 2011, p. 230). BLC is the type of language ability shared by all native speakers of a language. However, BLC is limited to speech reception and speech production and does not comprise the literacy skills of reading and writing. These are part of Higher Language Cognition (HLC), which contains understanding and production of “low-frequency lexical items or uncommon morphosyntactic structures” and includes “topics other than simple everyday matters” (Hulstijn, 2011, p. 231). Since HLC is developed primarily as a result of language use in academic and professional contexts and is thus a function of the levels of formal education, native speakers will differ significantly in their HLC. By the same token, the language proficiency of a typical HLL will be comparable to that of native speakers with limited formal education who do not use written language in everyday life (Zyzik, 2016, p. 22). However, there will also be considerable variability within the group of HLLs, depending on the level of formal education they have received in their heritage language. Having native-like skills in BLC will give HLLs an advantage over FLLs on oral and aural tasks; also, their receptive knowledge of the core vocabulary will surpass that of most FLLs (Zyzik, 2016, p. 30). However, a typical HLL will have no productive use of vocabulary beyond the mid-frequency range; that is, many HLLs will have partial knowledge of such words and will not necessarily be able to use them productively in the appropriate context (Zyzyk, 2016, p. 30).
It seems reasonable to assume that, since a C-test is conducted in the written mode, it will tap into at least some aspects of HLC. Moreover, a test taker’s ability to recognize the meaning of a gap largely drawing on their BLC resources is secondary to their ability to produce its correct written form (an HLC ability). Knowing that HLLs have language proficiency profiles where BLC clearly dominates, whereas FLLs normally acquire both components simultaneously, it must be assumed that the results of a C-test will put HLLs at a disadvantage, as they will not be able to display at least some of their BLC proficiency. The main goal of this study is to test this assumption.
Research questions
The following research questions (RQs) were posed in the study:
RQ1: Are heritage and foreign language learners equally successful in reconstructing the same C-test gaps?
RQ 2: In what ways do the answer patterns on specific gaps differ for heritage and foreign language learners?
Methodology
Instrument
The Russian C-test used in the study was developed and validated for use as a placement test at the language center of a large public university in Germany. It consists of five texts of between 67 and 100 words, with each text approximately corresponding to a different proficiency level of the CEFR (A1/A2–C1). 2 The texts cover different topics and are taken from authentic sources.
Three C-test texts 3 were selected for the purposes of the present study: the first C-text is about a Russian friend and his studies and corresponds approximately to levels A1–A2; the second text (presented in Figure 1) concerns the results of a research study on the influence of internet use on the brain activity of elderly people and corresponds approximately to levels B1–B2; the third text is about the emerging middle class in Russia and corresponds approximately to levels B2–C1. Thus, the first text is an example of private discourse; the second and the third texts belong to the field of public academic discourse. The difficulty measures of the C-test texts range from −.99 to +.70 logits based on the Rasch analysis. 4 Each text contains 20 gaps, thus 60 points was the maximum that the students could obtain on the three C-test texts used in the study.
Data
Participants in the study were 113 FLLs and 89 HLLs of Russian at two universities in the USA and Germany. The USA data (N = 65; NFLLs = 65) are taken from a larger test validation study conducted by Drackert (2015), whose participants attended Russian courses as an obligatory part of their studies at a private university in the USA. Most of them were native speakers of English, their ages ranging from 18 to 30. They were enrolled in seven different curricular levels, from the second semester of Russian to the advanced Russian course. Based on the results of the Russian Speaking Test (Center for Applied Linguistics, 1998) and/or local placement procedures, their levels of proficiency were estimated to range from Novice-Mid to Advanced (ACTFL levels), with two students at the level of Advanced-High and Superior.
The data from the German educational setting (N = 137; NHLLs = 89; NFLLs = 48) come from a database of C-tests completed as a placement test at the language center of a large German university between 2013 and 2015. The test takers were university students of various departments wishing to attend Russian courses from A1 to B2 as an optional part of their studies. Based on the institutional placement procedures (including the C-test), their proficiency levels ranged from A1 to C1. Most of them were German native speakers aged from 20 to 33.
Both FLLs and HLLs were identified based on their answers to a placement questionnaire regarding their family language and previous exposure to Russian. Those students who reported Russian to be their parents’ language and admitted using it with at least one family member were identified as HLLs. The group of HLLs within the sample was rather heterogeneous, with their formal education in Russian ranging from zero hours to the first one or two years of schooling in Russia. However, what distinguishes this heterogeneous group from the group of FLLs remains clear: learning to read and write in Russian with a delay as opposed to learning all four skills simultaneously. There was no statistically significant difference between the two groups (FLLs M = 22.31, SD = 13.21 and HLLs M = 25.25, SD = 16.09) in terms of their performance on the C-test scored using the traditional method as determined by a t-test: t(201) = −1.424, p = .156. Furthermore, there was no statistically significant difference between the students from the USA (M = 25.52, SD = 9.44) and Germany (M = 22.76, SD = 16.39), as determined by the t-test: t(201) = 1.263, p = .208. We are aware of the limitations of determining proficiency based on the C-test scores as compared to using an external instrument. However, we chose to investigate the construct of the C-test by analyzing learners’ mistakes to obtain a deeper understanding of the performance underlying these scores, which would not be possible if an external criterion were used.
Analyses and procedures
A range of qualitative and quantitative analytical approaches were employed in order to explore the research questions posed above. First, we analyzed whether individual C-test items function significantly differently for FLLs and HLLs (RQ 1) by examining differential item functioning (DIF) within an IRT 5 analysis using WINSTEPS software. DIF identifies the ratio of individuals at the same ability level (determined by the person’s measure on the logit scale) who answer a particular item correctly. If an item measures the same ability in the same way across groups, the groups should display the same success rate regardless of their nature. Items that differ in success rates for two or more groups at the same ability level are said to display DIF (Holland & Wainer, 1993).
When a DIF analysis is conducted, it can be hypothesized that an item has the same difficulty for two groups except for measurement error. Items that indicate bias have three statistical characteristics: the DIF contrast defined as the difference in difficulty of the item between the two groups should be at least 0.5 logits for DIF to be noticeable; biased items have a t-value of 1.96 or more, and a probability of .05 or less (Draba, 1977).
DIF occurs when some factor apart from the test construct affects the performance of one group but not the other. Hence, for this study, DIF analysis was conducted in order to investigate if the item difficulty is the same for HLLs and FLLs irrespective of how they learned Russian. To conduct the DIF analysis, we first coded learner answers on 60 C-test gaps for all participants in two populations using the classical scoring system. One point was given for the correct answer and zero points for an incorrect answer or a gap left blank.
DIF analysis does not explain why an item benefits one group over the other; it only shows that the difference is systematic to the groups and mathematically significant. Therefore, we additionally deployed error analysis on some of the items in order to look for explanations of DIF. To account for parallel and interdependent reconstruction of the meaning and form of an individual gap, we focused our error analysis on the C-test gaps which represent different parts of speech (nouns, verbs, adjectives, and conjunctions) and different grammatical categories (gender, number, case, and tense). Essentially, we chose the items where the learners not only had to recognize the meaning of the gap, but also had to demonstrate their ability to manipulate the form correctly within a syntactic context. For this reason, we excluded non-inflected functional words. Furthermore, the selected gaps were roughly comparable in terms of difficulty. Middle-range items were chosen because items that are too easy or too difficult fail to differentiate sufficiently between test takers.
Classification of learner responses
In order to investigate the patterns of answers on the analyzed items, the answers first had to be classified. Initially, each of the authors independently inspected learner answers on the chosen items and developed her own classification of answers. Subsequently, the two classifications were compared, differences were discussed and one common classification (presented in Table 1) was agreed upon.
Classification of learner responses. a
The examples of the answers to each of the analyzed gaps can be found under https://tinyurl.com/yxa2vezm.
As can be seen from Table 1, a total of five categories were identified from the analysis of learner answers. Gaps that were not completed were considered as belonging to category 0. Category 1 basically includes responses that indicate that test takers were not able to reconstruct the meaning of the gap correctly. For example, for the gap zapa
Responses in category 2 directly illustrate learners’ failed attempts at reconstructing the correct form of the gap. In other words, the meaning of the word and the part of speech are recognized but the grammatical form of the word is reconstructed incorrectly. For example, for the gap zapa
The two remaining categories (3-4) include learner answers in which both the meaning and the grammatical form of the word were reconstructed correctly. Category 3, however, comprises learner responses in which the lexeme was recognized and the word was manipulated correctly to fit into the grammatical context but there was a spelling mistake in the word root. In these cases, the test taker has demonstrated lexical and grammatical knowledge but not orthographical skills. For example, for the gap prepod
Using the developed classification, both authors together examined and coded a total of 3838 answers to the 19 gaps given by 202 learners. Subsequently, the distribution of the types of responses across two groups (RQ2) was analyzed using cross-tabulation. Since the groups were different in size, we calculated the results in % that is, we compared what percentage of the mistakes belongs to a certain category for two learner groups.
Results
Research question 1
Research question 1 investigated whether item difficulty of the individual C-test gaps varied for heritage and foreign language learners. Overall, as can be seen in Figure 2, the item difficulty for FLLs ranged from −6.18 logits (the easiest item) to +4.28 logits (the most difficult item). The range of item difficulty in the HLL group, as displayed in Figure 3, was between −7.14 logits and +3.21 logits. In other words, the easiest items in the C-test were generally easier for HLLs and the most difficult items were more difficult for FLLs.

Item–person map (FLLs).

Item–person map (HLLs).
The easiest items
7
for FLLs were item 2 o
The DIF analysis demonstrated that a considerable number of items behaved very differently for the two groups in terms of their difficulty. Out of 60 items, 28 showed DIF of more than 0.5 logits and a t-value of more than 1.96 (see Appendix 1). However, the biased items did not benefit only one group. In particular, as summarized in Appendix 2, a total of 12 items were easier for FLLs and 16 items were easier for HLLs. For example, gaps such as prepod
Generally, it can be said that all of the biased items which benefit FLLs appear in the first text (with the exception of jaz
In order to account for the difficulty of the C-test texts, we additionally estimated DIF for each self-contained C-test text as a super-item (Norris, 2008) scoring it on a scale of 1 to 20. As displayed in Table 2, a significant difference in performance (DIF contrast of 0.71 logits) between HLLs and FLLs was found only on C-test text 1 (item 1).
Results of the DIF analysis for C-test texts.
Note: Width of Mantel-Haenszel slice: MHSLICE = .010 logits.
Research question 2
Since DIF analysis does not explain the reason behind the variance in performance of the two learner populations, we compared actual answers of both groups using error analysis. Out of the 28 biased items we chose 12 gaps, half of them benefitting HLLs (stimul
Figure 4 represents answer patterns for the six items that benefit HLLs. For these items, HLLs submitted 26.37% more correct answers than FLLs (43.97% and 17.60%) and demonstrated the recognition of the meaning of the gap nearly twice as often as FLLs (31.50% vs. 58.25%), as reflected in the categories of “correct answer,” “spelling mistake,” and “wrong form (agreement)” considered together. Both groups experienced comparable difficulty in producing the correct form of the lexemes since 13.40% of the mistakes in both groups fall into this category.

Distribution of answer patterns for the items that benefited HLLs.
A different pattern emerges in the learner responses for the group of items which benefit FLLs. As can be seen in Figure 5, HLLs and FLLs recognized the meaning of these words equally often, namely in approximately over 70% of cases (the categories “correct answer,” “spelling mistake,” and “wrong form (agreement)” considered together). However, FLLs submitted 15% more correct answers (49.07% vs. 34.57%), whereas HLLs made more mistakes, not only in the grammatical form, but also in spelling in the word root (37.91% vs. 23.45%).

Distribution of answer patterns for the items that benefited FLLs.
In the group of unbiased items, the difference between the learner groups appears less pronounced (see Figure 6). This is to be expected since the items do not differ statistically significantly in their difficulty. It should be noted, however, that HLLs submitted correct answers more often than FLLs (39.90% vs. 30.82%) and made slightly fewer mistakes in the agreement (10.97% vs. 15.63%) but slightly more spelling mistakes in the word root (3.62% vs. 1.38%).

Distribution of answer patterns for unbiased items.
A comparison of the answer patterns represented in Figure 6 to those in Figures 4 and 5 shows that the amount of variation in the performance of HLLs between biased and non-biased items is less evident than that of FLLs. FLLs appear to benefit much more from items that favor them (49.07% correct answers as compared to 30.82% on unbiased items) than HLLs (43.97% correct answers as compared to 39.90% on unbiased items).
Discussion and conclusions
In this study, we aimed to scrutinize the construct of a C-test and the appropriateness of its use with heritage language learners by conducting a DIF analysis and a subsequent error analysis of a selection of the gaps for two learner populations: foreign and heritage language learners. According to previous research conducted exclusively with FLLs, a C-test measures general language proficiency. However, it has not been investigated whether a C-test measures the same aspects of GLP if taken by HLLs, a group that is widely represented in Russian classes in different educational contexts.
The DIF analysis demonstrated that a considerable number of items appear to present a differing degree of difficulty for the two learner populations. Several factors appear to play a role in making an item more or less difficult for one of the groups: the frequency of a lexeme and its semantic complexity, the frequency of a grammatical category, the semantic and syntactic context (text difficulty) in which the word appears, and the degree of discrepancy between the phonological and the graphical form of the word. Of these factors, only the last one appears to benefit FLLs, giving them an advantage over HLLs in reconstructing the words that are difficult to spell. At the same time, HLLs appear to be better at decoding more difficult words in more difficult texts. However, when the C-test texts are analyzed as super-items, this advantage is no longer evident, whereas FLLs still have a clear advantage as C-test text 1 is biased towards them, possibly because of a large number of orthographically nontransparent words.
The subsequent error analysis corroborated the findings of the DIF analysis as it identified two main types of mistakes which reflect either failed reconstruction of meaning or failed reconstruction of form, that is, a lack of two of the aspects of knowledge and skills necessary to solve a C-test gap. The reason behind the variance in performance between the two groups on the biased items can be summarized as follows: HLLs tend to be able to recognize a wider range of lexical items within the context of the C-test texts, whereas FLLs simply do not possess this knowledge. At the same time, HLLs appear to be less successful in reconstructing the form of the gaps, which can be explained by the deficits in orthographical competence of HLLs pointed out in previous research (Böhmer, 2011; Mehlhorn, 2016). Furthermore, the greater variation in the performance of FLLs between the biased and non-biased items can be explained by the fact that the advantage of FLLs is clearly reflected in their scores: being able to produce the correct form of the recognized lexeme translates into an answer that is scored as correct. In contrast, HLLs do not benefit significantly from their ability to recognize more lexemes than FLLs, as recognition skills alone are not reflected in the scoring method employed in this study.
The different patterns of the answers given by FLLs and HLLs appear to support the findings of Hastings (2002) that C-tests are differentially sensitive to different aspects of language ability, one of which is spelling. To complete a C-test gap correctly, a learner needs to have the item in their vocabulary, to identify the item correctly based on the context, and to produce its correct grammatical and orthographical form, thus making use of different components of both BLC and HLC (Hulstijn, 2011). When a C-test is scored dichotomously, the test scores appear to reveal predominantly HLC components of language proficiency; namely, as can be seen from the results of the error analysis, HLLs are not given credit for their recognition skills (BLC) when they are not able to spell correctly (HLC). However, if partially correct answers were to be included in the scoring, the proportion of BLC revealed in the test scores would likely increase. Thus, the C-test construct is particularly sensitive to the scoring method used.
Different answer patterns of FLLs and HLLs also indicate that a FLL who received, for example, 20 points on the dichotomously scored C-test is likely to have recognized and spelled 20 words correctly, whereas a HLL with the same score was probably able to recognize more gaps in terms of their meaning but made mistakes while actually writing down some of the words, i.e. reconstructing their form. As suggested by reviewers, we conducted a post-hoc analysis of 19 items using the scoring method proposed by Baur and Spettmann (2007) and awarded one point for partially reconstructed items from response categories 2 and 3 (see the supplemental file). The results show that the performance of HLLs on these items improves considerably when the answers are scored using the word recognition value. At the same time, recalculating the scores for FLLs does not show such an improvement in performance. Thus, the C-test construct appears to be not only sensitive to the scoring method used, but also fluid, depending on the proficiency profiles of the learners (Sigott, 2004).
This quality of the C-test construct, which obviously requires further research, has consequences for making decisions about test takers who are heritage learners of the language. A placement decision based on dichotomously derived scores would clearly underestimate the language proficiency of HLLs (specifically, its BLC components). If HLLs are placed into language courses specifically tailored to their needs (i.e. courses primarily targeting the development of HLC), then such a placement decision can be justified. However, if placed into regular foreign language courses, these learners will outperform FLLs in terms of oral and aural skills and implicit language knowledge in general (Zyzik, 2016) and will not fully benefit from the instruction. To avoid a possible bias and to obtain a more comprehensive proficiency profile for HLLs, an additional test that targets oral and aural skills (e.g. an oral, group, or paired interview, or an Elicited Imitation test) should be used. Alternatively, as suggested above, a partial credit scoring for the C-test, which would account for word recognition, could be used. However, this rating method would first need to be validated and would obviously reduce the practicability of the test format, especially in online environments.
The present exploratory study with a limited body of data does not attempt to answer the question of the C-test construct either extensively or comprehensively. However, it does allow interesting insights into possible fields of future research. Further studies on the C-test construct could, for example, use a C-test with a larger number of gaps, while simultaneously analyzing a wider variety of linguistic phenomena to obtain a more robust and multifaceted data set. Apart from increasing the scope of data, it could be investigated whether the results would be similar when applied to other heritage language learning groups, such as Spanish, Turkish, or Polish.
Ideally, each test should be developed for use with a specific target audience and the question of scoring taken into consideration at the stage of test construction and validation (Kane, 2006). The present study raised important concerns about the interpretations of C-test scores when assessing different learner types with which the C-test has not been previously investigated. As the use of the C-test format is becoming more widespread and being applied in wider and more diverse contexts, the question of its suitability for a certain learner group and of an appropriate scoring method should be addressed in each specific case. Further research could also explore the various ways to score C-tests and the implications of different scoring methods for construct interpretation, test reliability, and test use.
Supplemental Material
Supplemental_file_comparison_of_the_scoring_methods_DRackertandTimukova – Supplemental material for What does the analysis of C-test gaps tell us about the construct of a C-test? A comparison of foreign and heritage language learners’ performance
Supplemental material, Supplemental_file_comparison_of_the_scoring_methods_DRackertandTimukova for What does the analysis of C-test gaps tell us about the construct of a C-test? A comparison of foreign and heritage language learners’ performance by Anastasia Drackert and Anna Timukova in Language Testing
Footnotes
Appendix
Distribution of biased and non-biased items across three texts.
| Text | Biased |
Biased |
No bias |
|---|---|---|---|
| 1 | яз |
универ истори э ста прест Рос препод прекр учи франц хор |
зо о рабо инте сту о дев что20 (č |
| 2 | выясн про лю стимул лю раб се замед |
исслед |
эскпер универ уча л ч инте деятел чте возра кот ча |
| 3 | сре эл осно образ пра реб бан |
опред |
Acknowledgements
We are grateful to the editors and reviewers of this journal for their valuable and insightful feedback on earlier versions of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
