Abstract
C-tests are gap-filling tests mainly used as rough and economical measures of second-language proficiency for placement and research purposes. A C-test usually consists of several short independent passages where the second half of every other word is deleted. Owing to their interdependent structure, C-test items violate the local independence assumption of IRT models. This poses some problems for IRT analysis of C-tests. A few strategies and psychometric models have been suggested and employed in the literature to circumvent the problem. In this research, a new psychometric model, namely, the loglinear Rasch model, is used for C-tests and the results are compared with the dichotomous Rasch model where local item dependence is ignored. Findings showed that the loglinear Rasch model fits significantly better than the dichotomous Rasch model. Examination of the locally dependent items did not reveal anything as regards their contents. However, it did reveal that 50% of the dependent items were adjacent items. Implications of the study for modeling local dependence in C-tests using different approaches are discussed.
Introduction
C-tests are economical gap-filling tests designed to measure overall language proficiency in the first language (L1) and the second language (L2; Raatz & Klein-Braley, 2002). A standard C-test battery is composed of 4–8 short independent passages where the second half of every other word is deleted. Each passage contains 20–25 broken words, resulting in at least 80 gaps or items. Examinees are required to fill in the missing letters and are given 1 point for each correctly reconstructed word. The advantages of C-test over the cloze test is that in C-tests (1) only exact word scoring is possible, which results in higher reliability coefficients, (2) the rate and point of onset of deletions is fixed, which leads to a more stable construct across different C-tests, (3) several independent passages are used to avoid bias due to text familiarity, (4) the higher number of gaps in C-tests leads to a better representation of the elements of language, and (5) L1 speakers of the language score more than 90% on C-tests (Raatz & Klein-Braley, 2002).
C-tests are based on the reduced redundancy principle (RRP; Raatz & Klein-Braley 2002). The RRP postulates that natural languages contain redundant elements, and that L1 or competent speakers of a language can function in the language satisfactorily when redundancy is reduced. Therefore, distorted messages, like speech in a noisy party, can be understood by competent users of a language. The RRP has been employed as a theoretical basis to develop new tests or to account for some existing tests like the cloze test. In RRP tests, noise is introduced into linguistic messages, and examinees are required to process the language.
The strong theoretical basis of the C-test combined with substantial empirical evidence in support of its validity and reliability has made the C-test an appealing testing technique in L2 measurement. C-tests are widely used for research, placement, screening, and other purposes in second and foreign language contexts. The latest C-test bibliography (Grotjahn & Drackert, 2020) records 469 studies on different aspects of the C-test and 313 additional studies in which C-tests are used as measurement instruments in L2 acquisition research.
Item response theory analysis of C-tests
Item response theory (IRT) analyses of C-tests have shown that they fit different types of IRT models, especially the Rasch models (RMs; Forthmann et al., 2019). A challenge for analyzing C-tests with IRT models is their interdependent structure. Conditional independence (CI) or local item independence is a basic assumption of all IRT and RMs. Violation of CI results in biased parameter estimates and artificially high reliability estimates (Kreiner, 2007). CI stipulates that items should be conditionally independent on the latent variable. That is, when the impact of the latent variable is factored out, the items should be independent. The assumption of CI—along with undimensionality and monotonicity—is fundamental in IRT models. If CI is violated, that is, if the items remain correlated after factoring out the latent variable, it means that another unintended latent variable exists that contaminates the test scores and threatens validity. Therefore, violation of the CI and multidimensionality are two related issues and, thus, approaches for addressing local item dependence (LID) always try to model LID in terms of multidimensionality (Kreiner & Christensen, 2004).
To solve the problem of LID in C-tests, researchers have frequently employed the super-item approach (Forthmann et al., 2019). In this modeling strategy, the dichotomous items that belong to each passage are aggregated as passage scores or super-items (testlets), and the raw scores on the passages are entered into the RM analysis (Forthmann et al., 2019). That is, each passage is treated as a polytomous item with 20–25 categories, and an RM for polytomous items is used. The drawback of these models for scaling C-tests is that the information contained in the individual items is lost. By aggregating the items within passages into super-items, the number of items is reduced to the number of passages, and one gets no information about the gaps.
Another less commonly used approach to model LID in C-tests is the testlet response theory (TRT; Bradlow et al., 1999). TRT is a multidimensional bifactor IRT model where all the items (gaps) are forced to load on a general ability dimension, while they load on specific testlet dimensions too. TRT is a complex model and requires large samples (Eckes & Baghaei, 2015). An alternative strategy for modeling LID in C-tests is the loglinear Rasch model (LLRM) of Kelderman (1984). The purpose of this research is to demonstrate how the LLRM works for a C-test. This is a novel psychometric modeling strategy that has not been used for C-tests before.
Loglinear Rasch Model
The RM is built on four assumptions: (1) the existence of a unidimensional latent variable, (2) a sufficient raw score, (3) locally independent items, and (4) absence of differential item functioning (DIF; Kreiner & Christensen, 2004). Kreiner and Christensen (2004) wrote that Assumptions 1 and 2 are fundamental and should be met for a scale score to make sense, but Assumptions 3 and 4 “are technical and relate to the selection or creation of items” (p. 193). They argued that for a set of items, if Assumptions 1 and 2 hold, but 3 and 4 do not, the items and the scoring are sound, and one just needs to model the LID or remove biased items instead of discarding all the items and rejecting the hypothesis that a latent variable exists. Discarding items results in shorter, less reliable, and less valid tests. If the reason for misfitting items is LID or DIF, the problem can be solved by modeling instead of discarding items.
Kelderman (1984) introduced the LLRM in which the assumptions of CI and absence of DIF are relaxed by introducing a loglinear structure among the items and exogenous variables. In LLRM, interaction terms are added among locally dependent items to account for LID or between items and exogenous variables (like socioeconomic and educational background) to account for DIF. In the case of LID, this is equivalent to allowing the residuals of locally dependent items to correlate. LLRM incorporates uniform LID and uniform DIF while retaining the fundamental assumptions of the RM, namely, the existence of a latent variable and a sufficient raw score.
RMs satisfy specific objectivity, which means that comparing examinees becomes independent of the particular set of items used for measurement (Rasch, 1960/1980). When LID is accounted for by adding an interaction parameter, the total score is still a sufficient statistic for estimating person parameters, and person parameters can be separated from item parameters, and thus objectivity holds, but the limitation is that it is not possible to obtain item-free person measures by selecting random subsets of items. Kreiner (2007) referred to this as essential objectivity. Nevertheless, the items still measure the latent trait and there is no need to discard them, which is referred to as essential validity by Kreiner (2007).
Method
Test material and participants
Test material was composed of two C-test passages and 30 multiple-choice reading comprehension items. Each C-test passage contained 25 gaps, and each correct answer was scored 1, which led to a score of 0–25 for each passage and a sum score of 0–50 for the whole C-test. Only the responses to the 50 C-test items were analyzed in this study. Fifty minutes were allotted for the entire test without specifying separate time limits for the two sections. The test was administered as an achievement test at the end of an advanced reading comprehension course. Score data are available to the public for inspection in the IRIS database (Baghaei & Christensen, 2023). Participants (N = 320, 218 females) were 4th-semester university students studying English as a foreign language at Islamic Azad University in Mashhad, Iran. Their first, home language was Persian, and the age range was 19–33 (M = 20.87, SD = 3.42).
Data analysis
In the first step, the 50 individual gaps were analyzed with the dichotomous RM (Rasch, 1960/1980). In the next step, the locally dependent item pairs were allowed to correlate using the LLRM (Kelderman, 1984). Details about the LLRM are included in this article’s supplementary material. Locally dependent items were identified using Yen’s Q3 statistic (Yen, 1984), which is based on the correlation between item residuals. According to Christensen et al. (2017), the Q3 statistic is the most widely used statistic for detecting LID. An absolute Q3 value of .20 and above within each testlet was considered as evidence of LID (Christensen et al., 2017). Eight pairs of items were flagged as locally dependent. Interaction terms were added to the model to allow the dependent item pairs to correlate. The two models were compared in terms of overall fit, item fit, precision, and parameter estimates. All the analyses were run using the DIGRAM software program (https://biostat.ku.dk/DIGRAM/; Kreiner, 2003).
The two models, that is, the RM and the LLRM, were compared with their −2loglikelihoods (deviance) and information criteria. The difference between deviances of two nested models is chi-square distributed with the difference between the numbers of estimated parameters degrees of freedom. As Table 1 shows, the LLRM fitted significantly better than the RM, χ2 = 411.74, df = 7, p < .001. Akaike information criterion (AIC) and Bayesian information criterion (BIC) for the RM were 11,920 and 12,105, respectively, while for the LLRM, AIC and BIC were 11,583 and 11,794, respectively, indicating the better fit of the LLRM. Outfit mean square statistics showed that more items fit the LLRM. There were six items with outfit mean square values above 1.50 (Linacre, 2022) in the RM, and there were three in the LLRM.
Model fit and parameter statistics for the C-test items in the RM and LLRM.
Note: AIC: Akaike information criterion; BIC: Bayesian information criterion; RM: Rasch model; LLRM: loglinear Rasch model; θ: person parameter; δ: item parameter; M: mean; SD: standard deviation.
Person parameters from the two models were compared. The latent scale is anchored by fixing the average item location at zero. Each of the interaction parameters γ describe the additional correlation between a pair of items, and in the LLRM, the scale is anchored in exactly the same way. The restriction (Σibi = 0) is the same in the two models that are compared. So the two sets of parameters that are compared are on the same scale. The dichotomous RM yielded a wider ability range and a slightly higher reliability, which could be due to the dependency among the items. As noted earlier, local dependence among items results in spuriously high reliability coefficients (Zenisky et al., 2002). Although the person parameters from the two models had a correlation of .998, the mean of the absolute differences between the parameters from the two models was .40 (range = 0 to .49). For the majority of the examinees, person ability was overestimated in the dichotomous RM by around .40 logits, which can have severe ramifications when standard setting and pass/fail decisions are made.
Gamma coefficients (see Equation 4 in the supplementary material) which indicate the strength of local dependence were also evaluated. According to Nielsen and Kreiner (2013), gamma values smaller than .10 are low, between .10 and .20 are moderate, and greater than .20 are strong. As Table 2 shows, all the gamma coefficients are extremely strong. The association between items 11 and 28 is negative and a lot smaller in size compared with other item pairs. The fact that these two items are in two independent C-test passages may explain the reason. Kreiner and Christensen (2004) showed that in multidimensional item bundles, where each testlet or item bundle measures a different construct, gamma coefficients are positive for items nested in the same testlet and negative for items belonging to different testlets.
Gamma coefficients representing the strength of local dependence.
Conclusion
Analyzing C-tests with IRT models is problematic due to the interdependent structure of C-tests. The conventional strategy to circumvent this problem is to treat each passage as a super-item or testlet and analyze the test with a polytomous RM. The drawback of the super-item approach is that the information contained in the individual items is lost, and this leads to less accurate item and person parameter estimates (Eckes & Baghaei, 2015). Additionally, the polytomous models yield many step parameters that do not have a clear interpretation for C-tests.
In this research, an extension of the RM, namely, the LLRM, which accommodates local dependence by adding interaction terms between dependent items, is employed, and the findings are compared with the dichotomous RM, where LID is ignored. Findings showed that LLRM fits significantly better than the dichotomous RM. Further analyses also showed that when LID is ignored, person ability parameters and test precision are overestimated.
Examination of the items that were identified as locally dependent did not reveal anything in regard to their content. However, four pairs of the dependent items (out of a total of eight pairs) were adjacent items. Therefore, one can conclude that adjacency accounts for 50% of local dependence in the C-test analyzed in the current study.
One reviewer argued that a limitation of modeling LID with LLRM is that LID should first be identified and then modeled in a second step, which adds some arbitrariness to the model because, as in different samples, different items might be identified as being dependent. While this argument is valid, it should be noted that LID could be merely an effect of the specific sample being tested. Thus, researchers should examine the items’ content to make sure that there are strong substantive reasons for modeling LID.
The advantage of the LLRM is that LID is modeled when it is necessary. In other approaches for modeling LID, such as TRT or the super-item approach, LID is assumed to exist a priori. That is, it is assumed that the items that share a common prompt or are nested within a passage must be locally dependent, and the assumed LID is modeled accordingly. This approach is problematic because modeling negligible testlet effects makes the model unnecessarily complicated, risks capitalization on chance, and increases errors in parameter estimates (DeMars, 2012). Thus, De Mars recommended that when TRT is used for modeling LID, testlets should first be examined to ascertain that they share something beyond the target latent variable before aimlessly modeling LID.
A shortcoming of the LLRM is that there is some limitation on the number of dependencies that can be modeled. Adding an interaction parameter for each item pair increases the number of parameters dramatically as the number of items increases (for 10 items there are 45 pairs; for 20 items 190 pairs; and for 50 items 1225 pairs). However, for models with few pairs of locally dependent items, the estimated interaction parameters can easily be interpreted.
This study is the first attempt to analyze C-tests with the LLRM. The advantage of LLRM over the super-item approach is that all the information in the individual items is used to estimate person and item parameters. Future studies should compare the LLRM with the super-item approach and the multidimensional IRT models for addressing LID.
Supplemental Material
sj-docx-1-ltj-10.1177_02655322231155109 – Supplemental material for Modeling local item dependence in C-tests with the loglinear Rasch model
Supplemental material, sj-docx-1-ltj-10.1177_02655322231155109 for Modeling local item dependence in C-tests with the loglinear Rasch model by Purya Baghaei and Karl Bang Christensen in Language Testing
Footnotes
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Alexander von Humboldt Foundation in Germany is acknowledged for providing a grant to the first author for conducting this research.
Open practice
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
