Abstract
Many models of cognitive diagnosis, including the least squares distance model (LSDM), work under the conjunctive assumption that a correct item response occurs when all latent attributes required by the item are correctly performed. This article proposes a disjunctive version of the LSDM under which the correct item response occurs when at least one attribute is correctly applied. Also, under both the conjunctive and disjunctive versions of the LSDM, this article demonstrates an approach to estimating the conditional probability that (a) a specific pattern of p attributes, (b) exactly p attributes, and (c) at least p attributes will be correctly performed across locations on the logit scale in the item response theory under the one-, two-, or three-parameter logistic model. Such information can be useful for interpretations and decisions based on a person’s performance on attributes that govern the correct responses on binary items under unidimensional item response theory calibrations for assessment in education, psychology, and other fields.
Cognitive diagnosis deals with identification, validation, and analysis of latent attributes that underlie responses on items of assessment instruments in psychology, education, and other behavioral fields. Typically, the item responses are scored on a binary (1/0) scale. The attributes may refer to various latent characteristics such as cognitive processes, operations, and trait states. Cognitive diagnosis models (CDMs) integrate cognitive psychology and psychometrics (e.g., de la Torre, 2008, 2009; de la Torre & Douglas, 2004; DiBello, Stout, & Roussos, 1995; Dimitrov, 2007; Embretson, 1984, 1995; Embretson & Wetzel, 1987; Fischer, 1973; Gitomer & Rock, 1993; Henson & Douglas, 2005; Mislevy, 1996; Snow & Lohman, 1989; C. Tatsuoka, 2002; K. K. Tatsuoka, 1983, 2009; Templin & Henson, 2006). CDMs have been developed in a variety of frameworks such as item response theory (IRT), primarily Rasch modeling (e.g., Embretson, 1984, 1995; Fischer, 1973; Whitely & Schneider, 1981), combining IRT information with algebraic methods (e.g., Dimitrov, 2007; K. K. Tatsuoka, 1985), and nonparametric IRT, Bayesian modeling, and latent class modeling (e.g., de la Torre & Douglas, 2004; Henson & Douglas, 2005; Junker & Sijtsma, 2001; Mislevy, 1996; Templin & Henson, 2006; von Davier & Rost, 1997; von Davier & Yamamoto, 2004).
A key element in all models of cognitive diagnosis is the so-called Q-matrix. When a set of K attributes (e.g., cognitive operations and/or processes) is hypothesized to underlie the responses on J items, the Q-matrix is a J × K matrix with elements qjk = 1 if item j requires attribute k, and qjk = 0, otherwise (j = 1, . . . , J; k = 1, . . . , K). The identification of attributes that govern the responses on items of interest and correct specification of the respective Q-matrix are conditions of critical importance to the validity of results obtained under CDMs.
Most CDMs are conjunctive models as they are based on the assumption that a correct item response is produced when all attributes required by the item are mastered. Examples of conjunctive CDMs are the multicomponent latent trait model (Embretson, 1984), the LSDM (Dimitrov, 2007), the DINA (deterministic inputs; noisy “and” gate) model (de la Torre & Douglas, 2004; Haertel, 1989; Junker & Sijtsma, 2001; Macready & Dayton, 1977), and the NIDA (noisy inputs; deterministic “and” gate) model (Maris, 1999). For example, DINA is a stochastic conjunctive model in which the probability for individual i to correctly answer item j is conditioned on a conjunctive parameter, ξ ij , as follows:
where α ik = 1 if individual i has mastered attribute A k and α ik = 0 otherwise, and qjk is an element of the Q-matrix for item j and attribute A k . With this, the probability for correct item response is specified as follows:
where sj is the slip parameter (the probability that a master of all necessary attributes “slips” and answers the item incorrectly) and gj is the guess parameter (the probability that a nonmaster of at least one attribute answers the item correctly).
There are situations, however, in which the item responses are not governed by the restrictive assumption that all attributes must be correctly applied. Such situations may occur in assessments of cognitive, behavioral, psychological, and personality traits. In assessment of pathological gambling, for example, Templin and Henson (2006) proposed a disjunctive version of the DINA model, referred to as DINO (deterministic input; noisy “or” gate) model. Although the parameter ξ ij under DINA equals unity (ξ ij = 1) only when person i has mastered all attributes required by item j, under DINO this parameter is redefined as a parameter ω ij , which equals unity (ω ij = 1) when at least one of these attributes has been mastered (Templin & Henson, 2006). That is,
With this, the probability of correct item response under the DINO model is determined as
The Least Squares Distance Model
The LSDM (Dimitrov, 2007), which is of primary interest here, is a conjunctive CDM. However, unlike DINA (and other nonparametric models), the LSDM is based on information about item calibration under a unidimensional IRT model—the one-parameter (or Rasch), two-parameter, or three-parameter logistic models (1PL, 2PL, or 3PL). Unlike any other CDM, the LSDM does not require item score information, as long as IRT estimates of the item parameters are available. Specifically, using IRT estimates of the item parameters and the Q-matrix, the LSDM provides probability curves for individual attributes on the logit scale of IRT item calibration, as well as information about potential Q-misspecifications for individual test items. Under the LSDM, the probability of correct item response is presented as a product of the probabilities of correct processing of the attributes required by the item, that is,
where Xij is the binary (1/0) response of individual i on item j, θ i is the trait score (in logits) of individual i, A k is the kth attribute, and qjk is the element of the Q-matrix for item j and attribute k (qjk = 1 if item j requires attribute A k and qjk = 0, otherwise).
The LSDM is performed in steps as follows:
IRT estimates of the item parameters are obtained through IRT calibration or other sources (e.g., item banks or study reports).
A set of fixed θ values is selected to cover a targeted interval on the logit scale (say, from −4.0 to 4.0, with a step of 0.5).
The probability of correct item response, P(X = 1|θ), is computed for each θ value using the item parameter estimates under a specific IRT model of item calibration (1PL, 2PL, or 3PL).
Taking the natural logarithm on both sides of Equation (5) generates a norm of linear equations ||
The norm ||
The attribute probabilities across θ values are tabulated and graphed to obtain attribute probability curves (APCs).
The item characteristic curve (ICC) of each item is recovered through the product of the probabilities of the attributes required by the item to evaluate the Q-matrix for possible misspecifications.
LSDM results can be particularly useful in (a) validation screening of attributes required by IRT-calibrated items (e.g., taken from an existing IRT item bank) prior to test administration and (b) providing information about probability curves of attributes on the logit scale, thus keeping the analysis of items and their underlying attributes within the familiar and widely used logit scale in IRT. The analytic simplicity of the LSDM allows for a straightforward development of a disjunctive version of the LSDM, as an IRT-based analog to the disjunctive model DINO (Templin & Henson, 2006), and other extensions described in the following.
Purpose of the Study
The purpose of this article is to extend the conjunctive LSDM (Dimitrov, 2007) at two levels. First, to provide a disjunctive version of the LSDM under which the target response on an item (X = 1) may occur when at least one of the attributes associated with the item is correctly applied. The notation LSDM-C is used here to denote the conjunctive LSDM (Dimitrov, 2007), whereas LSDM-D denotes the disjunctive LSDM. The second goal is to upgrade the LSDM by estimating the conditional probabilities that (a) specific patterns of p attributes, (b) exactly p attributes, and (c) at least p attributes will be correctly performed by individuals at a given location, θ, on the IRT logit scale (p = 1, . . . , K). The proposed approach works under LSDM-C and LSDM-D, thus providing flexibility and richness of psychometric information that can be useful for diagnostic decisions about individuals and groups based on attribute-related criteria. All estimation procedures are executable by a computer program named LSDM-CD (Least Squares Distance Model of Cognitive Diagnosis). LSDM-CD is developed in the framework of MATLAB (MathWorks, Inc., 2010) but is also adapted to operate as a stand-alone application for both Windows and Linux operating systems. 1
Method
To avoid confusion in terminology, some clarifications on terms and notations used in this section are necessary. First, target response (X = 1) can be the correct item response in an achievement test, endorsement of a statement in a personality test, manifestation of an observed behavior, and so on. Also, depending on the context, an attribute is performed, applied, or processed correctly. Furthermore, attributes associated with an item (in the Q-matrix) can also mean that they are required by (or related to) the item.
Disjunctive LSDM
Let A1, A2, . . . , A K denote K attributes that are hypothesized to govern the target responses (X = 1) of individuals on a set of items that measure an underlying latent trait (or ability), θ. Under the conjunctive LSDM (LSDM-C; Dimitrov, 2007), the target response (X = 1) on an item is obtained if all attributes associated with the item in the Q-matrix are performed correctly. In this case, the probability of correct item response on item j by an individual i with a trait score θ i on the logit scale, P(X ij = 1|θ i ) is presented in Equation (5).
Under LSDM-D, the target response (X = 1) on an item is obtained if at least one of the attributes associated with the item in the Q-matrix is performed correctly. In this case, the probability of the target response (X = 1) to occur is estimated through the following transformations of Equation (5):
which, after simple algebra, yields the following analytic form of the LSDM-D:
Conditional Probabilities of Performing Attribute Patterns
Let
where ν(k) is the kth binary element in pattern
Conditional Probabilities of Performing Exactly p Attributes
Let
where
Conditional Probabilities of Performing at Least p Attributes
Let
where
To summarize, the conditional probabilities for correct performance of a specific pattern of p attributes, exactly p attributes, and at least p attributes are estimated through Equations (8), (9), and (10), respectively, where
Validity Issues
As both LSDM-C and LSDM-D use estimates of item parameters, the assumption is that these estimates are obtained under an adequate model–data fit of the respective IRT model——1P (or Rasch), 2P, or 3P model. Thus, testing for IRT model–data fit should be performed prior to using LSDM. When item parameter estimates are taken from, say, an IRT item bank or other sources, researchers should attend to evidence of IRT model–data fit provided with such sources. Assuming dependable estimates of the item parameters, the validity of LSDM-based results is concerned with the adequacy of the set of attributes and possible misspecification in the Q-matrix.
Regarding the conjunctive model (LSDM-C), Dimitrov (2007) outlined some empirical criteria related to validity of attributes and their association with individual test items. First, the APCs should exhibit an expected monotonic behavior, relative difficulty, and discrimination. For example, if LSDM-C analysis is conducted on a set of reading comprehension items, it is logical to expect that the APCs should increase with the increase of the underlying ability of reading comprehension. Second, possible misspecifications in the Q-matrix are signaled by the lack of acceptable LSDM-based recovery of the ICCs. Under LSDM-C, the ICC recovery for an item is obtained by multiplying the probabilities for correct processing of the attributed required by the item. The mean absolute difference (MAD) between the ICC of an item and its LSDM recovery represents an index of Q-matrix fit for this particular item (MAD = 0 indicates perfect ICC recovery). As suggested by Dimitrov (2007), MAD < .05 indicates an adequate ICC recovery, 05 ≤ MAD < .10 indicates a tolerable ICC recovery, and MAD ≥ .10 indicates a poor ICC recovery. Overall, this rule of thumb was supported by a series of simulated studies on LSDM (Dimitrov, Romero, & Ponsoda, 2008; Romero, 2010). More refined information in this regard, which is beyond the scope of this article, is provided by Romero (2010) under various simulation conditions for number of items, number of attributes, and relative difficulty of attributes.
The aforementioned rule of thumb for MAD values in checking for misspecifications in the Q-matrix applies also to the disjunctive model (LSDM-D). However, while the ICC recovery of an item under LSDM-C is based on Equation (5), the ICC recovery of an item under LSDM-D is based on Equation (7). It is important to emphasize also that the criterion of monotonic behavior of APCs, which is expected under LSDM-C, does not fully apply to LSDM-D. This is because under the LSDM-D, a positive item response (X = 1) is obtained if at least one of the attributes associated with the item in the Q-matrix is performed correctly. This implies, for example, that an increasing APC of a relatively easy attribute may decrease at high ability levels where patterns of attributes that are more likely to occur do not contain the easier attribute under consideration. Thus, given the element of randomness associated with the occurrence of “at least one” attribute in the manifestation of a target response (X = 1), the criterion for monotonic behavior of APCs is not strictly applicable under LSDM-D.
It is also important to note that the selection of LSDM-C versus LSDM-D is guided by the nature (conjunctive vs. disjunctive) of the attributes that are sufficient for obtaining the targeted response (X = 1) on individual items (i.e., all vs. at least one, respectively). In other words, the LSDM choice (LSDM-C or LSDM-D) is not a post hoc decision based on a criterion of “better data fit.” It seems fair to say that, when applicable, LSDM-C would be the proper choice in most education and psychological tests (e.g., ability or performance testing). The practical use of LSDM-D would be primarily in determining whether “at least” a given number of attributes are processed for classification decisions related to, say, level of depression, anxiety, and other psychological or behavioral traits (assuming unidimensionality of the trait). The LSDM-C and LSDM-D are illustrated next in Examples 1 and 2, respectively. The computer program LSDM-CD was used to render tabulated results, APCs, ICC recovery, and probability curves for specific pattern of p attributes, exactly p attributes, and at least p attributes (p = 1, . . . , K).
Example 1
This example illustrates an application of the conjunctive model (LSDM-C; Dimitrov, 2007) with the proposed extensions for estimating
Rasch Item Parameter Estimates and Q-Matrix for Algebra Equation Items and Five Attributes
Note.
Prior to conducting LSDM-C, it was found that the data fit the Rasch model under the conditional maximum-likelihood chi-square test for the entire set of items and the Wald z test for fit of individual items. Specifically, the chi-square value was not statistically significant, χ2(12) = 19.67, p > .05, and the Wald z values for individual items were within the acceptable range from −2.00 to 2.00. The item calibration was performed using the computer program for Rasch analysis LPCM-WIN (Fisher & Ponocny-Seliger, 1998).
Attribute Probability Curves
As described earlier, LSDM-C solutions of Equation (5) render estimates of probabilities for correct performance of attributes across θ levels on the logit scale (Dimitrov, 2007). For the data in Table 1, the APCs are shown in Figure 1. As expected under the LSDM-C, the APCs increase with the increase of the ability on the logit scale. Figure 1 also shows, for example, that A3 (removing denominators) is the most difficult attribute, whereas A1 (collecting terms) is, for the most part, the easiest attribute for the study population of examinees.

Attribute probability curves under the conjunctive least squares distance model (LSDM-C) analysis of algebra equations and five underlying attributes (A1 = collecting terms, A2 = removing parentheses, A3 = removing denominators, A4 = solving for a variable with nonnumerical coefficients, and A5 = processing nontrivial cases)
ICC Recovery
The ICC recovery of an item under LSDM-C is the product of the APCs for the attributes required by that item. In light of the criteria related to MAD between an ICC and its LSDM-C recovery, discussed earlier, the results for the 13 algebra items indicated an adequate recovery for nine items, with MAD < .05, and tolerable recovery for four items, with .05 < MAD < .09. For illustration, the ICC recovery of one item (Item 3, with MAD = .0256) is depicted in Figure 2. Overall, the expected monotonic behavior of the APCs and ICC recovery of items provide satisfactory evidence of LSDM-C validity for the purpose of this example.

Recovery of the item characteristic curve for Item 3 under the conjunctive least squares distance model analysis of five attributes for the algebra test items
Probability for Correct Performance of Attribute Patterns
The estimation of conditional probabilities for correct performance of specific attribute patterns is illustrated here for all possible patterns of three and four (out of five) attributes at two locations on the logit scale (θ = −1 and θ = 1). Table 2 shows LSDM-C estimates of conditional probabilities for correct performance of each attribute (A1, . . . , A5) at θ = −1 and θ = 1. The estimates of conditional probabilities of correct performance of all possible patterns of three and four attributes are provided in Table 3. They are obtained using Equation (8) with the probability estimates for individual attributes provided in Table 2. For example, let us take the first pattern of three attributes in Table 3,
Conditional Probabilities for Correct Performance of Attributes at θ = −1.0 and θ = 1.0 for the Algebra Test Items
Conditional Probabilities for Correct Performance of All Possible Patterns of Three and Four (out of Five) Attributes at θ = −1.0 and θ = 1.0 for the Algebra Test
The LSDM-C estimates of probabilities for correct performance of the patterns of four attributes in Table 3 are depicted in Figure 3. As can be seen, pattern (1 1 0 1 1), which does not contain the most difficult attribute (A3), is much easier than all other patterns across all ability levels on the logit scale, whereas pattern (0 1 1 1 1) almost never occurs.

Probability curves for patterns of four (out of five) attributes under the conjunctive least squares distance model analysis of five attributes for the algebra test items (the attributes, from left to right, in each pattern are A1, A2, A3, A4, and A5)
Probability of Performing Correctly Exactly p Attributes
Under LSDM-C, the probability of performing correctly exactly p attributes (i.e., any pattern of p attributes) is the sum of the conditional probabilities of correct performance of all possible patterns of p attributes (p = 1, . . . , K). For example, using the results in Table 3, the probability of performing correctly exactly four (out of five) attributes for persons located at θ = −1 on the logit scale is

Probability curves for correct performance of exactly p attributes (p = 0, 1, . . . , 5) under the conjunctive least squares distance model analysis of five attributes for the algebra test items
Probability of Performing Correctly At Least p Attributes
Under LSDM-C, the estimation of the probability of performing at least p (out of K) attributes,

Probability curves for correct performance of at least p attributes (p = 1, . . . , 5) under the conjunctive least squares distance model analysis of five attributes for the algebra test items
Example 2
This example illustrates the disjunctive model (LSDM-D) for simulated data. The item parameters (discrimination, a, and difficulty, b) of 25 items were generated, with a randomly selected from a uniform distribution, U(0, 2), and b randomly selected from the standard normal distribution, N(0, 1). Probabilities of correct item responses were generated under the IRT two-parameter logistic model, P(X = 1|θ, a, b) for values of θ from −5.0 to 5.0, with a “step” of 0.5. A Q-matrix for six attributes (A1, . . . , A6) associated with the 25 items was also generated to reflect the disjunctive rule of correct item performance under the LSDM-D. The rationale behind using simulations in this example was to produce data based on the disjunctive rule of correct item performance for an appropriate illustration of the LSDM-D. The resulting item parameters and Q-matrix are provided in Table 4.
Simulated Item Parameters and Q-Matrix for LSDM-D Analysis
Note. LSDM-D = disjunctive least squares distance model.
The logic of LSDM-C estimation of probabilities for specific patterns of p attributes, exactly p attributes, and at least p attributes (p = 1, . . . , K), illustrated in Example 1, carries over the estimation under LSDM-D. Keep in mind, however, that although the conditional probabilities of correct attribute performance, P(A k |θ) are estimated via Equation (5) under LSDM-C, they are estimated via Equation (7) under LSDM-D. The results in this example were obtained using the LSDM-D option in the computer program LCDM-CD. As expected for the simulated data in this example, the MAD values for LSDM-D recovery of the ICCs for all 25 items were below the cutting value of .05 for adequate recovery (the MAD values ranged from .0014 to .0421).
For space consideration, shown here are only probability curves for correct performance of exactly p attributes (p = 0, 1, . . . , 6; Figure 6) and correct performance of at least p attributes (p = 1, . . . , 6; Figure 7). As noted earlier, the information depicted in Figure 7 would probably be of main interest for classification of individuals into trait-related categories (e.g., depression or some pathological behavior). For example, using the disjunctive DINO model, Templin and Henson (2006) classified persons as probable pathological gamblers when the probability that such persons meet at least 5 (out of 10) criteria was .5 or higher. In Figure 7, the “cutting scores” on the ability scale for correct performance of at least p attributes, with a probability of .5 or higher, were found to be the following: (a) θ = −2.228 for p = 1, (b) θ = −1.281 for p = 2, (c) θ = −0.575 for p = 3, (d) θ = 0.301 for p = 4, (e) θ = 1.215 for p = 5, and (f) θ = 2.448 for p = 6.

Probability curves for correct performance of exactly p attributes (p = 0, 1, . . . , 6) under disjunctive least squares distance model analysis of simulated data on 25 binary items and six attributes

Probability curves for correct performance of at least p attributes (p = 1, . . . , 6) under disjunctive least squares distance model analysis of simulated data on 25 binary items and six attributes
Discussion
The LSDM of cognitive diagnosis (Dimitrov, 2007) is based on a conjunctive model in which a correct item response assumes correct performance of all attributes required by the item (see Equation 5). Unlike other CDMs, the LSDM does not require item score information as long as IRT estimates of the item parameters are available (e.g., from an item bank or published sources). LSDM renders estimates of probabilities of correct attribute performance and validity information about recovery of ICCs from hypothesized attributes across trait scores on the IRT logit scale.
This article provides a straightforward extension of the conjunctive model (LSDM-C) to a disjunctive version (LSDM-D) under which a target item response (X = 1) is obtained when at least one of the attributes associated with the item is correctly performed. This is achieved by applying the LSDM algorithm using Equation (7) for the relationships between the probability of a correct item response and the probabilities of correct performance of attributes associated with the item.
Further extensions of both LSDM-C and LSDM-D render estimates of probabilities that (a) specific patterns of attributes, (b) an exact number of attributes, and (c) at least a specified number of attributes will be correctly performed by individuals at a given location on the IRT logit scale. The LSDM-C and LSDM-D allow researchers to conduct analysis and interpretations of both test items and attributes in a unified framework of IRT calibration for unidimensional binary items. The relative simplicity of the LSDM-based estimation procedures allows for their straightforward coding in the syntax of computer programs for mathematics or statistical data analysis such as MATLAB (MathWorks, Inc., 2010) and R (R Development Core Team, 2009). In this article, all results (probability estimates, ICC recovery, etc.) were obtained using the computer program LSDM-CD.
In conclusion, given the variety of contexts and measurement scenarios for cognitive diagnosis, the choice of a specific model (e.g., LSDM-C or DINA; LSDM-D or DINO) in any particular case would depend on the purpose of the study, nature of the data, and so on. It is important to emphasize that, unlike other models of cognitive diagnosis, LSDM-based analyses render psychometric information about binary items and their attributes on the same scale—namely, the familiar logit scale in IRT. The LSDM (Dimitrov, 2007) and its psychometric extensions proposed in this article can be particularly useful in analyses and interpretations of binary items and underlying (cognitive, behavioral, or psychological) attributes within a unified framework of unidimensional IRT calibrations in psychology, education, and other disciplines. Possibilities for multidimensional extensions of the LSDM-C and LSDM-D can be examined in future research.
Footnotes
Notes
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
