Abstract
This study presents new models for item response functions (IRFs) in the framework of the D-scoring method (DSM) that is gaining attention in the field of educational and psychological measurement and largescale assessments. In a previous work on DSM, the IRFs of binary items were estimated using a logistic regression model (LRM). However, the LRM underestimates the item true scores at the top end of the D-scale (ranging from 0 to 1), especially for relatively difficult items. This entails underestimation of true D-scores, inaccuracy in the estimates of their standard errors, and other psychometric issues. The inverse-regression adjustments used to fix this problem are too complicated for regular applications of the DSM and not in line with its simplicity. This issue is resolved with the IRF models proposed in this study, referred to as rational function models (RFMs) with one parameter (RFM1), two parameters (RFM2), and three parameters (RFM3). The proposed RFMs are discussed and illustrated with simulated and real data.
A recently developed method of scoring and analyzing test data, referred to as D-scoring method (DSM; Dimitrov, 2016, 2017), is adopted in the assessment practice at the National Center for Assessment in Saudi Arabia and gaining attention in the field of educational and psychological measurement. The transparency, simplicity, and efficiency of the DSM have been demonstrated in numerous studies (e.g., Al-Mashary & Dimitrov, 2018; Dimitrov, 2016, 2017; Domingue & Dimitrov, 2015; Dimitrov & Luo, 2017a, 2017b, 2018; Han, Dimitrov, & Al-Mashary, 2018) and computerized applications such as the System for Automated Test Scoring and Equating Under the D-scoring Method (SATSE-D, Atanasov & Dimitrov, 2019a), System for Automated Test Assembly (SATA-D; Atanasov & Dimitrov, 2019b), and a self-sustained computer program for D-scoring, equating and item analysis, DELTA (Atanasov & Dimitrov, 2018).
The DSM has the transparency and simplicity of the classical test theory and it brings some key measurement features that are not available in previously used classical test theory approaches to test scoring and analysis. For example, (a) the test score of an examinee is based on his/her response vector and takes into account the expected difficulty of the test items for the population, (b) the examinees’ scores and item difficulties are on the same scale, and (c) an analytic item response function (IRF) model is used for the estimations of true scores on individual items and the entire test, conditional standard error of measurement (CSEM), and other useful measures. In addition, as shown by Dimitrov and Luo (2018), the DSM for binary items can be readily applied with ordered polytomous items and the results parallel their counterparts obtained under the graded response model (GRM) in item response theory (IRT; Samejima, 1969, 1996). The measurement and practical efficiency of the DSM was also shown in the framework of multistage testing (Al-Mashary & Dimitrov, 2018; Han et al., 2018).
A key role in the DSM plays the analytic modeling of item response functions (IRFs) on the D-scale. In the initial work on such modeling, the DSM-based IRFs were estimated using a logistic regression model (LRM; Dimitrov, 2017) as an analog to the IRF logistic model in IRT (e.g., Hambleton, Swaminathan, & Rogers,1991). However, as explained hereafter, the LRM tends to underestimate the item true scores at the top end of the D-scale (ranging from 0 to 1), especially for relatively difficult items. This problem naturally carries over the estimation of trues scores on the entire test, CSEM, parametric indices of item and person fit, and so forth. Therefore, the purpose of the present study is to offer new IRF models that avoid LRM-related estimation problems on the D-scale. Prior to presenting these new models, a brief description of the DSM and the previously used LRM for IRF is provided next.
D-Scoring Method
Under the DSM of a test with binary items, the D-score of a person is based on his/her response vector of item scores (1 = correct, 0 = incorrect) weighted by the expected difficulties of the respective items for a target population of examinees (Dimitrov, 2017). To clarify, if
where
Logistic Regression Model for IRFs on the D-Scale
After computing the D-scores via Equation (1), the probability of correct response on an item by person s, given the Ds score of that person on the D-scale (from 0 to 1), was previously estimated as a predicted item score,
where Ds is the independent variable (predictor), obtained a priori via Equation (1), whereas
Problem With the LRM
The IRF model in Equation (2) works well in estimating item true scores,

Item characteristic curves (ICCs) of three items under the two-parameter logistic regression model (LRM), (a) an easy item (Item 1: a = 2.5, b = 0.15), (b) a moderate difficulty item (Item 2: a = 2.70, b = 0.35), and (c) a difficult item (Item 3: a = 6.00, b = 0;80) on the D-scale.
In some practical applications of the LRM on the D-scale, the underestimation of item true scores has been avoided by using inverse-regression adjustments (e.g., Han et al., 2018), but this procedure is too complicated for regular applications of the DSM. To solve this problem, the present study offers new IRF models, referred to as rational function models (RFMs), with one parameter (RFM1), two parameters (RFM2), and three parameters (RFM3). Presented first is the RMF2 as it plays a central role in the IRF modeling on the D-scale and serves as a base for the RFM1 and RFM3. In measurement terminology, the graphical representation of the IRF for an item is referred to as item characteristic curve (ICC).
Rational Function Model With Two Parameters for IRFs on the D-Scale
The RFM2 is designed to fit data on items with different levels of discrimination; that is, to allow for crossing of ICCs on the D-scale (Figure 2). The proposed RFM2 has the following analytic form
where P is the probability of correct item response for a person with a score D (from 0 to 1), b is the item location—that is, the location on the D-scale where the probability of correct item response is 0.5 (50% chance of success; when D = b, Equation 3 provides P = 0.5), and s is a fit parameter for shape. It should be noted that the fit parameter s governs the shape of the IRF in fitting the observed item responses over the D-scale but it is not the slope of the IRF at the item location, b; that is, s is not the item discrimination parameter. As shown next, the item discrimination, a, is obtained as a function of the parameters b and s. Given the values of b and s in Equation (3), the IRF increases monotonically from 0 to 1 with the increase of the D-score from 0 to 1. Thus, the RFM2 meets the key measurement assumption of monotonicity on the D-scale. Also, P = 1 when D = 1 and, mathematically, P tends to 0 when D tends to 0. Thus, it is appropriate to set P = 0 when D = 0 in computing the IRF value at D = 0 under the RFM2 in Equation (3).

Item response functions (IRFs) of five hypothetical items with different pairs of fit parameters (b, s) under the rational function model with two parameters (RFM2).
Item Discrimination Under the RFM2
Just as it is in IRT, the item discrimination, a, is the value of the first derivative of the IRF (here, in Equation 3) at the location b on the D-scale; that is,
Thus, the item discrimination, a, is computed via Equation (4) after obtaining estimates of the parameters b and s on fitting the RFM2 in Equation (3) to the observed item scores (Tables 1 and 2).
Item Fit Parameters for Five Hypothetical Items.
Note. s = model fit parameter for shape, b = item location, a = item discrimination (obtained via Equation 4).
Rational Function Model With One Parameter for IRFs on the D-Scale
The RFM1 is obtained from the RFM2 by fixing the fit parameter for shape, s, to a prespecified value in Equation (3). In an analogy with the 1PL in IRT, the IRFs produced by the RFM1 do not cross. One can fix the fit parameter s to a desirable value keeping in mind that higher s values produce steeper noncrossing IRFs, that is, higher level of discrimination at the respective item locations on the D-scale. For illustration, given in Table 2 are the parameters of 15 items grouped in triplets, with the items in each triplet having the same location, b, but different fixed values of the fit parameter for shape (s = 1, 2, and 4). The ICCs of these items are shown in Figures 3, 4, and 5, for the groups of items with s = 1, 2, and 4, respectively.
Item Parameters of 15 Items Grouped in Triplets With the Same Location Parameter and Three Different Fixed Values of the Fit Parameter for Shape (s = 1, 2, 4) Under the RFM1.
Note. RFM1 = rational function model with one parameter. The item discrimination, a, is computed from the values of item location, b, and fit parameter, s, via Equation (4).

Item response functions (IRFs) of five items obtained under the basic rational function model with one parameter (RFM1; s = 1) with different item locations.

Item response functions (IRFs) of five items obtained under a rational function model with one parameter (RFM1) with the fit parameter for shape fixed at 2.0 (s = 2) and different item location parameters, b.

Item response functions (IRFs) of five items obtained under a rational function model with one parameter (RFM1) with the fit parameter for shape fixed at 4.0 (s = 4) and different item location parameters, b
By fixing s = 1 in Equation (3), we obtain the analytic form for the “basic” RFM1:
After estimating the item location, b, under the RFM1 in Equation (5), the item discrimination, a, is computed by setting s = 1 in Equation (4). Some psychometric properties under the RFM1 and RFM2 are described in the appendix.
As noted earlier, the ICCs under the RFM1 do not cross, as it is with ICCs under the 1PL (Rasch) model in IRT. However, under the latter, the item discrimination is constant (a = 1) for all test items, whereas it varies across item locations, b, under the RFM1 because it depends on the values of s and b (see Equation 4). However, the RFM1 discrimination parameters of two items are equal if their locations are symmetrical around the mean of the D-scale (0.5). Indeed, as the product b(1 −b) has the same value for such two items, their discrimination parameter, a, is the same according to Equation (4). In Table 2, for example, Items 1 and 13 (with s = 1) have the same discrimination (a = 1.562) because their b-values (0.80 and 0.20, respectively) are symmetrical around the mean of the D-scale. As another example, Items 6 and 12 (with s = 4) have the same discrimination (a = 4.502) as their b-values (0.667 and 0.333, respectively) are also symmetrical around the mean of the D-scale.
Basic RFM1
The notation RFM1 will be used hereafter for the basic RFM1 (s = 1) in Equation (5). If another fixed value for s is used (s≠1), this will be explicitly stated. Conceptually, the reason for using s = 1 for the basic RFM1 is the same as that for fixing the discrimination parameter (a = 1) in the 1PL (Rasch) model in IRT. Specifically, fixing a = 1 in the Rasch model is seen as appropriate because a value greater than 1 means that the item discriminates between high and low performance more than expected for an item of this difficulty, whereas a value less than 1 means that the item discriminates between high and low performers less than expected for an item of this difficulty. (Kelley, Ebel, & Linacre, 2002)
As Kelley et al. (2002) also stated, Rasch analysis requires items which provide indication of relative performance along the latent variable. It is this information which is used to construct measures. From a Rasch perspective, over-discriminating items are tending to act like switches, not measuring devices. Under-discriminating items are tending neither to stratify nor to measure. (p. 883)
It should be also noted that under the Rasch model, where the theoretical item discrimination parameter, a, is fixed at 1, the empirical item discriminations are not exactly equal and they are reported “posthoc” (as a type of fit statistic) in Rasch computer programs such as Winsteps (e.g., Smith & Wind, 2018). The amount of the departure of a discrimination from 1.0 an indication of the degree to which that item misfits the Rasch model.
An illustration of over-discrimination under the RFM1 is provided with the IRFs shown in Figures 4 and 5, for the RFM1 with s = 2 and 4, respectively. Test developers may decide to select items that fit a RFM1 with a fit parameter for shape, s, greater than 1 if the goal is to stratify people for selection or other purposes. However, under-discriminating items (when s < 1) cannot be useful in a proper measurement scenario. As illustrated in Example 1 in the following, the basic RFM1 (s = 1) works well for data simulated under the 1PL (Rasch) model, whereas RFM1 with s = 2 does not.
Rational Function Model With Three Parameters for IRFs on the D-Scale
In many test scenarios, especially with multiple-choice items, low-ability examinees tend to perform higher than expected on some items. In IRT this issue is addressed by using the 3PL model of IRF for binary items (Birnbaum, 1968) where, in addition to the item parameters of location, b, and discrimination, a, a third parameter, c, is used to account for guessing or other causes of higher-than-expected performance of low-ability examinees (e.g., see Hambleton et al., 1991; Han, 2012). Likewise, the RFM2 in Equation (3) is extended here to a rational function model with three parameters (RFM3) as follows
where c is the pseudo-guessing parameter and b and s are item parameters as with the RFM2.
After estimating the parameters b, s, and c, on fitting the RFM3 in Equation (6) to the observed scores on the item under consideration, the RFM3 discrimination parameter, a, is obtained as follows (compare with Equation 4 under the RFM2):
It should be noted that, just as it is with the 3PL model in IRT, the item location, b, under the RFM3 is the location on the D-scale where the probability of correct item response is computed as
An illustration of IRF for an item under the RFM3 as well as the IRF of the same item obtained under the RFM2, is provided in Figure 6. This item was selected from a set of 72 binary items representing the verbal part of the General Aptitude Test administered to 27,078 examinees in Saudi Arabia. Given are the expected difficulty of this item (δ = 0.445) and the estimates of its parameters under the RFM2 (a = 1.746, b = 0.405) and RFM3 (a = 1.793, b = 0.453, c = 0.141).

The item response function (IRF) fit to observed responses on Item 4 (δ = 0.445) under two models, RFM2 (a = 1.746, b =0.405) in the left panel, and RFM3 (a = 1.793, b = 0.453, c = 0.141) in the right panel, based on the scores of 27,078 examinees on a verbal test that consists of 72 binary items from which the item for this illustration (Item 4) was selected. RFM2 = rational function model with two parameters; RFM3 = rational function model with three parameters.
Checking for Item Fit Under RFMs
The checking for item fit under RFMs is conducted here by using an index for item fit, referred to as mean absolute difference (MAD). This index represents the mean of absolute differences between the examinees’ observed scores (1/0) on the item and their expected values obtained under the RFM model of choice (RMF1, RFM2, or RFM3). That is,
where
Example 1
This example illustrates the RFM1 with two fixed values of the fit parameter for shape, s = 1 (basic RFM1) and s = 2, for higher level of item discrimination. For technical convenience, the data were simulated under the 1PL (Rasch) model in IRT, but the item analysis in this example is entirely under the RFM1 (s =1 and s = 2) on the D-scale. The data consist of simulated binary (1/0) scores on 20 items for a large sample of examinees (N = 2,983). The computations were performed using the computer program DELTA (Atanasov & Dimitrov, 2018). This program performs analyses under the RFM1, RFM2, and RFM3, allowing the user to fix the model-fit parameter for shape, s, at a desirable value (s > 0) under the RFM1.
The results related to the purpose of this example are summarized in Table 3. Based on the criterion for item fit (MAD < 0.07), the basic RFM1 (s = 1) provides a tenable fit for all 20 items, whereas none of the items exhibits a tenable fit under the RFM1when s = 2. As explained earlier, the item discrimination, a, varies across item locations, b, even though the fit parameter for shape, s, is fixed. However, if two items are symmetrically located around the mean of the D-scale (0.5), they have the same value for the discrimination parameter, a.
Estimates of Item Parameters for 20 Binary Items and Their Fit Index (MAD) with the Simulated Data in Example 1 Under the RFM1 with Two Fixed Values of the Fit Parameter for Shape (s = 1, s = 2).
Note. RFM1 = rational function model with one parameter; MAD = mean absolute difference; δ = expected item difficulty (δ
Example 2
This example illustrates the RFM1 and RFM2 on real data consisting of (1/0) scores on 25 items of the Reading Comprehension Section of the General Aptitude Test administered to 7,782 high-school graduates in Saudi Arabia. The data were analyzed under the RFM1 (s = 1) and under the RFM2 (s freely estimated) using the computer program DELTA. The results are summarized in Table 4. The examination of the item-person map produced by DELTA revealed an adequate match between the distributions of expected item difficulties, δ, and examinees’ ability scores on the D-scale (see Figure 7).
Estimates of Item Parameters for 25 Binary Items and Their Fit Index (MAD) for Reading Comprehension Data Under the Basic RFM1 and RFM2 in Example 2.
Note. RFM1, rational function model with one parameter; RFM2 = rational unction model with two parameters; MAD = mean absolute difference; δ = expected item difficulty; a = item discrimination;b = item location. Given in boldface are values that signal item misfit (MAD > 0.07).

Item-person map for data on the Reading Comprehension Section (25 items) of the General Aptitude Test from its administration to 7,782 high school graduates in Saudi Arabia.
Based on the criterion for item fit (MAD < 0.07), the results in Table 4 indicate a tenable data fit for all items under the RMF2, whereas there are five misfit items (3, 12, 14, 15, and 22) under the RFM1. For illustration, the IRF fit to the observed responses on Item 12 is depicted in Figure 8, with a misfit under the RFM1 (MAD = 0.090) and a tenable fit under the RFM2 (MAD = 0.028). Based on the results for item fit, the use of RFM2 over RFM1 is recommendable for the data in this example.

IRF fit to the data on Item 12, with a misfit signaled under the RFM1 (MAD = 0.090), left panel, and a tenable fit under the RFM2 (MAD = 0.028), right panel (see also Table 4). IRF = item response function; RFM1 = rational function model with one parameter; RFM2 = rational function model with two parameters; MAD = mean absolute difference.
Conclusion
Under the DSM for binary test items, the score of an examinee is based on his/her response vector on the items weighted by their expected difficult for a target population of examinees. The resulting D-scale ranges from 0 to 1 (for ease of interpretation in test reports the D-scores are presented on a scale from 0 to 100). The D-score of an examinee shows what proportion (or %) of the ability required for total success on the test is demonstrated by that examinee. The simplicity and efficiency of the DSM in psychometric procedures of test scoring and equating, item analysis, and their automated implementation have been evidenced in previous studies and practical applications in large-scale assessments at the National Center for Assessment in Saudi Arabia.
The estimation of examinees’ true scores on test items via the use of an item response model on the D-scale is a key feature of the DSM. In an initial effort on this matter, a two-parameter LRM was used (Dimitrov, 2017) as an analog to the 2PL model in IRT. However, the IRFs under the LRM tend to underestimate the item true score at the top end of the D-scale, especially for relatively difficult items. This is because the logistic IRF for such items tends to get closer to its upper asymptote (P = 1) beyond the right end of the interval on the D-scale (from 0 to 1). Thus, the use of the LRM on the D-scale (from 0 to 1) as an analog to the IRT logistic model (logit scale, from −∞ to +∞) is not psychometrically efficient. To solve this problem, the present study offers new models for IRF on the D-scale, referred to as rational function models (RFMs), with one parameter (RFM1), two parameters (RFM2), and three parameters (RFM3).
In some aspects the RFM1, RFM2, and RFM3 are similar to the 1PL, 2PL, and 3PL, respectively, in IRT. For example, (a) the IRFs do not cross under the RFM1 and 1PL (Rasch) model in IRT, (b) the IRFs may cross under the RFM2 and the 2PL model in IRT, and (c) the RFM3 and the 3PL model in IRT introduce a third parameter, c, to account for pseudo-guessing. However, it is important to emphasize that there are critical differences between the RFMs in the DSM and their logistic counterparts in IRT. Under IRT models, the person parameter (ability, θ) and item parameters (e.g., a, b, c in 3PL) are unknown and estimated simultaneously by using complex maximum-likelihood or Bayesian procedures (e.g., Baker & Kim, 2004; Fox, 2010). In contrast, under the RFMs, the person parameter (D-score) is estimated a priory (via Equation 1) and then used as an independent variable in the regression model of choice (RFM1, RFM2, or RFM3) to estimate the respective item parameters. Furthermore, the item discrimination, a, is not incorporated directly into the RFM but, instead, it is computed “post hoc” as a function of the RFM regression parameters for location and shape (see Equations 4 and 7). Also, the RFM1 and 1PL (Rasch) model are similar in that they do not allow for crossing IRFs, but under the 1PL the slopes of the IRFs are equal across all item locations on the IRT logit scale, whereas they are equal only for pairs of items with symmetrical locations around the mean of the D-scale under the RFM1. Thus, the RFMs under the DSM do not represent direct analogs to the IRT logistic models for binary items. Moreover, the estimation of D-scores does not require the use of an IRF model, which can be very useful in testing scenarios where issues of sample size, computational complexity, or strictness of assumptions can make the DSM preferable to IRT-based analysis.
Recommendations for the Practice
The selection of a proper model (RFM1, RFM2, or RFM3) is important for the accuracy of results and their valid interpretations in the DSM framework. As shown in Examples 1, the RFM1 performs very well when the data comply with the 1PL (Rasch) model in IRT. Under both the RFM1 and 1PL (Rasch) models, the methodological conception is that “the data must fit the model.” Conversely, under the RMF2 or RFM3 and their 2PL or 3PL counterparts in IRT, “the model must fit the data.” Thus, the RFM1 should be the model of choice from the start when the goal is to select items that “fit the model” in order to benefit from RFM1 features such as (a) simplicity of estimation procedures and computations, (b) nonintersection of IRFs, (c) parallel slopes of IRFs for pairs of items with symmetrical locations around the mean of the D-scale, (d) simple formulas for useful psychometric properties (see the appendix), and so forth. The RFM2 is suitable for cases where the data do not fit the RFM1, so IRFs may cross over the D-scale. A useful practice is to examine the estimates of the fit parameters for shape, s, across all items and it they are very close to a specific value (say, s≈1.0), one can fix s to that value thus reducing the RFM2 to the simpler model, RFM1.
The RFM3 typically provides better fit to the data compared with the RFM2, as it accounts for pseudo-guessing effects, but this comes with a “price” of problems associated with the use of the third parameter, c, in terms of (a) lower estimation accuracy; (b) interpretation of the item location—the place on the scale where the probability of correct item response is (1 +c)/2, that is, not 0.5; and (c) inappropriateness of comparing item difficulties (locations on the scale) when the third parameter, c, varies across items. For the case of 3PL in IRT, these problems are extensively discussed in a simulation study by Han (2012), but they remain under the RFM3 on the D-scale as well, so just ignoring them can be very damaging to the validity of RFM3-based results and their practical use.
In conclusion, the proposed RFM1, RFM2, and RFM3 models for IRFs of binary items extend and enhance the capability of the DSM and brings benefits of conceptual and technical simplicity and efficiency to the theory and practice of measurement and large-scale assessments in education and psychology.
Footnotes
Appendix
Author’s Note
The expected (true) D-score,
where
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
