Abstract
To reduce the chance of Heywood cases or nonconvergence in estimating the 2PL or the 3PL model in the marginal maximum likelihood with the expectation-maximization (MML-EM) estimation method, priors for the item slope parameter in the 2PL model or for the pseudo-guessing parameter in the 3PL model can be used and the marginal maximum a posteriori (MMAP) and posterior standard error (PSE) are estimated. Confidence intervals (CIs) for these parameters and other parameters which did not take any priors were investigated with popular prior distributions, different error covariance estimation methods, test lengths, and sample sizes. A seemingly paradoxical result was that, when priors were taken, the conditions of the error covariance estimation methods known to be better in the literature (Louis or Oakes method in this study) did not yield the best results for the CI performance, while the conditions of the cross-product method for the error covariance estimation which has the tendency of upward bias in estimating the standard errors exhibited better CI performance. Other important findings for the CI performance are also discussed.
Item response theory (IRT) is a widely used latent variable modeling method in educational and psychological measurement. Due to its versatility, IRT has been used for many applications such as test construction, differential item functioning, and test equating (de Ayala, 2009; Yen & Fitzpatrick, 2006). A major reason for IRT’s popularity was due to successful realizations of estimation methods for IRT models. For the item parameter estimation of IRT models, the marginal maximum likelihood (MML) method with the application of the expectation-maximization (EM) algorithm (Bock & Aitkin, 1981) and its variant where item parameters take priors in the context of the MML EM method (e.g., Mislevy & Stocking, 1989; Mislevy, 1986; Tsutakawa, 1992) are the most frequently used methods in research and practice.
When estimating an IRT model, the point estimates of model parameters are the first major interest (e.g., item difficult and item discrimination [or slope] parameter estimates). Uncertainty estimates such as confidence intervals (CIs) of model parameters, however, are important and may be of another interest in IRT applications because the interval estimates take into account sampling fluctuations, accommodating the extent of sampling variability in the point estimation. In estimating an IRT model with the MML method, the person ability parameters are assumed to take a certain distribution a priori (or simply stating that they are assumed to follow a distribution from the traditional frequentist point of view), such as normal distribution, while the item parameters do not take any prior distributions. In this “pure” MML approach, the item parameter estimates are the maximum likelihood (ML) estimates and their large-sample theory standard errors (SEs) are obtained typically from an observed or expected information matrix. The 100 ×(1
where
In the MML estimation, typically no prior distributions are used for item parameters. However, when a sample size is small, or to prevent unreasonable estimates or nonconvergence, prior distributions for item parameters may be used in the estimation process. For example, an application of the 2-parameter logistic model (2PLM) may exhibit unreasonable item slope estimates (e.g., very large or extremely uncommon values) or nonconvergence for several items due to Heywood cases. Also, in applications of the 3-parameter logistic model (3PLM), it is common to find the use of prior distributions for the pseudo-guessing parameters to reduce the issue of nonconvergence due to the model complexity by adding the pseudo-guessing parameters. When the item prior distributions are used in the MML-EM approach, the estimation method may be called a marginalized Bayesian EM approach, where the point estimates of item parameters are the marginal maximum a posteriori (MMAP) and the SE of item parameter is replaced by the posterior standard error (PSE). MMAP and PSE are the terms adopted for the MML-EM approach. They are to differentiate from MAP and posterior standard deviation (PSD) from a posterior distribution in a Bayesian approach using Markov Chain and Monte Carlo (MCMC). In the Bayesian modeling with MCMC, a credibility (or posterior) interval, which is similar to CI can also be obtained directly from the simulated posterior distribution where the 100
This interval estimation may be treated as an approximation of the 95% CI in the marginalized Bayesian EM approach to formula (1). This CI construction assumes that
Popular IRT programs often uses MML-EM as default, providing the marginalized Bayesian EM approach as an option. When item priors are employed, practitioners end up with the point estimates of items parameters (MMAPs) and their PSEs, and are more likely to resort to formula (2) to construct an interval estimate and use it as 95% CI. This study investigates the performance of the interval estimation procedure, formula (2) as 95% CI for the 2PLM and the 3PLM applications when item priors are taken under the marginalized Bayesian EM estimation approach. Currently, the literature does not show the behavior of formula (2) in the marginalized Bayesian EM estimation context. The 2PLM estimation usually works well without item priors, but there exist some situations, especially in some small samples, which use priors for item slope parameters because of improper solutions or nonconvergence in the item slope parameter estimation. Therefore, the focus of our investigations is on the item slope parameters in the 2PLM estimation when taking priors. For the 3PLM applications, the investigation centers on the pseudo guessing parameters when taking priors, given the common practice of using priors for the pseudo guessing parameters. Also in this study, the distributions of item slope and pseudo-guessing parameters for data generation are not the same as the prior distributions for the item slope and the pseudo-guessing parameters in the 2PLM and 3PLM estimation. (More details on the item prior distributions are provided later in the “Method” section.) In real data analysis, it is highly likely that chosen prior distributions do not match the true unknown distributions of parameters exactly. Given this reality, the parameter simulation distributions were chosen not to be identical with the specified prior distributions for the item slope and the pseudo-guessing parameters. In this sense, this study is also a sensitivity analysis for using widely known prior distributions, examining the degree of robustness of them as well.
The term SE is used from now on to denote SE in the MML-EM or PSE for simplicity. Also, the interval estimation procedures (1) and (2) will be referred to as, simply, CI construction procedures although formula (2) may be more properly called as posterior CI. As mentioned, in the MML-EM and the marginalized Bayesian EM approach, an SE method is required separately. Because SEs play a pivotal role in the CI construction procedures, they are reviewed briefly in the next section.
Methods for Standard Error Estimation
The descriptions of SE methods here are brief. Readers interested in more detailed technical treatment of those are referred to the references in the descriptions later. SE values are computed by taking the square root of the diagonal elements in the error covariance matrix which is the inverse of an information matrix. Because the EM approach only yields a complete-data information matrix in the process of estimation as a byproduct, methods to estimate an observed information matrix (or error covariance matrix) have been proposed in the IRT literature. In the MML-EM estimation context, those include the “cross-product” method (Meilijson, 1989; Pawitan, 2001.), the supplemented EM algorithm (S-EM: Cai, 2008; Meng & Rubin, 1991), the sandwich covariance matrix (White, 1982; Yuan et al., 2014), the Oakes method (Chalmers, 2018; Oakes, 1999; Pritikin, 2017)), and the Louis method with a refined computational approach (Liu & Chalmers, 2020; Louis, 1982). 1 Of these, the most studied methods up to now appears to be S-EM and the cross-product, the latter of which tends to be included as part of the comparisons with other major SE method(s) in the studies due to its relative simplicity and computational efficiency. The performance of S-EM was investigated by Cai (2008), Tian et al. (2013), Paek and Cai (2014), Pritikin (2017), and Chalmers (2018). The performances of the sandwich covariance, the Oakes, and the Louis were not fully documented yet by many studies in the IRT literature. As far as the authors are aware, Yuan et al. (2014) is the only published study on the performance of the sandwich procedure in IRT. To formulate the sandwich estimator, Yuan et al. combined the cross-product method in the IRT literature and the Louis (1982) information form. For the Louis (1982) method, a refined computational approach was proposed by Liu and Chalmers (2021). Their approach greatly enhanced the computational efficiency of the Louis procedure, demonstrating the refined approach’s potential across IRT and other latent variable models.
With respect to finite difference methods to approximate the observed information matrix, the Oakes method with finite difference applications (Chalmers, 2018; Pritikin, 2017) has been shown to be as good as or better than S-EM. More concretely, Pritikin (2017) implemented the Oakes method with the forward difference numerical approximation, comparing it with three types of S-EM procedures: original S-EM by Meng and Rubin (1991), a refined S-EM by Tian et al. (2013), and another refined S-EM by Pritikin (2016). Chalmers (2018) used the three numerical approximation methods (forward, central, and Richardson’s extrapolation) for the Oakes procedure and compared them with the refined S-EM by Tian et al., concluding that the central difference or the Richardson extrapolation method worked best. The Oakes procedure with numerical approximation was preferred to S-EM by both Chalmers (2018) and Pritikin (2017) for its accuracy and computational efficiency (calculation speed and reducing the occasions of nonconvergence in the calculation of S-EM error covariance matrix). Three SE methods were included in this study: cross-product denoted as CRP, Louis, and Oakes.
As mentioned before, the CRP method is considered as the most straightforward computationally relative to the Louis and the Oakes method. It is the outer products of gradients. The Louis observed information procedure is more complex than the CRP involving the second-order partial derivative, but it could provide a better sample-based approximation of the information matrix than CRP (see Yuan et al., 2014). The Louis observed information can be expressed as two components and one of the two is equivalent to the CRP (Liu & Chalmers, 2020; Yuan et al., 2014). Under the correct model specification assumption, as sample size N increases, CRP and the Louis observed information matrices becomes equal, the CRP component in the Louis method approaching zero (Yuan et al., 2014). The Oakes method, as the S-EM method,
2
is built upon the missing information principle, that is, the observed information matrix (
In the Oakes procedure, when items take priors, their first- and second-order derivatives (i.e., gradients and Hessian matrix) of the priors are added in the process of obtaining the complete data information and the missing data information to obtain the observed information. For the Louis procedure, Liu and Chalmers (2020) showed that its information, when item priors are used, is equal to the Louis information without item parameter prior plus the Hessian of the item parameter priors, that is,
Method
Simulations were conducted to assess the CI procedures, formula (1) in the MML-EM and formula (2) in the marginalized Bayesian EM approach.
IRT models used in this study are the unidimensional 2PLM and 3PLM. The 3PLM item response function is
where
The item response data were generated following the standard IRT data generation method. The steps to generate item response data are as follows. For a response to the jth item by the ith person, (a) compute
The simulation factors were SE methods, sample sizes, test lengths, and prior distributions for item slope parameters in the 2PLM and for the pseudo guessing parameters in the 3PLM. CIs were constructed by fitting correct models with and without item priors (i.e., 2PLM fit to 2PLM data and 3PLM fit to 3PLM data). The number of replications per condition was 1000. Data simulation and IRT model estimation with different SE methods under the MML-EM and the marginalized Bayesian EM approach where item priors are taken, were implemented using the “mirt” package (Chalmers, 2012) in the R statistical language and environment for statistical computing (R Core Team, 2020).
3
Convergence was checked using the default check provided by the “mirt” program and its second-order test of a unique modal solution. In addition, unreasonable or improper solutions were flagged if
SE Methods
As mentioned in the section of “Methods for Standard Error Estimation,” cross-product denoted as CRP, Louis, and Oakes methods were included in this study. CRP is computationally most straightforward and used as a default method in some IRT programs (e.g., IRTPRO by Vector Psychometric Group, LLC., 2021). The Louis method appeared to regain attention recently as an SE method in IRT (Liu & Chalmers, 2020) although its performance is not yet completely and fully documented yet (for scenarios where item parameter takes its prior as in this study). 4 Past studies (Chalmers, 2018; Pritikin, 2017) indicated Oakes performed the same or better than S-EM. The Oakes method is a default method in the “mirt” program. Note that although the three SE methods were originally included, the results showed that the Oakes and the Louis produced virtually the same results, therefore, in the presentation of the results, the Louis method is omitted.
Test Length
When data followed the 2PLM, a 20-item test and a 40-item test were simulated. Given that the 3PLM is popular in large-scale educational achievement testing which tends to have a relatively long test, 30-item and 50-item tests were simulated for the 3PLM data.
Sample Size
In the 2PLM data, sample sizes of 200, 900, and 1600 were used. The sample sizes for the 3PLM data were 1000, 1700, and 2400. Some small-scale preliminary analyses were conducted with smaller sample sizes (e.g., less than 200 for the 2PLM and 1000 for the 3PLM). Based on some preliminary analysis, large sample sizes (900 or 1600 for the 2PLM application and 1700 or 2400 for the 3PLM) were selected to examine the extent of the improvement in the approximation of the CR procedure, formula (2).
Prior Distributions for Item Parameters
When the 2PLM was fitted to the 2PLM data, the item slope parameters (

Prior Distributions For Item Slope and Pseudo Guessing Parameters
Descriptive Statistics for Prior Distributions and for Parameter Simulation Distributions.
Note. 2PLM = 2-parameter logistic model; 3PLM = 3-parameter logistic model.
Prior distributions of the slope parameters are positively skewed. LogN(0, 1) is flatter than logN(0, 0.5) and the difference between the mean (solid line) and the mode (dotted line) is larger in the logN(0, 1) than logN(0, 0.5). The expectation E(X), the SD, and the mode of
In the 3PLM fit to the 3PLM data, four types of prior distributions were used for the pseudo-guessing parameters (
Parameter Values for Data Generation
For both 2PLM and 3PLM, the item slope parameters (
Note again that, for the reason explained in the introduction, the prior distributions for the item slope and the pseudo-guessing parameters are not the same as those of item slope and pseudo-guessing parameters for data generation. (See the bottom panel in Table 1.)
Evaluation of CIs
For quantifying the performance of a CI procedure, coverage rate (CR) was used as outcome. CR is defined as
where
Results
The results from the Oakes SE method were virtually the same as those from the Louis SE method. The description of the results below, therefore, does not include the Louis SE method. Readers should be aware that the Louis SE method is always implied whenever the Oakes SE method is addressed. For those who are also interested in the item parameter recovery and the standard error estimation performance, their summaries are available upon request.
2PLM Data
In fitting the 2PLM to data, the slope parameters took lognormal(0, 0.5), lognormal(0, 1) or no prior distribution, while the item difficulty did not take any priors. Figure 2 provides CR for the slope parameters.

CR for the Item Slope Parameter (With Its Prior) in the 2PLM.
A larger sample size (N = 900 or 1600) increased the performance of CR, reducing bias (the deviation from .95). A short test, TL (test length) = 20, produced better performance of CR than a long test (TL = 40). The differences made by the SE methods and the prior distributions decreased in a larger sample and in a shorter test. The CRP SE method always produced larger CR than the Oakes SE method. When no prior is taken, the overestimation of CR in the CRP SE method is clearly noticed compared to the CR of the Oakes method. This is consistent with the past result that CRP tends to overestimate SE (Paek & Cai, 2014). Between the two prior conditions, the lognormal(0, 1) prior condition produced better CR performance than the lognormal(0, 0.5).
The percent bias of CR ( =
Percent Bias of CR for the Item Slope Parameter in the 2PLM.
Note. CR = coverage rate; 2PLM = 2-parameter logistic model; TL = test length.
Comparing the no-prior conditions of the two SE methods, the percent bias is larger in the CRP than the Oakes. Across all test lengths and sample sizes except for N = 1600 and TL = 20, the Oakes method with logN(0, 0.5) produced the largest size of bias, that is, the largest underestimation of CR, entailing significant sizes of bias. The logN(0, 0.5) prior condition, when combined with TL = 40, showed bias significantly deviating from .95 across all sample sizes. The logN(0, 1) prior condition yielded three cases of significant underestimation of CR with the Oakes SE method across all combinations of TL and N, while yielding none with the CRP SE method. The top four largest underestimation of CR (by −11.5% to −4.3%, which means CR = 83.5% to 90.7%) were found when the Oakes SE method was used with the lognormal prior distribution. The largest underestimation of CR was observed when a test was large (TL = 40), the sample size was small (N = 200), and the Oakes SE method was used with the logN(0, 0.5) prior distribution. Although the Oakes SE method is a better performer for the SE estimation than the CRP, as shown in the literature, the performance of the CI construction by formula (2) with the Oakes seems to suffer when item priors are used together. On the contrast, the overestimation tendency of SE in CRP appears to help with the performance of CR when the item prior distributions are used.
To check if there was any impact of the item slope priors on the CR of the item difficulty parameter which did not take any priors, the CR of the item difficulty parameters was examined. The CR for the item difficulty performed well in general, showing it is within the margin of error except for the small sample size of N = 200. The plot of CR and the percent bias of CR are shown in Figure 3 and Table 3.

CR for the Item Difficulty Parameter (With No Prior) in the 2PLM.
Percent Bias of CR for the Item Difficulty Parameter in the 2PLM.
Note. CR = coverage rate; 2PLM = 2-parameter logistic model; TL = test length.
Note again, the item difficulty did not take any priors. CR and the percent bias in Figure 3 and Table 3 are presented in each of the simulation conditions for the item slope parameters. (For example, logN(0, 0.5) on the x-axis (denoted as LN(0, 0.5)) does not mean that the item difficulty took the lognormal prior distribution.). The same pattern of the relatively higher CR for the CRP than the Oakes SE method was observed. The differences between CRP and Oakes decreased as the sample size increases and the test length decreases, which were observed in the item slope parameter case. When the sample size is N = 200, the range of the overestimation of CR in CRP was from 1.7% to 3.1%, meaning CR ranged from 96.7% to 98.1%. This overestimation bias of CR for item difficulty parameters in CRP is again compatible with the past report of the overestimation tendency of CRP in the literature. The item difficulty parameters did not take priors and the CR of the item difficulty with the Oakes method performed better than the CRP method. Overall, in the 2PLM estimation, the impact of the item slope priors on the CI of the item difficulty appears to be little except for a relatively small sample size (N = 200).
3PLM Data
In the estimation of the 3PLM, the pseudo-guessing parameter (

CR for the Pseudo-Guessing Parameter (With Its Prior) in the 3PLM.
Percent Bias of CR for the Pseudo-Guessing Parameter (With Prior) in the 3PLM.
Note. CR = coverage rate; 3PLM = 3-parameter logistic model; TL = test length.
The impacts of TL and N were not consistent across different prior conditions. For the two beta prior distributions, increasing N reduced CR underestimation, but that was not always the case in the logitnormal prior conditions. The SE method showed the similar pattern as in the 2PLM. The CRP SE method tended to produce a larger CR than the Oakes method. In both the logitN(–1.4, 0.5) and logitN(–1.4, 1) prior conditions, the CR bias of the CRP SE methods were all within the margin of error. The magnitudes of the CR bias for the Oakes SE method were within the margin of error in the logitN(–1.4, 1) prior condition only. The use of the beta priors resulted in very large underestimation of CR while the logitnormal priors yielded much improved results. When logitN(–1.4, 0.5) was used with the Oakes SE method, the CRs were slightly off from the margin of error, especially when TL = 50 and N = 1700 and 2400, whose CRs were .936 (–1.4% bias) and .933 (−1.7% bias).
The CR and the percent bias of the item slope parameter which did not take any priors are shown in Figure 5 and Table 5.

CR for the Item Slope Parameter (With No Prior) in the 3PLM.
Percent Bias of CR for the Item Slope Parameter (Without Prior) in the 3PLM.
Note. CR = coverage rate; 3PLM = 3-parameter logistic model; TL = test length.
In the conditions where the pseudo guessing parameter took the beta priors, the CR of the item slope parameter showed very large underestimation of CR (by min = 2.6% and max = 13.6% underestimation), falling outside of the margin of error. As before, the CRP SE method condition produced higher CR than the Oakes CR. The impact of the overestimation tendency of the CRP SE method on CR was also observed in the logitN(–1.4, 1) prior condition, making the CR fall outside the margin of error. The CR of Oakes SE method conditions fell within the margin of error for both logitN(–1.4, 0.5) and logitN(–1.4, 0.5) prior conditions. No uniformly consistent pattern of CR bias was observed across all levels of N and TL. The differences between the CRP and the Oakes SE method conditions were smaller in TL = 30 than TL = 40. The average difference of the percent CR (= 100 × CR) between CRP and Oakes SE methods was 0.8% (min = 0.6% and max =1.4%) in TL = 30% and 1.3% (min = 0.7% and max = 1.7%) in TL = 50.
For the item difficulty parameter which did not take any priors, Figure 6 and Table 6 show the CR and the percent bias.

CR for the Item Difficulty Parameter (With No Prior) in the 3PLM.
Percent Bias of CR for the Item Difficulty Parameter (Without Prior) in the 3PLM.
Note. CR = coverage rate; 3PLM = 3-parameter logistic model; TL = test length.
Overall, the patterns of the CR bias followed those of the item slope parameter which were mentioned above. The biggest impact factor was the prior distribution of the pseudo-guessing parameter as before. In the conditions where the beta priors were taken for the pseudo-guessing parameter, CR of item difficulty was underestimated to a very large extent and its percent bias ran the gamut from −11.1% to −30%. When the pseudo-guessing parameter took either logitN(–1.4, 0.5) and logitN(–1.4, 1) priors, all the CRs of item difficulty in both CRP and Oakes SE method conditions fell within the margin of error, though the tendency of the higher CR in the CRP than in the Oakes SE method was consistently observed as in all other parameter CIs. The differences of CR for the item difficulty between CRP and Oakes SE method across different test lengths were slightly smaller than those for the item slope parameter (by 0.1% in each test length condition).
Summary and Discussion
The MML estimation is the most popular method to obtain the IRT model item parameters. As IRT model complexity increases and to reduce the occurrences of nonconvergence or to make sure more reasonable parameter estimates, prior distributions for item parameters were made available in the MML estimation approach. With the use of priors for item parameters, the point estimates and their standard errors are MMAP are PSE. For the end-users of IRT programs which use the MML method with priors for item parameters, the estimates of MMAP and PSE are available as output from the estimation. When an interval estimation such as CI is desired, a straightforward approach to construct 95% CI with the available output of MMAP and PSE in the marginalized Bayesian EM approach is to use
When data follow the 2PLM, the performance of the CI improved as N increases and TL decreases. The CR of the CRP SE method was higher than that of the Oakes SE method. This was observed for all the CIs for all parameters. This seems to be due to the overestimation tendency of the CRP SE estimation, which is reported in the past literature. The differences of CR between CRP and Oakes SE method conditions decreased as N increases and TL decreases. For the item slope parameter with its lognormal prior distribution, the Oakes SE method condition produced poor performance compared to the CRP SE method. However, when no prior distribution was used, the Oakes SE method condition yielded more accurate results of CR than the CRP SE method for both item slope and difficulty parameters. The worst condition for CR for both SE methods was where a test is large (TL = 40) and a small sample size (N = 200) was used. When the lognormal prior for the item slope parameter was combined with N = 200 and TL = 40, the underperformance of the CI for the item slope parameter was the most severe one in the Oakes SE method condition (e.g., −11.5% in the percent CR bias, that is, CR = 0.835).
When data follow the 3PLM, the pattern of the higher CR in the CRP SE method than the Oakes SE method was repeated as in the 2PLM applications. The use of the different priors for the pseudo-guessing parameter had the largest effects on the performance of CI for all item parameters. The conditions where the beta prior distributions for the pseudo-guessing parameter were employed exhibited very poor performance of CI, showing the CR bias ranging from −68% to −34.3% for the pseudo-guessing parameter; from −14.3% to −2.6% for the item slope parameter; and from −30.8% to 16.4% for the item difficulty parameter.
Across 2PLM and 3PLM data simulation conditions, the strongest factor affecting the CI performance was the use of prior distributions. Compared to relatively strong prior conditions, weak prior conditions (
In addition to this general observation, several peculiar observations which were not anticipated deserve attention. First, across 2PLM and 3PLM applications, the performances of CIs for the item parameters which did not take any priors (item difficulty parameter in the 2PLM and both item slope and difficulty parameters in the 3PLM) were consistently poor when the item slope in the 2PLM and the pseudo-guessing parameter in the 3PLM took, especially, their relatively strong prior distributions. Specifically, in a small sample size of N = 200 in the 2PLM estimation, the CR of item slope parameter with the lognormal(0, 0.5) prior and the CR of the item difficulty parameter with no prior distribution were mostly out of the margin of error. In the 3PLM estimation with a relatively small sample size of N = 1000, all beta priors and the logitnormal(−1.4, 0.5) prior for
The CI in this study was constructed by formula (2). Formula (2) is an approximation to the CI for the MLE. The results of this study can also be interpreted to indicate that the CI procedure using formula (2) is not working well unless a sample is sufficiently large, that is, the assumed normal approximation in formula (2) using MMAP and PSE is not very good unless N is large enough. Given that the MML estimation with priors of item parameters are used in practice to estimate IRT models, the CI construction by formula (2) could be often adopted when CIs are desired. The impact of using the priors for item parameters in small samples on the CI construction by formula (2) could be very negative and severe as shown in N = 200 with lognormal(0, 0.5) for the item slope parameter in the 2PLM estimation and N 2400 with beta priors—B(5, 17) and B(2, 5)—for the pseudo-guessing parameter in the 3PLM. The observed underperformance of the CI procedure by formula (2) does not mean that PSE underestimates the variability of the sampling distribution of MMAP. As mentioned before, PSE tends to be smaller than the MLE’s SE. The reason for PSE smaller than MLE’s SE is that the sampling distribution of MMAP itself has less variability than that of MLE. As a sample size increases, the extent of shrinkage in MMAP reduces and PSE approximates the MLE SE more closely in theory. One may, therefore, expect that the sampling distribution of MMAP get closer to that of MLE as sample sizes get larger, thereby
To construct less biased CI using formula (2) with MMAP and PSE from the marginalize Bayesian EM approach, general recommendations based on this study include using a large sample and a weak prior distribution. In real data analysis, the first choice of increasing a sample size is often not possible and the second option is expected to be the choice more often. This study clearly demonstrated the extent of bias which stems from the CI construction procedure by formula (2) when priors which are widely known and used in research and practice are adopted for item parameters under the marginalized Bayesian EM approach. Though the above recommendations sound plain, one should be careful when dealing with a small sample size and the chosen priors because the CI constructed by formula (2) can be severely biased. Specifically, when there is no information or reasonable conjecture to make a choice for a prior distribution of the pseudo-guessing parameters, the use of beta priors with the same distribution parameter specifications used in this study does not seem to be an easy recommendation to be made in the 3PLM estimation when N
The study used the priors known in the literature or used in practice. The choice of a particular distribution and its parameters for the selection of priors of item parameters may not be the same as in this study. Future endeavors may include different priors, for example, normal distributions for the item slope parameter or the same but different specifications of the parameters of the prior distributions. In terms of the CI construction by formula (2), the use of a weaker prior is generally preferable to the extent that there is no estimation issue (e.g., nonconvergence or unreasonable estimates) incurred by the weaker specification of the prior. For practitioners who are more likely to stick to those widely known priors as used in this study, weaker prior specifications such as beta distributions with wider SD and the logitnormal with wider SD, for example, logitN(–1.4, 2) may be tried. (It should be noted that using a very weak priors increase the chance of convergence problems in small samples or complex models such as the 3PLM.) Exploring diverse and weaker forms of prior distributions for the investigation of the CI construction by formula (2) in the MML estimation with priors for item parameters seems to be another line of future applied research. Exploring diverse sample sizes and test lengths may also be considered at the same time. This study used up to N = 2400, which was not successful to improve the performance of CI by formula (2) in the 3PLM estimation with some prior distributions. It would be interesting to examine the effects of larger sample sizes and weaker specifications of the beta priors simultaneously, which could reveal the conditions that improve the performance of CI constructed by formula (2). Also, investigations at the item level to identify the associations of particular constellations of item parameters with better or worse CI performance may be pursued.
Finally, we would like to remind readers that our study is a simulation study which can reflect real data applications as much as the study design covers important aspects of real data analysis. The points addressed above should be cautiously applied as the characteristics of real analysis deviates from those incorporated in this study.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
