Abstract
The traditional application of the Montreal Cognitive Assessment uses total scores in defining cognitive impairment levels, without considering variations in item properties across populations. Item response theory (IRT) analysis provides a potential solution to minimize the effect of important confounding factors such as education. This research applies IRT to investigate the characteristics of Montreal Cognitive Assessment items in a randomly selected, culturally homogeneous sample of 1,873 older persons with diverse educational backgrounds. Any formal education was used as a grouping variable to estimate multiple-group IRT models. Results showed that item characteristics differed between people with and without formal education. Item functioning of the Cube, Clock Number, and Clock Hand items was superior in people without formal education. This analysis provided evidence that item properties vary with education, calling for more sophisticated modelling based on IRT to incorporate the effect of education.
Keywords
Detection of dementia is a clinical and research imperative. The current estimate of dementia prevalence worldwide is 47 million, projected at over 130 million by 2050, with a larger growth in low- and middle-income countries (Alzheimer’s Disease International, 2013; Livingston et al., 2017). The majority of people with dementia are undiagnosed (Prince, Bryce, & Ferri, 2011). Pharmacological and nonpharmacological interventions have been recommended to delay cognitive decline in mild dementia (National Institute for Health and Clinical Excellence, 2011); and early detection appears to be associated with milder cognitive impairment (Tang et al., 2016). However, until disease-modifying treatment or targeted prevention strategies become available, clinical diagnosis remains the mainstay in the detection of dementia. This strains the mental health professional workforce and health care costs substantially. An efficient and cost-effective triage system is needed, which could include data-informed decision-making strategies, such as symptoms or complaints patterns, and cognitive screening tools with validated population-specific cutoff values (Xu et al., 2017).
The Montreal Cognitive Assessment (MoCA) is one of the most commonly used cognitive screening tools for dementia, first published in 2005 (Nasreddine et al., 2005). A PubMed/MEDLINE search using the keyword combinations “(Montreal Cognitive Assessment[Title/Abstract] OR MoCA[Title/Abstract]) AND (validation[Title/Abstract] OR psychometric properties[Title/Abstract])” yielded 129 results regarding its application to various populations, suggesting considerable differences in item characteristics and cutoff scores for detecting cognitive impairment in these populations. These studies often fall into one of three categories: (a) smaller-sample studies investigating criterion validity against clinical judgment in differentiating between normal, mild cognitive impairment, and dementia groups; (b) norm studies using a statistical approach to determine the range of normal functioning based on an assumption of normally distributed total test scores, with reference to 2, 1.5, and 1 standard deviation(s) from the mean; and (c) validation studies either examining construct validity using confirmatory factor analysis (CFA) or item/domain properties using item response theory (IRT).
A challenge with these validation methods is the great variability across populations with different education levels, which mandates multiple cutoff values that complicate screening efficiency with increased potential for human error. For example, a population-based study in a mainland China sample of over 8,000 older persons recommended very different cutoff scores (13 vs. 24) for illiterate individuals versus those with 7 or more years of education (Lu et al., 2011). Similarly, a Hong Kong study of over 1,700 participants recommended cutoff values ranging from 9 to 25, depending on age and education levels. The authors concluded that conventional single cutoff scores risk misclassification and cautioned against a one-size-fit-all approach (Wong et al., 2015). A potential source of the difference in cutoff values is item bias. As stated by Balsis, Choudhury, Geraci, Benge, and Patrick (2018), item bias refers to the situation where participants in one group are more likely to endorse an item than participants of another group due to some characteristic of the item that is not related to the latent construct. When conducting cognitive assessment for dementia (including MoCA and other inventories), a key concern is whether sociodemographic factors (particularly educational level) influence item responses across the population; that is, whether a participant is more likely to score on an item due to his or her educational background rather than cognitive ability. Classical test theory emphasizes total scores, which are calculated with each test item equally weighted. In MoCA, to account for the effect of the important confounding factor of education, one or two extra points are often added to the total score. This kind of approach is somewhat arbitrary and does not consider the properties of each item in relation to confounding factors affecting performance.
IRT, an important representative of modern psychometric theory, provides statistically robust methods in this context to minimize the effect of education in the testing of true cognitive ability, without the need to develop multiple versions of the test. IRT has been developed and widely used in the field of education to improve the precision of ability and achievement measurement, such as the Graduate Record Examination (Embretson & Reise, 2013). Although the value of IRT in clinical assessment is gradually being recognized, its application in improving the measurement precision of cognitive assessment is still limited (Balsis et al., 2018; Thomas, 2011). A PubMed/Medline search conducted in July 2018 identified only four studies of MoCA in which IRT analysis was employed. The key findings of these studies are briefly summarized here. In 2011, a Taiwan team applied IRT to a mixed sample of 207 participants comprising patients diagnosed with Alzheimer’s disease, people with mild cognitive impairment, and healthy controls (Tsai et al., 2012). They concluded that the best performing domain was the frontal domain, which includes attention, concentration, working memory, and abstract thinking. Another Chinese sample study conducted separate IRT analyses on Cantonese- and Mandarin-speaking participants from the Chinese American Eye Study (Zheng et al., 2012). They found that the executive function domain performed the best. It should be noted that both of these studies reported only combined domain level results and did not report the performance of each item. To shorten the administration time of MoCA, Roalf et al. (2016) employed IRT and computerized adaptive testing in 1,850 individuals with and without neurodegenerative disease using a short form of MoCA (s-MoCA) consisting of eight items with high discrimination and appropriate difficulty. Clock Drawing (from the visual/executive domain) and Serial subtraction (attention domain) were found to be the most informative of the eight items. Using the same methodology, a Czech group developed a short-form assessment from the standard Czech version of MoCA (Bezdicek et al., 2018). Among the eight items they identified, six overlapped with s-MoCA. Trail Making and Clock Drawing were revealed as the best items. To summarize, although results are specifically related to the characteristics of the population being studied, certain items, such as Clock Drawing and Serial Subtraction, have consistently demonstrated superior performance. It is worth noting that although both the Chinese American and Czech studies have raised concerns about the potential effect of age, gender, education or culture on measurement precision, investigation was only carried out in the total score level. More advanced item level analysis represented by differential item functioning (DIF) has yet to be conducted.
In the study reported in this article, we applied IRT analysis to investigate the characteristics of individual MoCA items in a culturally homogeneous sample of older persons with diverse educational backgrounds. We propose to account for item bias by measuring DIF for individual items between participants with different educational levels. The implications of using these two approaches on categorizing cognitive impairment levels are discussed.
Method
Participants
Data were available from the baseline assessment of a longitudinal study on health and well-being of Cantonese-speaking older persons in Hong Kong (Liu et al., 2017). Participants were recruited from 12 public rental estates using stratified random sampling. Older tenants in each estate were classified into three age strata: from 65 to 74 years, 75 to 84 years, and 85 years and older. Participants were randomly sampled from each strata with target sample sizes set to 50, 60, and 70, respectively, from each estate. Individuals who were older than 85 years were purposely oversampled as they are likely to be frailer and have higher levels of cognitive impairment. The assessments were conducted by trained interviewers between July and November 2014. Participants were asked to self-report any diagnosed chronic diseases from a list of 31 conditions. Those with a known dementia diagnosis or a known psychiatric disorder were excluded from this analysis. Relevant data were retrieved from 1,873 participants in this study. Ethical approval was obtained from the Review Board of the Human Research Ethics Committee for Non-Clinical Faculties at The University of Hong Kong.
Montreal Cognitive Assessment
The Chinese (Cantonese) MoCA (version 7.0) was used in this study, which has been translated/adapted and validated in Hong Kong (Chu, Ng, Law, Lee, & Kwan, 2015). It covers eight domains, including short-term memory, visuospatial ability, executive function, attention, concentration, working memory, language, and orientation (Nasreddine et al., 2005). Items from each domain can be found in Table 1. These domains are theoretical domains informed by clinical observations rather than empirically identified dimensions informed by psychometric theories. The MoCA total score, ranging between 0 and 30, is the sum of scores on each test item, calculated with each item equally weighted. In this study, no additional points were given to low education participants.
Sample Characteristics by Educational Level.
Education Level
Educational level was recorded in the following categories: no formal education, primary school, junior school, high school; postsecondary (non-degree); and college and above. Here no formal education means that the participant has never attended school of any kind. Since only about one fifth of the participants received an education at junior school or above, educational level was recategorized as “any formal education” (yes/no).
Dimensionality Analysis
Latent dimensionality of MoCA was evaluated using exploratory factor analysis (EFA) and CFA. Up to five dimensions were considered. The principal axis method with oblique oblimin rotation was used, together with the polychoric correlation matrix. A large ratio (>3) of the first to second eigenvalues was used to determine if a unidimensional interpretation of the scale was appropriate for proceeding with the item response analysis. For the EFA, mean residuals equal to 0, a standard deviation of the residuals less than 0.05, and a goodness-of-fit measure equal to or higher than 0.9 were considered a good fit (Reise et al., 2011). As simultaneous identification of multiple factors and observation of evidence for a unidimensional structure is very common, we further explored the extent to which “multidimensional data yield univocal scale scores” using a confirmatory bifactor model as proposed by Reise, Moore, and Haviland (2010). We estimated a bifactor model with a single general factor and several orthogonal group factors suggested by the previous EFA model. Polychoric correlations and the diagonally weighted least squares estimator were used for estimating the bifactor model (Yang-Wallentin, Jöreskog, & Luo, 2010). According to Rodriguez, Reise, and Haviland (2016), the explained common variance (ECV) can be used to determine whether the measures are essentially unidimensional and, hence, subsequent IRT analysis can be carried out. They have suggested that when ECV > 0.7, the measure can be regarded as essentially unidimensional. The bifactor model was estimated using the lavaan package of the statistical software R (R Core Team, 2017; Rosseel, 2012). The remaining analyses, including the ECV calculation, were conducted using the psych package (R Core Team, 2017; Revelle & Zinbarg, 2009).
Item Response Theory Analysis
Similar to CFA models, IRT views the constructs of interest (e.g., cognitive ability) as latent variables that cannot be observed directly. IRT models are built on the assumption that the probability of a participant responding to an item correctly is a function of two sets of parameters: their position on the latent continuum of interest (person parameter) and the properties of the item (item parameters). This relationship is typically depicted by item characteristic curves (Embretson & Reise, 2013). To numerically describe this function, one or more parameters will be employed depending on which IRT model has been adopted. Details of different types of IRT models are documented by Thomas (2011). Two basic parameters are item difficulty and item discrimination, which are related to the probability of correct response and to what extent the item can differentiate participants with different levels of the construct, respectively (DeMars, 2010). The two-parameter logistic model, a widely utilized model which allows unique difficulty and discrimination parameters for each item, was employed in this study.
Other key concepts in IRT are item information and test information. Item information is inversely related to standard error of measurement. High information is equivalent to lower standard errors and more precise latent variable estimates. Information is a function of the value of the latent variable and, hence, is different for participants with different abilities. Test information refers to the sum of the information of each item and is related to the precision of measurement of the latent variable based on all the items.
Given that several items of MoCA are polytomous, the graded response model (GRM; Samejima, 2016) and the generalized partial credit model (GPCM; Muraki, 1992), which are both extensions of the two-parameter logistic model for dichotomous variables, were considered and compared. The GPCM can be viewed as a series of dichotomous models, where the probability of responding in one category as opposed to responding in an adjacent category is modeled (de Ayala, 2009). The GRM, meanwhile, models the probability of obtaining a certain item score or higher. We estimated the IRT models using marginal maximum likelihood with the R package mirt (Chalmers, 2012). The Bayesian information criterion (BIC) was used as the model selection criterion (Schwarz, 1978). The model with lower BIC was selected. After selecting the functional form of the IRT model, education level was used as a grouping variable to estimate multiple-group IRT models.
As mentioned by Kim, Cohen, Cho, and Eom (2019), detection of DIF can be viewed as a model selection problem. In our context, this means that we consider different multiple-group models that allow for differences in the item parameters between the groups. To investigate potential DIF, we used a model selection procedure that compared the BIC from multiple-group IRT models with all possible levels of invariance at the item level. The model with the lowest BIC was selected. This approach has been used in many previous studies (Cho, Suh, & Lee, 2016; Kim et al., 2019). With the estimated values of the parameters from the multiple-group model, we calculated the ability estimate for each individual. In this analysis, the latent variable, namely cognitive ability, was estimated conditional on the estimated item and distribution parameters, and these parameters were thus treated as fixed and known. For the latent variable estimation, we used the posterior mean, also known as the expected a posteriori estimator (Bock & Aitkin, 1981).
Fit Statistics
To evaluate the IRT model fit, the limited information goodness-of-fit test statistic was used to compute root mean square error of approximation (RMSEA). An RMSEA value of less than 0.05 is an indicator of good model fit (Cai & Monroe, 2014; Maydeu-Olivares & Joe, 2006). Item fit was assessed using Orlando and Thissen’s S-χ2 item fit index for polytomous IRT models (Kang & Chen, 2008).
Reliability Estimation
Reliability is generally defined as the ratio of true-score variance to observed-score variance. The true score is “the hypothetical average of the observed scores that would be obtained” if measurement was repeated an infinite number of times (DeMars, 2010). Cronbach’s alpha is the most extensively used reliability measure in applied research. However, it is well recognized that Cronbach’s alpha is only the lower bound to the reliability as its assumptions almost never hold. Other commonly used alternatives for measuring reliability that originate from factor analysis are McDonald’s ω and maximal reliability ρ (Bentler, 2007; McDonald, 2013). In this study, we used the IRT-based reliability coefficient to investigate the consistency of the MoCA scores (Cheng, Yuan, & Liu, 2012). Since explicit expressions for the reliability of IRT multiple-group ability estimates and total scores are not available in the literature, we used Monte Carlo simulation to estimate them. In each replication, we independently generated two sets of item responses for each group based on the estimated item and distribution parameters. Pearson correlations of the resulting total scores and ability estimates for the two groups between two sets of item responses were then calculated. The means of the Pearson correlations across 1,000 replications were then taken as estimates of the reliabilities of the total scores and ability estimates for each group. Details of the simulation procedure are presented in the supplemental document (All supplemental materials are available in the online version of the article.).
Classification
Cutoff values for neuropsychological testing have been proposed for different levels of cognitive impairment. According to the Diagnostic and Statistical Manual of Mental Disorders–Fifth edition, cognitive performance criteria for major neurocognitive disorders can be conceptualized as 2 standard deviations below the group mean, or approximately the 2nd percentile (American Psychiatric Association, 2013). Using Petersen’s revised diagnostic criteria of mild cognitive impairment, the criterion can be conceptualized as 1.5 standard deviations from the group mean, or approximately the 7th percentile (Petersen et al., 2009). These percentiles of a normal distribution have been used for classification with the total scores for groups of individuals with similar levels of education (Wong et al., 2015). We use the 2nd and 7th percentiles to explore the implication of using cutoff values derived from the total scores and cutoff values generated from IRT ability estimates on the classification properties.
Specifically, we investigated to what extent the methods provided the desired classification rates (i.e., 2% for major neurocognitive disorder and 7% for mild cognitive impairment) and the consistency between the methods with respect to the classification of individuals. In the context of IRT models, the participant’s cognitive ability is treated as a latent variable that follows a normal distribution. For the multiple-group analysis, the group with formal education is set as the reference group with the latent construct following a standard normal distribution. The latent construct of the group without formal education was assumed to follow a normal distribution with mean and standard deviation estimated from the data. For the IRT ability estimates, cutoff values are the theoretical percentiles of the normal distributions.
Results
Of the 1,873 participants, 851 (45.4%) had no formal education. In Table 1, we compared the basic demographics and MoCA scores of participants with and without formal education. The independent sample t test and Chi-square test were used to compare the means of continuous variables and percentages of categorical variables in the two groups, respectively. Participants who had no formal education were significantly older (81.5 ± 7.3 vs. 77.5 ± 8.1 years), more likely to be women (64.3% vs. 45.3%), and scored much lower in MoCA (15.9 ± 5.9 vs. 20.9 ± 5.3).
Traditional Item Statistics
Table 2 shows the traditional item psychometric properties for the total sample, including item-total correlations and the response proportions for each possible outcome per item. Subgroup response proportions are available in Supplementary Table 1. For the total sample, the raw and standardized Cronbach alphas of the entire scale were 0.79 and 0.83, respectively. The item-total correlations ranged from 0.33 (Clock Shape) to 0.72 (Short-term Memory) with an average of 0.54. A preliminary examination of the response proportions showed a potential ceiling effect of the item Clock Shape as more than 90% of the participants answered the item correctly.
Abbreviated Item Content and Descriptive Statistics for the Chinese (Cantonese) MoCA (Total Sample, N = 1,873).
Note. MoCA = Montreal Cognitive Assessment.
Dimensionality Assessment
A series of EFA models were estimated to investigate the appropriateness of a unidimensional interpretation of the scale and the results are summarized in Table 3. The eigenvalue ratio of the first to second eigenvalues for the single-factor model was 4.75. The single-factor model has mean residuals equal to 0, standard deviation of residuals of 0.08, and a goodness-of-fit measure of 0.96. The three-factor model has mean residuals of 0, standard deviation of residuals 0.04, and a goodness-of-fit of 0.99. Heywood cases were encountered with four and five-factor models and, hence, the results of these models were omitted. Using the factor structures identified based on the three-factor EFA, we fitted a bifactor model with three specific factors. The comparative fit index, Tucker–Lewis index, and RMSEA values were 0.984, 0.977, and 0.037, respectively. The ECV for the general factor equals 0.76, indicating that the common variance can be viewed as unidimensional. These results collectively support a unidimensional interpretation of the scale.
Factor Loadings from the Exploratory Factor Analyses of the Chinese (Cantonese) MoCA (Total Sample, N = 1,873).
Note. MoCA = Montreal Cognitive Assessment. Numbers in bold indicate the highest factor loading among the three factors.
Item Response Theory Analysis
IRT models were fitted to the data assuming a single latent variable. The initial model selection using a single group indicated that the GRM was the better fitting model, with BIC equal to 41036.95 compared with 41116.67 for the GPCM.
As shown in Figure 1, the test information is in general higher for the no formal education group. Figure 2 shows the item information functions of all items. Most items are substantially informative. Several items provided considerably different amounts of information for people with different cognitive ability. The Trail Making item provided the most information when the participant’s cognitive ability was slightly higher than average, but very little information when the participant’s cognitive ability was more than 1 standard deviation below the average. The items Clock Shape, Naming, Digit Span, Serial Subtractions, and Orientation were more sensitive in discriminating participants who had lower cognitive abilities. Items such as Logical Memory and Verbal Fluency provided relatively less information.

Test information function of the Chinese (Cantonese) MoCA by education level.

Item information functions of the Chinese (Cantonese) MoCA items using both single-group and multiple-group analyses across levels of cognitive ability.
Using educational level as the group variable, the model selection procedure for the multiple-group models identified three items that had DIF with respect to education, namely Copy of the Cube, Clock Number, and Clock Hand. In these items, the information provided is higher for the no formal education group compared with the group with some formal education, with the most substantial differences observed in the items Clock Number and Clock Hand. For people whose cognitive abilities are about 1 standard deviation below the group averages, the Clock Number item provided more than three times more information for the no formal education group compared with the group with some formal education. The item characteristic curves are presented in Figure 3. Figures 2 and 3 also show that the three items function differently for the two groups.

Item characteristics curves of the Chinese (Cantonese) MoCA by education level.
According to Petersen et al. (2009), mild cognitive impairment is defined as 1.5 standard deviations below the group mean. Multiple-group item response analysis showed that for the group without formal education, the mean and standard deviation estimates of cognitive ability were −1.08 and 1.09, respectively. We hence fixed cognitive ability at −1.5 and −2.715 for the group with and without formal education, respectively, and ranked the item from high to low by item information in Table 4. There were considerable differences among values of item information, ranging from 0.09 to 0.62 in the group with formal education and 0.01 to 0.66 in the no formal education group. Orientation, Serial subtraction and Naming consistently provided high levels of information for both groups. Trail Making and Cube consistently provided relatively low levels of information. It should be noted that since item information can be very different given different level of cognitive ability, the ranking provided here is limited to a very specific situation.
Item Information by Educational Level, Ranked From High to Low.
Fit Statistics
The RMSEA based on the statistic of the multiple-group model was 0.042 (95% confidence interval [0.039, 0.045]), indicating a good model fit. The S-χ2 statistics for each item in each group with their corresponding p values are summarized in Table 5. The test statistic originally identified three misfitting items in each group based on a significance level of 0.05. However, if a Bonferroni correction is applied to account for the problem of multiple comparisons, only the Clock Shape item in the group with no formal education is significantly misfitting, based on the adjusted significance level of 0.0018. We further examined the empirical item characteristic curve plot for the misfitting item, which compares the empirical response function with the IRT model-based response function. The empirical plot is shown in Figure 4. It is seen that the misfit only occurs at the lower end of the continuum of cognitive ability and that the misfit is not severe.
S-χ2 Item Fit Indices and the Corresponding p Values.

Empirical plot of the Clock shape item.
Reliability
Because of the different properties of three of the items and the differing distributions of the latent variable in participants with and without formal education, the reliabilities of the total scores and the ability estimates differed between the two groups. For the group without formal education, the reliability of the total scores is estimated at 0.79 and that of the ability estimates at 0.83. For the group with formal education, the estimated reliability of the total scores is 0.73 and that of the ability estimates is 0.78. Hence, according to the multiple-group IRT model, the Chinese (Cantonese) MoCA is more reliable for the group without formal education and scoring the MoCA using ability estimates is more reliable than using the total scores for both groups.
Classification
Table 6 shows the cutoff values for the 2nd and 7th percentile and the corresponding percentages of participants classified in each category using either the total score or the ability estimate. In the group with formal education, an ability estimate cutoff value of −2.054 was suggested for the 2nd percentile, indicating that a participant will be classified as meeting the criterion for major neurocognitive disorder if his or her ability estimate is more than 2.054 standard deviations lower than the group mean. An examination of the percentage of each category revealed that using the ability estimates actually matched the desired percentages (i.e., 2% and 7%) for the group with formal education. The classification consistency between the total score and ability estimates is high, ranging from 94.9% to 99.8%.
Cutoff Values and Corresponding Percentage of Participants Classified in Each Category Using Total Score and Ability Estimates.
Discussion
To our knowledge, this is the first study to investigate the potential use of an IRT approach to enhance measurement precision in a commonly used dementia cognitive screening test, MoCA, focusing on the effect of education on item performance. Our results showed that IRT is appropriate when applied in MoCA, and individual item performance varied in providing information to differentiate levels of cognitive ability, particularly when formal education is considered. Ability estimates from the MoCA provide an alternative to unweighted total scores in improving measurement precision of participants according to their cognitive ability, considering different effects of education on individual item functioning. More reliable cutoff values for classification can be obtained when the criterion variable becomes available.
Our findings have three practical and research implications. First, our findings from a large and culturally/ethnically homogenous population indicated the validity of MoCA in assessing a single latent construct of global cognition, while the disagreement between the true distribution and the assumed normal distribution showed that cutoff values are a source of error that can be addressed statistically. In line with many previous studies, our analyses showed adequate psychometric properties and classification accuracy of MoCA (Koski, Xie, & Finch, 2009; Lam et al., 2013; Tsoi, Chan, Hirai, Wong, & Kwok, 2015). This provided a basis for its further application in triage systems for the detection of dementia, and in monitoring changes in global cognition over time. In a recent Cochrane review involving 9,422 participants, the authors concluded that there is currently insufficient evidence regarding the clinical utility of MoCA in detecting dementia, as most studies included small samples of case-control design, and a low specificity of ≤0.60 is noted when a single cutoff score of ≥26 (for normal cognition) is used (Davis et al., 2015). Apart from highlighting the inadequacy of using a single cutoff, which has been discussed by other authors, our results demonstrated the measurement precision advantage by adopting statistically guided methods such as IRT to tackle the noise introduced by confounders, such as education (Wong et al., 2015).
Second, we noted an interesting phenomenon of superior performance of MoCA in those without formal education, as compared with those who had received formal education. Similar results have recently been noted in another study using receiver operating characteristic curve analysis, which showed an area under the curve for mild neurocognitive disorder of 0.90 in people with primary education or below, while that for people with secondary education or above was 0.66 only (Liew, Feng, Gao, Ng, & Yap, 2015). Our item information analysis by education showed that three items, all within the visuospatial domain, make the major contribution to this difference. While cognitive processes involved in these items are likely complex, including spatial planning, conceptualization, visuomotor coordination, attention, and symbolic representation, education, and learning experience have known effects on both Copy of the Cube and Clock Drawing Test (Julayanont, Phillips, Chertkow, & Nasreddine, 2013). The role of learning in three-dimensional form perception is well recognized, which may explain the better performance of the Copy of the Cube in our sample with no formal education who are relatively free from a learning effect (Sinha & Poggio, 1996). On the other hand, a commonly reported feedback from our sample with no formal education is the novelty of the experience in using a pen/pencil to draw a figure to represent a complex visual image. Such an ability—to rapidly learn a new task from instructions—has been underresearched and to which recent attention has been drawn called under a proposed term “rapid instructed task learning” (Cole, Laurent, & Stocco, 2013). Whether the performance of a new task according to instruction per se adds to the cognitive ability requirement of the Clock Number and Clock Hand (but not Clock Shape) items, thereby increasing their item information in those with no experience learned from school, should be further investigated.
Third, modern psychometric techniques, represented by IRT, should be more widely applied to improve the measurement precision of important clinical scales such as the MoCA (Balsis et al., 2018; Thomas, 2011). The issue of nonignorable variation in MoCA item characteristics across populations with difficult sociodemographic backgrounds has been well recognized by both researchers and clinicians. To address this challenge, modified versions of MoCA have been developed (“MoCA Basics”), which target participants who are illiterate or with low education. On the other hand, electronic cognitive assessments (e.g., an App for use on a tablet) are becoming more available, which streamline screening by allowing automatic scoring using a complex algorithm and multiple cutoff values. IRT analysis, which can also be incorporated into electronic assessment Apps, thus provides a potential solution to minimize the effect of education in the testing of true cognitive ability, without the need to develop multiple versions of the test.
This study has several limitations. First, ability estimates are sample specific due to the self-adjusting nature of the analyses, meaning that the cutoff values generated from this study cannot be extrapolated directly to another study sample. However, our findings demonstrated the feasibility of the method, which can be readily adopted in other populations by incorporating ability estimates in future electronic versions of MoCA. In populations where there are insufficient data for initial IRT analyses, our results could serve as reference values as new sample-specific data are being collected. Second, we do not have information external to the test in this sample as a gold standard to test against for diagnostic accuracy. Although the purpose here is to estimate the distribution of cognitive ability in a norm sample, with reference to the standard deviation criteria proposed in Diagnostic and Statistical Manual of Mental Disorders and by Petersen et al., inclusion of clinical and/or neuropathological information in future studies will allow further testing of measurement precision using IRT analysis (American Psychiatric Association, 2013; Petersen et al., 2009). Finally, because of the skewness in education level in our sample, and the potential qualitative difference between older persons with or without formal education, we did not look at marginal effects of increasing levels of education. Continued data collection will allow teasing out the minute effects of education on MoCA test and item information.
This study used only baseline data from a 4-year longitudinal study. Although the longitudinal information was not utilized in this current analysis, follow-up data can be used for future studies once the data become available. By combining the longitudinal information, IRT analysis can be applied to investigate measurement invariance. Cutoff values of IRT ability estimates for people with different educational levels can be refined using the criterion variable created using the follow-up information.
The key concept in IRT is the formulation of the probability, conditional on the latent variable, of obtaining each response category of an item in a questionnaire or test. As clinical researchers try to increase predictive value and clinical utility of MoCA, such as by developing an education-adjusted MoCA memory index score, IRT appears a useful addition to resolve the issue of characterizing true cognitive ability (Julayanont, Brousseau, Chertkow, Phillips, & Nasreddine, 2014).
Supplemental Material
Supplementary – Supplemental material for Applying Item Response Theory Analysis to the Montreal Cognitive Assessment in a Low-Education Older Population
Supplemental material, Supplementary for Applying Item Response Theory Analysis to the Montreal Cognitive Assessment in a Low-Education Older Population by Hao Luo, Björn Andersson, Jennifer Y. M. Tang and Gloria H. Y. Wong in Assessment
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
