Abstract
The Montreal Cognitive Assessment (MoCA) is a widely used cognitive screening tool in stroke. As scoring the visuospatial/executive MoCA items involves subjective judgement, reliability is important. Analyzing data on these items from A Very Early Rehabilitation Trial (AVERT), we compared the original scoring of assessors (n = 102) to blind scoring by a single, independent rater. In a sample of scoresheets from 1,119 participants, we found variable interrater reliability. The match between original assessors and the independent rater was the following: trail-making 97% (κ = 0.94), cube copy 90% (κ = 0.80), clock contour 92% (κ = 0.49), clock numbers 89% (κ = 0.67), and clock hands 72% (κ = 0.46). For all items except clock contour, the independent rater was “stricter” than the original assessors. Discrepancies were typically errors in original scoring, rather than borderline differences in subjective judgement. In trials that include the MoCA, researchers should emphasize scoring rules to assessors and implement independent data checking, especially for clock hands, to maximize accuracy.
Keywords
Introduction
Cognitive impairment and dementia are both prevalent after stroke (Linden, Skoog, Fagerberg, Steen, & Blomstrand, 2004) and have a marked deleterious impact on quality of life (Cumming, Brodtmann, Darby, & Bernhardt, 2014). Yet assessment of cognitive function is often neglected in stroke research: Cognitive outcome measures are included in <5% of stroke studies (Lees, Fearon, Harrison, Broomfield, & Quinn, 2012) and <2% of acute stroke treatment trials (Anderson, Arciniegas, & Filley, 2005). While comprehensive neuropsychological testing can be time consuming, brief cognitive screening tools are available that are valid in stroke populations (e.g., Mini-Mental State Examination [MMSE], Montreal Cognitive Assessment [MoCA]; Cumming, Churilov, Linden, & Bernhardt, 2013).
All cognitive screening tools have a degree of subjectivity in their interpretation and scoring, and therefore reliability is an important consideration. One study demonstrated that a large group of general practitioners trained to deliver and score the MMSE gave significantly higher scores than a single experienced psychologist delivering and scoring the MMSE in the same participants (Fabrigoule, Lechevallier, Crasborn, Dartigues, & Orgogozo, 2003). Reliability can be enhanced by standardizing administration guidelines, having clear scoring instructions, and thoroughly training assessors. The realization that the MMSE, first developed in 1975 (Folstein, Folstein, & McHugh, 1975), was prone to variability in administration and scoring prompted the development of a “standardized” MMSE in 1991. The standardized version included clear instructions on administration and scoring and significantly improved both inter- and intrarater reliability (Molloy, Alemayehu, & Roberts, 1991). The MoCA was developed in 2005 (Nasreddine et al., 2005) and includes well-defined administration and scoring guidelines. It has a greater focus on visuospatial and executive items than the MMSE. The MoCA contains a clock drawing task, which has a strong reputation as a multifaceted and highly informative cognitive task (Shulman, 2000). A categorical rating of the clock drawing task is sensitive to mild Alzheimer’s disease, even in nonexperienced raters (Vyhnalek et al., 2017). Yet despite its clear administration guidelines, problems have been reported with inter- and intrarater reliability in the scoring of the MoCA clock draw task (Price et al., 2011).
Any problems with reliability may be exaggerated in large clinical trials that recruit across many sites and thus have many different outcome assessors. The reliability of scoring on the MoCA clock draw and the other visuospatial/executive items (trail-making and cube copy) in the context of a multicenter trial has not been investigated. A Very Early Rehabilitation Trial (AVERT) was a Phase III randomized controlled trial of earlier (within 24 hours) and more frequent mobilization after stroke (Bernhardt, 2015). AVERT recruitment (n = 2,104) covered 56 different hospital sites in 5 different countries, spanning the years 2006 to 2014. In 2008, the MoCA was added as a 3-month outcome, and in 2011, we reported initial feasibility results (Cumming, Bernhardt, & Linden, 2011). The aim of the current analysis was to use AVERT data to determine the interrater reliability of scoring the visuospatial/executive items from the MoCA.
Method
Study Design
AVERT was a pragmatic, parallel-group, single-blind, multicenter, international randomized controlled trial. The trial had ethical approval from the relevant human ethics committee at each participating site. Eligible participants were aged 18 years or older and were recruited within 24 hours of a confirmed stroke. Exclusion criteria included premorbid disability, early deterioration, palliation, other serious illness or coronary condition, and falling outside set physiological parameters (e.g., blood pressure, heart rate, temperature). Participants were randomly assigned to receive usual stroke unit care alone or very early and more frequent mobilization in addition to usual care, stratified by hospital site and stroke severity. The intervention period lasted 14 days or until discharge from the acute stroke unit, whichever was sooner. The primary outcome was favorable outcome at 3 months poststroke, measured using the modified Rankin Scale (Bonita & Beaglehole, 1988). More details of the study rationale, design, and statistical analysis have been published previously (Bernhardt et al., 2015).
Procedure
Recruitment to AVERT began in 2006. During 2008, the trial management committee approved a protocol revision that added a cognitive screening tool (the MoCA) to the 3-month outcome assessment. Blinded assessors were trained in the administration and scoring of the MoCA. Each assessor was taken through the detailed instructions that are provided by the developers of the MoCA, and always had a copy of these instructions to refer to in their case report form completion manual. In addition, each assessor was given access to a video demonstrating MoCA administration and scoring, recorded by an experienced cognitive neurologist (TL). The blinded assessors were allied health professionals, with the majority being physiotherapists and others being occupational therapists or speech pathologists. The MoCA has been designed so that it can be administered by people without professional neuropsychological qualifications. It takes approximately 10 minutes to complete, is scored out of 30 (with 1 point added for low education), and contains sections on visuospatial/executive, naming, memory, attention, language, abstraction, and orientation. First introduced as a screening tool with high sensitivity and specificity for mild cognitive impairment (Nasreddine et al., 2005), its validity has been demonstrated in stroke populations (Cumming et al., 2013; Pendlebury, Mariz, Bull, Mehta, & Rothwell, 2012). We used the original version of the MoCA but modified its format so responses could be machine-read in our TeleForm (Verity Inc., Sunnyvale, CA) system. It consisted of 2 pages, with the 3 visuospatial/executive items (trail-making, cube copy, and clock drawing) on a page of their own. The pages contained check boxes for each item; blinded assessors at each site marked these boxes to indicate correct or incorrect responses. On completion of the 3-month assessment, as with all other case report forms, MoCA forms were faxed or emailed back to a central repository in Melbourne for checking and processing. If the 3-month assessment was conducted by telephone, the MoCA was not administered. If the participant’s English language fluency led to problems completing the English version of the MoCA—most frequent at the Singapore and Malaysian hospital sites—the MoCA version in their primary language was used and results transcribed on to our case report form. In the case of Bahasa Malay there was no existing MoCA version, so we created and validated one (Sahathevan et al., 2014). If the MoCA was not administered, or was attempted but not completed, reasons for the missing data were recorded. These reasons could be (a) communication problems, (b) unable to use pen/paper, (c) visual problems, (d) telephone follow-up, (e) refused, or (f) other. Of the 2,104 participants recruited to AVERT, 1,189 had complete MoCA data at 3 months (see Results for details on missing data).
MoCA Scoring
Once the incoming MoCA data were checked and uploaded to the database, total score was automatically calculated from the number of marked “correct” boxes. For the visuospatial/executive items, in addition to the checked boxes, we had access to the raw data in the form of a scanned copy of each participant’s attempts at trail-making, cube copy, and clock draw. Five points are available for these items (1 point each for correct trail-making, cube copy, clock contour, clock numbers, clock hands at 10 past 11). As part of our quality assurance process, we extracted all available raw data on these items to evaluate the reliability of the original scoring. A single experienced rater (DL), who was blind to the original scoring, scored each of the 3 visuospatial/executive items independently. The first step of our analysis was to compare the original scoring (as sent in from the hospital sites) to the secondary blind scoring (from single independent rater DL). We also wanted to probe whether any discrepancies between original and secondary scoring were due to clear errors or to borderline differences in subjective judgement. To do this, once the secondary scoring was complete, all discrepancies were inspected unblinded to the original score by the independent rater (DL) and another rater (TC). At this step, we made a consensus decision about whether the discrepancy could have resulted from a difference in subjective judgement (e.g., clock hands of similar length being marked correct, despite the scoring rules requiring that the minute hand must be clearly longer) or a clear error (e.g., clock hands pointing to the wrong time being marked correct). If it was the former, original scoring was retained (and thus the item was no longer discrepant). If it was the latter, score was changed to that of the independent rater (and the item remained discrepant). Once this was done, discrepancies were recalculated.
Statistical Analysis
Discrepancies between the original assessors and the independent rater were reported using descriptive statistics. Kappa values with 95% confidence intervals were calculated to quantify interrater reliability. Levels of agreement were classified as “none to slight” (0.01-0.20), “fair” (0.21-0.40), “moderate” (0.41-0.60), “substantial” (0.61-0.80), or “almost perfect” (0.81-1.00; Landis & Koch, 1977). All analyses were performed using SPSS version 20.
Results
A total of 2,104 participants were recruited to AVERT. Of these, 317 participated prior to the 2008 introduction of the MoCA, 136 had died prior to 3-month assessment, 6 were lost to follow-up, and 456 had partially or completely missing data on the MoCA. Of the 456 with missing MoCA data, 76 had no reason recorded. Of the other 380, 143 (38%) had communication problems, 66 (17%) were unable to use pen/paper, 18 (5%) had visual problems, 157 (41%) had telephone follow-up, 51 (13%) refused, and 43 (11%) were listed as “other.” Of the remaining 1,189 with complete MoCA data, we were able to review the raw data on the visuospatial/executive items from 1,119 and blind score them (the other 70 files were unavailable due to a retrieval fault in the TeleForm system). These 1,119 case report forms were originally administered and scored by 102 different assessors across 54 hospital sites. The English version of the MoCA was used in 52 out of 54 sites, while the Singapore (n = 73) and Malaysian (n = 47) sites used the relevant language version (predominantly Chinese or English).
The match between original and secondary scoring ranged from 97% (on the trail-making item) to 72% (on the clock hands item; see Table 1). Secondary scoring was stricter for all items except clock contour. Interrater reliability was the following: for trail-making, κ = 0.94 (95% confidence interval [CI] 0.92-0.96); for cube copy, κ = 0.80 (95% CI 0.76-0.83); for clock contour, κ = 0.49 (95% CI 0.40-0.58); for clock numbers, κ = 0.67 (95% CI 0.62-0.73); and for clock hands, κ = 0.46 (95% CI 0.42-0.51). Following consensus review, discrepancies between raters were reduced (see Table 1). Discrepancies were attributable to borderline subjective judgements—rather than original scoring errors—in 12 out of 36 cases for trail-making, 30 out of 112 for cube copy, 11 out of 90 for clock contour, 28 out of 127 for clock numbers, and 63 out of 316 for clock hands. Once these adjustments had been made (with the “benefit of the doubt” given to the original scorer in borderline cases), the match between original and secondary scoring ranged from 98% (on the trail-making item) to 77% (on the clock hands item). Data are shown in Table 1.
Discrepancies Between Original and Secondary Raters on Each Visuospatial/Executive Item, Before and After Unblinded Consensus Review.
Note. Orig0_New1 = original score incorrect, secondary score correct; Orig1_New0 = original score correct, secondary score incorrect.
An example of a discrepancy is shown in Figure 1. For this scoresheet, original scoring had all 5 items as correct. In secondary blind scoring, both cube copy and clock hands were scored incorrect. In unblinded consensus scoring, cube copy was scored incorrect (violates the requirement for lines to be parallel) but the original scoring of clock hands as correct was retained (hands not clearly different lengths, but close enough to be borderline). Figures 2 and 3 provide additional examples of the clock drawing and cube copy items, with reference to the scoring guidelines.

Example visuospatial/executive scoresheet.

Example clocks.

Example cubes.
To evaluate differences in scoring accuracy between sites, we summed the number of discrepancies for individual participants across the 5 items and calculated a mean number of discrepancies per participant for each site. Of the 54 sites, 30 sites (accounting for 80% of the participants) had a mean number of discrepancies between 0.34 and 0.82 out of 5. At the upper end of the distribution, there were 4 outliers with high mean discrepancies (between 1.50 and 1.71 out of 5), and these were all smaller sites (<10 participants). In terms of geographical region, Australia and New Zealand (24 sites, 575 participants, mean = 0.49, SD = 0.71) had fewer discrepancies than the United Kingdom (28 sites, 405 participants, mean = 0.73, SD = 0.78) and the 2 Asian sites in Singapore and Malaysia (120 participants, mean = 0.82, SD = 0.90).
Discussion
In the AVERT trial, scoring reliability of the visuospatial/executive MoCA items was variable. Percentage match between the original scoring (across 102 different assessors) and a single blinded independent rater was 97% for trail-making, approximately 90% for cube copy, clock contour, and clock numbers, but only 72% for clock hands. Most of the discrepancy for the clock hands item was due to it being originally scored as correct in cases of it being clearly incorrect. After accounting for the discrepancies between original and secondary scoring that could be considered borderline subjective judgements, the percentage match for clock hands remained relatively low (77%). In terms of interrater reliability between original and blind scoring, kappa agreement levels ranged from “almost perfect” for trail-making (0.94) to “substantial” for cube copy (0.80) and clock numbers (0.67) to “moderate” for clock contour (0.49) and clock hands (0.46).
Some have argued that the MoCA criteria for scoring the clock draw item leave greater scope for subjective interpretation than other scoring systems, such as the Cosentino criteria (Price et al., 2011). This may be true, but our findings suggest that the major interrater reliability issue is incorrect application of the scoring guidelines, not differences in subjective judgement. In the MoCA scoring criteria for clock hands, it is specified that the minute hand must be clearly longer than the hour hand for the item to be scored correct. We noted many cases of hands pointing to the 10 and the 2 (i.e., potentially correctly placed) that were scored correct by the original assessors even though the hands were of similar length. While it may seem appropriate to be generous and give a point in these cases, proper application of the scoring criteria mean they must be scored incorrect. Another clock hands scoring criterion—that the source point of the hands must be centrally located—was also occasionally overlooked by original assessors, leading to additional false positives. Data from our blinded assessors, many with limited MoCA experience, suggest that additional training may be required in the scoring of the clock hands item.
Clock contour, while not featuring as many discrepancies as clock hands, also had only “moderate” agreement between raters. This was due to a tendency for original assessors to score the item incorrect when it was actually correct. According to the contour scoring criteria, minor distortion is acceptable and no specifications are made relating to size. There were not many borderline cases, with match between raters only increasing one percentage point (91.9% to 92.9%) after consensus. For clock numbers, the discrepancies were more equally balanced and there were more borderline cases (the match increased from 88.5% to 91.0% after consensus). In comparison with the clock draw items, higher agreement was found for the trail-making and cube copy items, which is probably attributable to their simpler and more self-evident scoring criteria.
With the exception of clock contour, all 5 items were scored more “strictly” by the independent rater than by the original assessors. This may reflect a tendency for clinicians (the AVERT blinded assessors had backgrounds in allied health) to score people’s MoCA performance according to potential or assumed capacity rather than according to the specific (but somewhat arbitrary) scoring rules. If this is the case, studies using clinicians to administer the MoCA may typically underestimate cognitive impairment in their samples. Clear scoring rules are important; the reliability of psychometric tools such as the MoCA is a product of their consistent administration and rigorous application of scoring criteria. The AVERT trial included face-to-face training of assessors, a detailed case report form completion manual, and a custom-made video tutorial for administering and scoring the MoCA. Our findings, however, indicate that there is still room for improvement in scoring reliability for the visuospatial/executive items.
Several limitations of the current study should be noted. We did not have audio or video recordings of the assessment session, so the blind scoring of the independent rater was limited to what appeared on the case report form. This can be important; for example, we noticed that several consecutive clock draw items from a particular site with the hands set to 3 o’clock were marked correct. We contacted the site, and these participants had in fact been asked to set the time to 3 o’clock instead of 10 past 11, so the items were scored correct. Different language versions of the MoCA were sometimes used at the Singapore and Malaysian sites, and this may have negatively influenced test–retest reliability. Relative to Australia/New Zealand, however, the Asian and U.K. sites both had a higher number of discrepancies, indicating that language cannot be the sole explanation. While our data set was very large, the sample may not be fully representative, with 456 AVERT participants having partially or completely missing MoCA data. Of the 1,189 participants with complete MoCA data, there were 70 for whom we could not extract scoresheets.
We recommend that in trials including the MoCA, researchers should emphasize the scoring rules to assessors. We think that, in addition to the clearly written MoCA scoring criteria, visual examples of correct and incorrect items (similar to those contained in the standardized MMSE; Molloy et al., 1991) would increase reliability. For research trials, it may be necessary to establish criteria for demonstrating competence in MoCA scoring, in the same way that other clinical scales have certification procedures. Implementing independent data checking, especially for the clock hands item, will further increase accuracy.
Footnotes
Acknowledgements
We thank the AVERT Collaboration investigators for all their hard work and dedication. This study was made possible by the support of many people not specifically named here, particularly the blinded assessors at each hospital site.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This project had no specific funding support. AVERT was initially supported by the National Health and Medical Research Council of Australia (JB, Grant Numbers 386201, 1041401), with additional funding from Chest Heart and Stroke Scotland (JB, Res08/A114); Northern Ireland Chest Heart and Stroke; Singapore Health (JB, SHF/FG401P/2008); the UK Stroke Association (JB, TSA2009/09); and the UK National Institute of Health Research (JB, HTA Project 12/01/16). The Florey Institute of Neuroscience and Mental Health acknowledges the support received from the Victorian Government via the Operational Infrastructure Support Scheme.
