Abstract
The Trail Making Test (TMT) is used as an indicator of visual scanning, graphomotor speed, and executive function. The aim of this study was to examine the TMT relationships with several neuropsychological measures and to provide normative data in community-dwelling participants of 55 years and older. A population-based Spanish-speaking sample of 2,564 participants was used. The TMT, Symbol Digit Test, Stroop Color–Word Test, Digit Span Test, Verbal Fluency tests, and the MacQuarrie Test for Mechanical Ability tapping subtest were administered. Exploratory factor analyses and regression lineal models were used. Normative data for the TMT scores were obtained. A total of 1,923 participants (76.3%) participated, 52.4% were women, and the mean age was 66.5 years (Digit Span = 8.0). The Symbol Digit Test, MacQuarrie Test for Mechanical Ability tapping subtest, Stroop Color–Word Test, and Digit Span Test scores were associated in the performance of most TMT scores, but the contribution of each measure was different depending on the TMT score. Normative tables according to significant factors such as age, education level, and sex were created. Measures of visual scanning, graphomotor speed, and visuomotor processing speed were more related to the performance of the TMT-A score, while working memory and inhibition control were mainly associated with the TMT-B and derived TMT scores.
Keywords
The Trail Making Test (TMT) was developed by Partington and Leiter in 1938 as a divided attention test, and was originally part of the Army Individual Test Battery (Partington & Leiter, 1949) used by the U.S. Army (Army Individual Tests Battery, 1944). It was later included in the Halstead-Reitan Neuropsychological Battery (Reitan, 1955a). The TMT consists of two parts (TMT-A and TMT-B). In TMT-A, the respondent is asked to connect randomly arranged circles containing numbers from 1 to 25 following the number sequence, and doing it as quickly as possible. The task in TMT-B is similar to TMT-A, but the respondent has to alternate between numbers and letters. Apart from TMT-A and TMT-B direct scores, which is the time used to complete each part, other derived scores are used, such as the difference score (TMT-B − TMT-A) and the ratio score (TMT-A/TMT-B). Neuroanatomically, TMT performance activates the right inferior medial frontal cortex, as well as other nonfrontal brain regions, including the left precentral, angular, and medial temporal gyrus, and the intraparietal sulcus (Jacobson, Blanchard, Connolly, Cannon, & Garavan, 2011). Overall, there is evidence that multiple cognitive functions are involved in the TMT execution, and that it requires the activation of different brain regions (Oosterman et al., 2010; Terada et al., 2013).
From a clinical point of view, the TMT is widely used as an indicator of brain dysfunction (Armitage 1946; Larrabee, Millis, & Meyers, 2008; Reitan, 1955b; Reitan, 1959, Reitan & Wolfson, 1993; Strauss, Sherman, & Spreen, 2006). It is used in multiple circumstances, such as brain injury (Heled, Hoofien, Margalit, Natovich, & Agranov, 2012; Lange, Iverson, Zakrzewski, Ethel-King, & Franzen, 2005) or diverse neurodegenerative diseases, including Alzheimer’s disease (Ashendorf et al., 2008), dementia with Lewy bodies (Ferman et al., 2006), or Huntington’s disease (Lemiere, Decruyyenaere, Evers-Kiebooms, Vandenbussche, & Dom, 2004; O’Rourke et al., 2011). Moreover, the TMT has been seen to have ecological validity and distinguishes healthy elders from neurological patients (Bell-McGinty, Podell, Franzen, Baird, & Williams, 2002; Dawson, Anderson, Uc, Dastrup, & Rizzo, 2009; Emerson et al., 2012). Longitudinal studies have shown that the TMT is capable of forecasting clinical and functional changes in patients with Alzheimer’s disease (Chen et al., 2001), depression (Potter et al., 2013), mild cognitive impairment (Chapman et al., 2011; Ewers et al., 2012; Gomar, Bobes-Bascaran, Conejero-Goldberg, Davies, & Goldberg, 2011), brain injury (Hanks et al., 2008), and stroke (Wiberg, Kilander, Sundström, Byberg, & Lind, 2012). Moreover, in older adults, TMT performance has been related to physical decline and higher risk of mortality (Vazzana et al., 2010).
Originally, cutoffs were used to distinguish healthy participants from those with a cognitive impairment (Matarazzo, Wiens, Matarazzo, & Goldstein, 1974; Reitan, 1959; Reitan & Wolfson, 1993), but the effect of sociodemographic variables required the use of normative data stratified by age and education level (Horton & Roberts, 2003; Spreen & Strauss, 1998). The effect of age and education on the TMT performance is widely known. Thus, the time required to perform the test increases with age (Cavaco et al., 2013; Ernst, Warner, Townes, Peel, & Preston, 1987; Giovagnoli et al., 1996; Ivnik, Malec, Smith, Tanglos, & Petersen, 1996; Mitrushina, Boone, Razani, & D’Elia, 2005; Rasmusson, Zonderman, Kawas, & Resnick, 1998; Salthouse et al., 2000; Stuss, Stethem, & Poirier, 1987; Tombaugh, 2004; Wecker, Kramer, Wisniewski, Delis, & Kaplan, 2000) and with low education levels (Bornstein, 1985; Bornstein & Suga, 1988; Ernst et al., 1987; Giovagnoli et al., 1996; Ivnik et al., 1996; Seo et al., 2006; Stuss, Stethem, Hugenholtz, & Richard, 1989; Tombaugh, 2004). The effect of sex on TMT performance is controversial, and while some studies report no effect (Ivnik et al., 1996; Lucas et al., 2005; Mitrushina, Boone, Razani, & D’Elia, 2005; Tombaugh, 2004), others report differences only in TMT-A performance (Cavaco et al., 2013; McCurry et al., 2001; Seo et al., 2006), and others only on TMT-B performance (Bornstein, 1985; Cavaco et al., 2013; Ernst et al., 1987; Wiederholt et al., 1993).
TMT normative data are available for several countries and cultures (Bezdicek et al., 2011; Cangoz, Karakoc & Selekler, 2009; Cavaco et al., 2013; Hashimoto, 2006; Hester, Kinsella, Ong, & McGregor, 2005; Seo et al., 2006; Tombaugh, 2003), including Spain (Del Ser et al., 2004; Peña-Casanova, Quiñones-Ubeda, et al., 2009; Periáñez et al., 2007). Published normative data differ depending on the type of sample used as a reference (Ashendorf et al., 2008; Giovagnoli et al., 1996; Lezak, Howieson, & Loring, 2004; Mitrushina, Boone, Razani, & D’Elia, 2005; Periáñez et al., 2007; Soukup, Ingram, Grady, & Schiess, 1998; Strauss et al., 2006), and therefore, using the appropriate normative data is crucial to avoid diagnostic errors (Fernández & Marcopulos, 2008).
Different cognitive abilities underlie in the execution of the two parts of the TMT. An important number of studies have investigated this by examining the association between the TMT and other neuropsychological or cognitive measures (reviewed by Sánchez-Cubillo et al., 2009). Most of these studies suggest that graphomotor speed and visual scanning play an important role in the completion of both TMT-A and TMT-B, while executive function components—such as working memory, inhibition control, or set-switching abilities—are more specifically involved in TMT-B performance. However, the existing findings differ in the relative contribution that these factors make to the direct and derived scores. Most of the studies had limitations, such as relatively small samples and different types of samples, and included patients and healthy volunteers (mostly students or older adults selected with nonprobabilistic sampling procedures). Moreover, age has an important effect in the cognitive abilities involved in TMT performance (e.g., Salthouse, 2011), and the results of studies conducted with young participants may not be applicable to elder populations. In addition, previous studies have shown that clinical and healthy groups differ in the correlations between TMT scores and other cognitive measures. For example, Misdraji and Gass (2010) found a greater contribution of graphomotor speed to TMT performance in impaired than in nonimpaired participants, thus decreasing the influence of other factors such as visual scanning or mental shifting efficiency. The low statistical power caused by sample sizes and the limited external validity of convenience samples used in most of the studies have contributed to produce divergent results.
Taking into account the wide use of the TMT in clinical practice, the aim of this study was to explore the relationships of the direct and derived scores of the TMT with other related neuropsychological measures and to provide normative data based in a large population-based sample of individuals aged 55 years and older.
Method
Study Population
The participants come from the Regicor study, which is a prospective and population cohorts study about cardiovascular risk factors (http://www.regicor.org/). Specifically, it includes three population-based cohorts of inhabitants of the province of Girona (Catalonia, Spain) who were recruited in the years 1995, 2000, and 2005 (Grau et al., 2007). The reference population included all the inhabitants in the province of Girona aged 35 to 74 years in the year they were recruited, and they were selected from the census. The rate of participation of the three cohorts was of 72.4%, 70.0%, and 73.8%, respectively. Between the years 2008 and 2013, the participants of the three cohorts were contacted once again to carry out the clinical follow-up. For the present study, we selected all the participants who accepted to participate in a complementary protocol about cognitive functions in individuals aged 55 years and older. The primary exclusion criterion was the presence of visual, hearing, or motor deficit that may compromise the correct execution of the neuropsychological tests. In order to remove individuals with suspected cognitive impairment, we excluded the individuals requiring more than 300 seconds to complete the TMT-B form. The study was reviewed and approved by the Institutional Review Board of the Girona Healthcare Institute, and the data included in this article were obtained in compliance with the Helsinki Declaration. The informed consent was signed by all the study participants.
Procedure and Instruments
The data collection process of the Regicor study was performed in the health care centers of each participating municipality. There, the examining team received the participants between 8:00 a.m. and 11:00 a.m., after they had been informed by postal mail and they received a telephone call. The examination was standardized and all the participants were examined as follows: drawing of blood to start, an electrocardiogram, a carotid echo-Doppler, measurement of weight and height and, finally, a clinical interview and completion of several health questionnaires. Several self-reported and objective cardiovascular health data were collected: hypertension, diabetes mellitus, myocardial infarction, angina, obesity, depressed mood based on Patient Health Questionnaire–9 (Kroenke, Spitzer, & Williams, 2001; Manea, Gilbody, & McMillan, 2015), body mass index, supine and diastolic blood pressure and high density lipoprotein, low density lipoprotein, and triglycerides). After a small pause, the neuropsychological examination was carried out by examiners who were previously trained. All tests were administered in the same order to decrease test effects during the neuropsychological assessment.
Trail Making Test
The TMT materials consisted of two parts (TMT-A and TMT-B). The TMT-A consisted of a standardized page on which the numbers 1 to 25 are scattered within circles, and the participants were asked to connect the numbers in order as quickly as possible. Similarly, the TMT-B consisted of a standardized page that included the numbers 1 to 13 and the letters A to L. The participants were instructed to draw lines connecting numbers and letters in order, alternating numbers and letters. Before starting the test, a practice trial of six items was administered to the participants to make sure that they understood both tasks. When a participant made an error during the test performance, the examiner pointed it out and explained it, then guided the participant to the last circle completed correctly, and requested to continue with the task. A maximum time of 300 seconds was allowed before discontinuing the test. Direct scores of TMT were the time in seconds taken to complete each task (A and B). The TMT difference (TMT-d = TMT-B − TMT-A) and the TMT ratio (TMT-r = TMT-B/TMT-A) scores were also calculated.
Symbol Digit Test (SDT)
The SDT (Spanish version; Wechsler, 1999) was used to assess visual search, and perceptual and graphomotor speed. A coding key showed nine abstract symbols, each paired with a number. Below the key, series of symbols were presented, and the participants were asked to write down the corresponding numbers, as quickly as possible. The number of correct substitutions during a 90-second interval was used as the score.
Stroop Color–Word Test (SCWT)
The SCWT (Spanish version; Golden, 1994) was used to evaluate inhibition control. A page with 100 color words (blue, green, and red) printed in nonmatching colors was shown to the participants, who were asked to name the colors disregarding the verbal content of the words. The number of words from which the participants named the color during a 45-second interval was used as the score.
Digit Span Test (DST)
The DST (Spanish version; Wechsler, 1999) was used to evaluate working memory. The participants were required to repeat series of digits that became gradually longer. The maximum digit span that the participants were able to repeat in direct and reverse orders constituted the forward (DST-f) and backward (DST-b) scores, respectively.
Verbal Fluency
The participants were asked to say words beginning with the letter “P” (phonemic verbal fluency—pVF), and to say words pertaining to the “animals” category (semantic verbal fluency—sVF). They had to do it as quickly as possible, and the number of words produced during 1 minute was scored for both phonemic and semantic verbal fluency.
MacQuarrie Test for Mechanical Ability–Tapping Subtest (t-MQTMA)
The t-MQTMA (MacQuarrie, 1925) is a specific task to measure hand–eye coordination and muscular control and was used to assess visuomotor processing speed. The participants were asked to put three pencil dots in each of a number of circles as fast as possible. The number of dots during 30 seconds was used as the score.
Statistical Analysis
A descriptive analysis of the demographic and health characteristics of the participants and the neuropsychological examination was carried out by means of absolute and relative frequencies for qualitative variables and by means of central tendency and dispersion measures for quantitative variables. A bivariate analysis of the association of each neuropsychological test was performed using the Pearson product-moment correlation coefficient.
The assessment of the association between the TMT scores and the other neuropsychological measures was done using exploratory factor analyses (EFAs) and multiple regression analyses. Four EFAs were used to explore the association of each TMT score with the other tests. The extraction method used was the principal component analysis. Prior to performing the EFAs, the conditions of application were made by calculating Bartlett’s test of sphericity and the Kaiser–Meyer–Olkin measure of sampling adequacy. Factors were selected if their eigenvalue was greater than 1, and tests were included in a factor if its factor loading was ≥0.30. Four regression models (enter method) were computed to assess the construct validity of the TMT. These models included the TMT-A, the TMT-B, the TMT difference, and the TMT ratio scores as dependent variables, and the scores of the SDT, SCWT, DST-f, DST-b, pVF, sVF, and t-MQTMA were entered as independent variables. To determine the effect of age, sex, and educational level in the execution of the TMT scores, four regression models (backward stepwise method) for each of the TMT scores were conducted. The contribution of each independent variable to the total variance explained by each regression model was calculated using the beta coefficient product by Pearson’s correlation coefficient for each independent variable with the dependent variable (Guilford & Fruchter, 1973). To avoid multicollinearity issues, the variable age (67 years) was centered and the quadratic term of the variable age was then calculated, and the possible interactions were analyzed. The variable sex was coded (male = 0, female = 1) and the variable education level was recorded according to the years of education of the participants, up to 8 years of education (LE [level of education] low), between 9 and 12 years of education (LE average) and more than 12 years of education (LE high). Two dummy variables were created (LE low and LE high) using LE average as a reference category to include in regression models. For each adjusted model, the necessary requirements for applying a regression analysis were determined: homoscedasticity, normal distribution of residuals, absence of multicollinearity, and absence of extreme values. Homoscedasticity was determined by means of Levene’s test of variance homogeneity for the residuals of the score predicted by the stratified regression model according to the grouping per quartiles of the predicted score. The normal distribution of the residuals was determined by means of the Kolmogorov–Smirnov test and the visual inspection of the histogram. The presence of multicollinearity was evaluated by means of the variance inflation factors which must not be above 10 (Belsley, Kuh, & Welsch, 1980). Cook’s distance was calculated to identify possible influential cases defined from values equal to or above 1 (Cook & Weisberg, 1982).
To develop normative data, two different strategies were applied. The first used the traditional method based on the score stratification according to the variables that presented with a significant effect on the execution of the TMT (Mitrushina, Boone, Razani, & D’Elia, 2005) and the second one based on regression models (Van Breukelen & Vlaeyen, 2005). According to the traditional method, the normative values were calculated for each score of the TMT (5, 10 to 90, and 95 percentiles and the associated scaled scores). In order to maximize the number of participants who contribute to the normative distribution for each mean value of the age interval, the strategy of interval superposition was adopted (Pauker, 1988). Hence, each mean age value provided norms for participants of that age plus or minus 1 year, except for cases fewer than 61 years and above 82 years of age. The age range for each mean value was of 10 years, so that the mean age value of 60 years included the interval of participants between 55 and 65 years of age, the mean age value of 63 years included the interval between 58 and 68 years, and so forth. The age distribution allowed for the calculation of normative data for the following nine groups of mean age values: 55 to 61, 62 to 64, 65 to 67, 68 to 70, 71 to 73, 74 to 76, 77 to 79, 80 to 82, and 82 and above years.
According to the method based on regression models, raw scores in the TMT of one individual become the standardized residual values, following three steps. First, the predicted values by the regression model are calculated; second, the residual values are calculated (ei = score observed − predicted score); and third, the residual values are standardized (Zi = ei/SD [residual]). Standardized residual values are interpreted according to a normal distribution table. Data processing and analysis was conducted using PASW Statistics version 19 (SPSS; Chicago, IL, USA) for Windows. All the analyses were bilateral, and a p value of less than .05 was considered to be statistical significant.
Results
The sample consisted of 2,564 individuals of 55 years of age and above, and the final rate of participation was of 76.7% of the candidates (Figure 1). The mean age of the participants was of 66.4 years (SD = 8.0), and 52.40% were women. Regarding the education level, 51.9% had a low educational level, 28.0% had an average level of education, and 21.1% had a high level. In general, individuals who did not participate (rejected and discontinued) were older and had a lower level of education than those who agreed to participate and completed the TMT (71.6 years [SD = 8.58] vs. 66.4 [SD = 8.0], Z for Mann–Whitney U = −10.4; p < .001; 72.2% vs. 51.9% low level of education, χ2 = 50.3; df = 2; p < .001).

Study participation flowchart.
Table 1 shows the demographic and health characteristics of the study participants stratified by age groups and the level of education.
Demographic and Health Characteristics of the Study Participants Stratified by Age Groups and Level of Education.
Note. BMI = body mass index; SBP = supine blood pressure; DBP = diastolic blood pressure; HDL = high-density lipoprotein; LDL = low-density lipoprotein.
Table 2 shows the descriptive statistics for the TMT scores and the neuropsychological measures included in the analysis of TMT construct validity. The descriptive analysis of direct and derived TMT scores showed skewed distributions. Thus, all scores were logarithmically transformed before performing further statistical analyses.
Descriptive Statistics of the Neuropsychological Tests Included in the Analyses of TMT Construct Validity.
Note. TMT-A = Trail Making Test–A; TMT-B = Trail Making Test–B; TMT-d = Trail Making Test difference; TMT-r = Trail Making Test ratio; SDT = Symbol Digit Test; SCWT = Stroop Color–Word Test; DST-f = Forward Digit Span Test; DST-b = Backward Digit Span Test; pVF = Phonemic Verbal Fluency; sVF = Semantic Verbal Fluency; t-MQTMA = MacQuarrie Test for Mechanical Ability1–Tapping subtest.
The bivariate analysis of the association between the neuropsychological tests scores showed coefficients that indicated significant and moderate to large associations between tests scores (Table 3).
Bivariate Correlation Matrix (Spearman’s Rank Correlation Coefficient).
Note. TMT-A = Trail Making Test–A; TMT-B = Trail Making Test–B; TMT-d = Trail Making Test difference; TMT-r = Trail Making Test ratio; SDT = Symbol Digit Test; SCWT = Stroop Color–Word Test; DST-f = Forward Digit Span Test; DST-b = Backward Digit Span Test; pVF = Phonemic Verbal Fluency; sVF = Semantic Verbal Fluency; FFT = finger tapping test; t-MQTMA = MacQuarrie Test for Mechanical Ability–Tapping subtest.
The EFAs were performed on the direct and derived scores of the TMT and the tests included in the neuropsychological examination. The application indices of EFA for each TMT score were positive and are shown in Table 4. In order to interpret the results, a Promax oblique rotation, which allows for a correlation to be made between the factors, and the Kaiser normalization were performed. Table 4 shows the number of factors extracted in each EFA, and the proportion of variance explained for each one.
Factor Loadings (Principal Component Analysis Extraction Method).
Note. TMT-A = Trail Making Test–A; TMT-B = Trail Making Test–B; TMT-d = Trail Making Test difference; TMT-r = Trail Making Test ratio; SDT = Symbol Digit Test; SCWT = Stroop Color–Word Test; DST-f = Forward Digit Span Test; DST-b = Backward Digit Span Test; pVF = Phonemic Verbal Fluency; sVF = Semantic Verbal Fluency; t-MQTMA = MacQuarrie Test for Mechanical Ability–Tapping subtest.
p Value. High factor loadings groups are highlighted in bold.
All the necessary requirements to apply the regression analyses were met on the direct and derived log scores of the TMT. The Levene test for the residuals of each quartile TMT scores did not reject the assumption of homogeneity of the variances and the result of the application of the Kolmogorov–Smirnov test showed p values greater than .135 for residuals of each TMT score. Variance inflation factor values were not above 2.869 in any of the regression models. The presence of extreme values was not seen, as the highest Cook’s distance was of 0.037. Table 5 shows the results of the regression models that included the demographic and neuropsychological predictors. The coefficient of determination ranged between 0.616 (TMT-B) and 0.077 (TMT-r). The SDT was the measure that explained the greatest amount of variance of the TMT-A, the TMT-B, and the TMT-d scores. The DST-b was also associated with the TMT-B and the derived scores, and the SCWT also contributed to the TMT-B, the TMT-d, and the TMT-r scores.
Standardized Regression Coefficients (Sdβ) and Variance Explained by Variable (ΔR2) of the Neuropsychological Measures on log Transformed TMT Scores.
Note. Sdβ = standardized regression coefficient; ΔR2 = variance explained by variable (Pearson’s correlation * Sdβ); TMT-A = Trail Making Test–A; TMT-B = Trail Making Test–B; TMT-d = Trail Making Test difference; TMT-r = Trail Making Test ratio; SDT = Symbol Digit Test; SCWT = Stroop Color–Word Test; DST-f = Forward Digit Span Test; DST-b = Backward Digit Span Test; pVF = Phonemic Verbal Fluency; sVF = Semantic Verbal Fluency; t-MQTMA = MacQuarrie Test for Mechanical Ability–Tapping subtest.
Coefficient of determination (R2 = regression models adjusted for age, sex, and education level).
p < .05.
Table 6 presents the results of regression models for the different TMT scores including the demographic variables. Age, sex, and educational level were associated to the scores of TMT-A, TMT-B, and TMT-d. The TMT-r score was only related to the educational level. Age was the variable that had a greatest variance percentage explained in TMT-A regression model, while the level of education was the variable with the greatest variance percentage explained in TMT-B, TMT-d, and in TMT-r regression models. None of the TMT scores was affected by the quadratic term of age, and no significant interactions between the educational level and the age were detected for any TMT score.
Stepwise Multiple Linear Regression of Sex, Age, and Education on Log Transformed TMT Scores.
Note. TMT-A = Trail Making Test–A; TMT-B = Trail Making Test–B; TMT-d = Trail Making Test difference; TMT-r = Trail Making Test ratio; LE = level of education; β = regression coefficient; SE(β) = standard error of β; Sdβ = standardized regression coefficient; ΔR2 = variance explained by variable (Pearson’s correlation * Sdβ).
R2 = coefficient of determination.
p < .05.
Using data of regression models to obtain normative values requires the users to carry out some calculations. The normative data are obtained from the values predicted by the regression model in combination with the standard deviation of residuals. First, the score predicted by the regression model must be calculated based on the characteristics of age, educational level, and sex by means of the regression coefficients presented in Table 6. Second, the residual value must be calculated, which corresponds to the difference between the predicted value and the obtained value (ei = obtained value − predicted value). Third, it is possible to transform the standard deviations of the residual values (Table 7) into Z values (Zi = ei/SD [residual]). For example, for a 75-year-old man with a high educational level, the predicted Log TMT-A score would be of 3.850 (3,758 + [0,051 * 0] + [8 * 0,017] + [−0,044 * 1] + [0,226 * 0]). If the obtained TMT-A score was 61 seconds (Log TMT-A = 4.110), the residual value would be of −0.261 (ei = 4.110 − 3.850) and the standardized residual value would be −1.265 (= −0.261/0.206) which corresponds to a p value of .102 and to the 10th percentile.
Standard Deviations of Residuals for the Predicted Log TMT scores.
Note. TMT-A = Trail Making Test–A; TMT-B = Trail Making Test–B; TMT-d = Trail Making Test difference; TMT-r = Trail Making Test ratio.
Supplemental material (available online at http://journals.sagepub.com/doi/suppl/10.1177/1073191115602552) includes the normative data of the TMT direct and derived scores. Tables S1 to S4 provide normative data obtained by the traditional method. The 5, 10 to 90, and 95 percentiles, scaled and T scores are stratified in accordance with the relevant variables detected in the regression models. Each table comprises age intervals superimposed with mean age values in 3-year intervals (age ± 1 year). To be used, the age group with the mean age value closest to the age of the participant for whom the score is to be interpreted must be chosen. The scores included in the tables allow positioning the performance of the participants compared with their normative group based on percentiles. Tables S5 to S8 present normative data from the method based on regression models for 5-year groups to make their use easier. For their clinical use, if the person that is to be evaluated does not have the exact age used for the creation of the table (55, 60, 65, . . . ), it must be rounded to the closest age of reference. The supplemental material also includes an Excel file (TMT scoring file.xls) to calculate the percentiles and the standard deviations of the direct and derived TMT scores obtained based on regression models according to age, sex, and level of education.
Discussion
This study aimed to evaluate the relationships of the TMT direct and derived scores with measures of visual search, graphomotor speed, visuomotor processing speed, inhibition control, working memory, and verbal fluency, and to provide normative data for participants aged 55 years and older based on a large population-based sample.
The TMT has been hypothesized to reflect a variety of cognitive processes that fit into the executive functioning construct. However, this construct is an umbrella that consists of several subcomponents, and there is no consensus regarding which and how many subcomponents there are and how they should be measured (Jurado & Rosselli, 2007; Koziol, 2014). We used a double approach to investigate how the TMT direct and derived scores were associated with a set of measures related with the executive functioning. The EFAs allowed us to test if each TMT score load on a single factor that included all the measures of executive functioning or if they were mediated by different cognitive processes and loaded into different factors. Complementarily, the multiple linear regressions allowed us to quantify which tests were more involved in the variability of each TMT measure, reporting a potential task-specific subcomponent relation between these measures.
The TMT-A loaded into a factor that included all the tests except the DST scores (forward and backward) that were the only nontime-dependent measures. In addition, the 42.4% of the variance of the TMT-A score was explained by visual search and graphomotor speed as assessed by the SDT (34.5%), and by visuomotor processing speed as assessed by the t-MQTMA (7.9%). Taken together, these findings agree with previous studies that support the TMT-A score as a measure of attention and speed that requires visual scanning, graphomotor speed, and visuomotor processing speed as main cognitive functions (Ehrenstein, Heister, & Cohen, 1982; O’Rourke et al., 2011; Sánchez-Cubillo et al., 2009). The inclusion of a measure of hand–eye coordination and muscular control such as the t-MQTMA instead of a pure measure of motor function such as the Finger Tapping Test allowed us to explain a higher percentage of variance of the TMT-A. For example, in the study by Sánchez-Cubillo et al. (2009), the Finger Tapping Test did not reach statistical significance, and visual scanning and graphomotor speed were the unique predictors, explaining 11.2% of the TMT-A variance. Instead, our results indicate that the visuomotor processing speed is responsible of the 7.9% of the overall variance of the TMT-A.
The TMT-B loaded into a single factor that included all the tests, and the multiple regression analysis predicted the 61.6% of the TMT-B score through the SDT score (31.5%), the DST-b (8.3%), the SCWT score (5.1%) and the t-MQTMA score (4.0%) as main predictors after adjusting for age, sex, and education level. Thus, according to our results, the capacity to connect by pencil lines numbers and letters in alternating order requires mainly visual scanning and graphomotor speed in order to find and connect the numbers and letters randomly arranged in the page, and in a minor degree it requires working memory (ability to manipulate information as measured by the DST-b) to remember the correct number and letter in the sequence, inhibition of habitual response when it comes to alternate the sequence between numbers and letters (as measured by the SCWT), and visuomotor processing speed in order to execute the task as fast as possible. Our results disagree with some previous works that suggest that working memory could explain more variance of TMT-B than visual scanning and graphomotor speed (Crowe, 1998; Sánchez-Cubillo et al., 2009). However, these studies used small convenience samples of students (Crowe, 1998) or healthy old volunteers (Sánchez-Cubillo et al., 2009), and multivariate analyses were not adjusted for educational level, which may explain this discrepancy. Moreover, with a different version of the TMT task, Salthouse (2011), also demonstrated in a large sample of 1.056 adults ranging from 18 to 98 years that speed and an index of fluid cognitive ability were more important functions than working memory in the execution of the TMT-B task.
The TMT-d and the TMT-r derived scores are attempts to elucidate the supplementary task requirements of TMT-B by removing the variance attributable to the visual scanning and graphomotor speed. According to our EFAs results, the TMT-d score loaded into a single general factor such as the TMT-B score. However, the EFA including the TMT-r produced a two dimension factor structure, with a speed-related factor that included the SDT, SCWT, t-MQTMA, pVF, and sVF and a factor related to working memory and attention that included the TMT-r score. Interestingly, the multiple linear regression analysis results showed a similar pattern of the TMT-B for the TMT-d, but with a minor participation of visual scanning and graphomotor speed, and the effect of visuomotor processing speed was removed in this case. Instead, for the TMT-r score only working memory and inhibition of the automatic response where the main cognitive abilities involved. None of the visual, graphomotor, or visuomotor speed tasks achieved statistical significance, suggesting that the TMT-r score may be a purest measure of executive functioning in terms of working memory and inhibition control capacity. Previous studies have reported a lack of correlation between this score and various cognitive measures (Corrigan & Hinkeldey, 1987; Sánchez-Cubillo et al., 2009), while others have suggested that it may be related to set-switching abilities (Arbuthnott & Frank, 2000; Ríos, Periánez, & Muñoz-Céspedes, 2004), a cognitive mechanism that was not assessed in this study.
When assessing the relationship between neuropsychological measures, it is important to bear in mind the problem of task purity, especially regarding measures related to the executive functioning (Weiskrantz, 1992). There are several subcomponents of executive functioning such as working memory, inhibition, planning, shifting, and to date there is not an agreement regarding the definition of these subcomponents. Besides, up to 18 different definitions of executive functioning have been reported (Wasserman & Wasserman, 2013). Generally, executive functioning is described as a set of cognitive functions that include abilities of goal formation, planning, carrying out goal-directed plans, and effective performance. This broad dimension of the executive functioning presupposes that this set of functions requires a large variety of cognitive skills that may be quite independent from frontal cortex, and the tests used to assess executive functions are contaminated by the presence of other nonexecutive functions (Rabbitt, 1997).
In this context, our results suggest that the TMT, traditionally defined as a measure of attention, speed, and mental flexibility, has two direct scores (TMT-A and TMT-B) with merged cognitive requirements and two derived scores (TMT-d and TMT-r) that represent purer measures of executive functioning.
The wide use of TMT required the publication of normative data for several languages and versions (Bezdicek et al., 2011; Cangoz, 2009; Cavaco et al., 2013; Hashimoto, 2006; Hester et al., 2005; Lezak et al., 2004; Mitrushina, Boone, Razani, & D’Elia, 2005; Seo et al., 2006; Steinberg, Bieliauskas, Smith, & Ivnik, 2005; Strauss et al., 2006; Tombaugh, 2004). The need of normative data in neuropsychology is especially important in old people, since this group has an increased risk of having cognitive impairment, and since the TMT is a useful test in the process of diagnosing dementia (Ashendorf et al., 2008; Ferman et al., 2006). Only one study has provided normative data in Spain using a population-based sample, but was restricted to people older than 70 years (Del Ser et al., 2004). Other studies have been based in alternative sampling designs. Periáñez (2007) used 223 healthy controls in three age groups (69 participants in the range of 16-24 years, 89 in the range of 25-54 years, and 65 participants in the range of 55-80 years). Peña-Casanova, Blesa, et al (2009), in the project NEURONORMA, provided normative data, stratified by age and education level, using a convenience sample of 354 participants attended at neurology offices in Spain, with no cognitive disorders, and of 50 years and above.
Because the TMT-B requires a specific knowledge (alphabet) that is acquired during schooling, restrictions were applied during data collection, as recommended by CERAD (The Consortium to Established a Registry for Alzheimer’s Disease; Morris, Heyman, & Mohs, 1989), and as performed in other studies (Seo et al., 2006). Refusal or discontinuation of the TMT-B was associated with increased age and low education level of the participants. As much as 11.0% of participants were excluded for requiring more than 300 seconds to complete the TMT-B.
Our results indicate that TMT-A, TMT-B, and the TMT-d scores are associated with age, education level, and sex, but the TMT-r score is only associated with education level. In agreement with previous normative studies, TMT performance decreases with age, and improves with increasing education level. However, the influence of these variables on TMT-A and TMT-B is different (Bornstein, 1985; Bornstein & Suga, 1988; Cangoz, Karakoc, & Selekler, 2009; Ernst et al., 1987; Giovagnoli et al., 1996; Horton & Roberts, 2003; Ivnik et al., 1996; Mitrushina, Boone, Razani, & D’Elia, 2005; Peña-Casanova et al., 2009; Rasmusson et al., 1998; Salthouse et al., 2000; Seo et al., 2006; Spreen & Strauss, 1998; Steinberg et al., 2005; Stuss et al., 1987; Stuss et al., 1989, Tombaugh, 2004; Wecker et al., 2000). Thus, our results are in accordance with those reporting a greater variance in TMT-A explained by age, and greater variance in TMT-d score, and TMT-r score explained by education level. No specific interaction between age and education level was seen in any TMT score, which is also consistent with previous studies (Seo et al., 2006).
Regarding sex, women displayed a worse performance in TMT-A, TMT-B, and TMT-d score than men, showing a decreased effect of age and education level, which had also been previously reported (Cangoz et al., 2009; Giovagnoli et al., 1996; Ivnik et al., 1996; Seo et al., 2006). In part, this is due to the lower education level of women in our cohort of 55 years and older, when compared with men. Other studies performed in the same geographical area (Girona province) obtained similar results using cohorts of this age or older (López-Pousa et al., 2004). On one hand, there are studies showing minimal or no effect of sex, so they do not recommend to adjust the data according to sex (Bezdicek et al., 2012; Hashimoto et al., 2006; Hester et al., 2005; Ivnik et al., 1996; Lucas et al., 2005; Mitrushina, Boone, Razani, & D’Elia, 2005; Tombaugh, 2004; Zalonis et al., 2010). In contrast, other authors report a worsened performance of either men or women, of the whole TMT (Cangoz et al., 2009; Cavaco et al., 2013), of TMT-A (McCurry et al., 2001; Seo et al., 2006), of TMT-B (Bornstein, 1985; Ernst et al., 1987; Wiederholt et al., 1993) or of derived scores (Cangoz et al., 2009). Due to the size of our study sample, and in spite of its modest effect, we included sex in the stratification of normative data.
Reporting two different systems to generate normative data may generate a clinical dilemma regarding which values should be chosen for a particular case. We provide normative data using the traditional method and a regression-based approach because both methods have their advantages and limitations, either in terms of ease of use or in terms of methodology. With the aim of obtaining reliable normative data, we used the interval superposition strategy with the traditional model of stratification of the variables age, education level, and sex. This strategy allowed us to obtain large sample sizes, of more than 19 participants, even for the group of 82 years and older. The complementary use of the regression method increases the reliability of the normative data obtained for groups with reduced sample sizes due to stratification, and provides norms with more continuous and smooth values (Van Breukelen & Vlaeyen, 2005; Van der Elst et al., 2006; Zachary & Gorsuch, 1985). The only difference between both methods is the level of precision; thus, using the traditional method, the clinician obtains a percentile range, while using the regression-based method the clinician obtains a specific percentile value. We recommend the use of the regression-based method for individuals with advanced age and/or high educational level because, typically, in normative studies, these groups have limited sample sizes. Although some calculations are needed, this method increases the reliability of the normative values.
Limitations and strengths of this study must be considered when interpreting the results. This was a nested study in a broad population-based study of cardiovascular risk factors, and time constrictions did not allow us to include neither more cognitive measures nor a measure of general intelligence. The latter has been shown to be an important predictor for several neuropsychological functions and may be used as stratification factor for normative data (Salthouse, 2005). Moreover, not all the cognitive functions described to be related with the TMT execution were included. For example, the task-switch ability that has been associated to the TMT-B and TMT-d performance (Arbuthnott & Frank, 2000; Sánchez-Cubillo et al., 2009) was not included in this study. However, the inclusion of a measure of inhibition control such as the SCWT may palliate this deficit. In this sense, the inhibition of habitual responses in order to make other responses or “switching” from one sequence of responses to another, appear to be logically, and operationally, very similar concepts (Lowe & Rabbitt, 1997). Some of the oldest participants with lowest education refused or discontinued the performance of the test, which may produce a participation bias and may overestimate the normative values for the older participants with low education level. An important strength of this study is the size of the sample, and the fact that it is a population-based sample, which guarantees the reliability of the normative data hereby presented for all genders, age groups, and education levels. Besides, the random selection of the municipalities (rural and urban areas), as well as of the participants, provides highest external validity.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Instituto de Salud Carlos III under Grant number PI09/02591.
