Abstract
This study compared prorated Boston Naming Test (BNT-P; omitting the noose item) and standard administration (BNT-S) scores in physical medicine and rehabilitation patients (N = 480). The sample was 34% female and 91% White with average age and education of 46 (SD = 15) and 14 (SD = 3) years, respectively. BNT-P was calculated by summing correct responses excluding item 48 and estimating the 60-item score with cross multiplication and division. BNT-P and BNT-S scores were compared via concordance correlation (CC) coefficients; reflected and log transformed data were examined with equivalence tests. BNT-P and BNT-S scores showed almost perfect agreement (CC = .99). Transformed scores demonstrated equivalence (±1.1 points). Raw and scaled score differences were 0 in 88% and 96% of cases, respectively. Race and ethnicity accounted for item 48 outcomes while controlling for age and education. Findings support the utility of prorated BNT scores in rehabilitation patients.
The Boston Naming Test (BNT; Kaplan et al., 1983) remains one of the most commonly used confrontation naming measures by clinical neuropsychologists (Rabin et al., 2016). However, questions persist concerning the test’s cultural fit and norms, particularly in reference to ethnicity (Harry & Crowe, 2014; Horwitz & McCaffey, 2010; Werry et al., 2019). The test has long been scrutinized for apparent performance discrepancies across non-White individuals, with particular emphasis on African American populations. These differences remain after controlling for education, occupation, literacy, and medical history (Werry et al., 2019). These differences are also apparent among both older and younger African Americans (Na & King, 2019), which suggests a need for closer analysis of racial norms, item-level functioning, and administration of the BNT.
Despite development of several cultural revisions of the BNT, a survey of members of the National Academy of Neuropsychology indicated that only 7.6% of clinicians reported using these cultural adaptations of the BNT. Additionally, fewer than half of the survey respondents made adjustments for ethnicity (Bortnik et al., 2013). This is significant, as the norms provided with the test are limited regarding ethnicity and geographic location. Use of such normative data may not be appropriate for non-White groups, and may result in misdiagnosis of dementia or mild cognitive impairment in older adults (Werry et al., 2019). In addition, some studies have indicated that while establishing age, education, and gender norms among African American populations can help reduce misdiagnosis (Manly et al., 1998; Werry et al., 2019), it may remain insufficient because educational and cultural experiences can vary greatly within groups. Item response theory-based approaches have identified redundancy in BNT items and indicated the feasibility of creating a BNT version using items with equidistant difficulty parameters and high discriminability that would not demonstrate differential item functioning between White and African American examinees (Pedraza et al., 2011). To our knowledge, such a form has not been developed. Therefore, ongoing consideration of the cultural impact of BNT items is imperative to improving assessment accuracy.
An overlooked area of concern regarding the cultural appropriateness of the BNT is the inclusion of Item 48, the noose item. The noose, or hangman’s knot, has a long history as a method of execution in many countries where its image may serve as a reminder of the grim realities of capital punishment. In the United States, however, the noose also brings to mind the sordid history of lynching by which more than 4,700 individuals in the country were killed between 1880 and the 1960s (Potok et al., 2007).
As practiced during its heyday, lynching was brutal, vicious, and routine. More often than not, to be lynched was to be hanged. Not only were people hanged, but they were often beaten, burned, and bludgeoned before or after and their bodies left on display. The rope helped facilitate this display, holding the body high off the ground for all to see. (Shuler, 2014, p. 5)
The noose is a powerful reminder of this part of U.S. history, “remains a terrifying symbol, and continues to be used by racists to intimidate African-Americans” (Potok et al., 2007). The rationale for continued use of a symbol of racial violence (Elassar, 2020; McTaggart, 2014) in an objective measure of cognitive ability warrants further consideration. Empirically, performance discrepancies observed between African American and White examinees on the item (Pedraza et al., 2009) provide psychometric challenges to its utility. Pedraza and colleagues found differential item functioning on 12 BNT items altogether, suggesting that examining test scores as a whole may be insufficient for understanding the factors contributing to group (specifically, racial group) differences. It is also possible that use of an item with connotations of racial violence may hinder performance or introduce stereotype threat (Thames et al., 2013).
Potential solutions include administering other measures of visual confrontation naming (e.g., Neuropsychological Assessment Battery [NAB] Naming; Stern & White, 2003) or using one of the BNT short forms that do not include Item 48 (Katsumata et al., 2015). NAB Naming is promising in that it includes color stimuli and has shorter response windows than the BNT, but it remains limited by a ceiling effect that may be more pronounced than that of the BNT (Harry & Crowe, 2014; Sachs et al., 2016). In addition, some findings raise questions about the sensitivity of the measure to cognitive dysfunction. Messerly and Marceaux (2020) examined NAB Naming in a mixed clinical sample of whom 64% demonstrated cognitive impairment. They observed significant differences in NAB Naming between groups diagnosed with major neurocognitive disorder/dementia, mild neurocognitive disorder, and no cognitive impairment, but median scores in the groups were 28, 30, and 31, respectively. The average NAB Naming score in the whole sample was within normal limits (M = 48.1, SD = 10.2) and “the majority of participants [earned] perfect or near perfect scores of 30 or 31” (p. 414). Pulsipher et al. (2013) compared clinical groups with unilateral left and right hemisphere stroke with matched control participants on multiple NAB measures including Naming. They reported group differences consistent with expectation in that the control group scored higher on NAB Naming than both stroke groups and the right hemisphere stroke group scored higher than the left hemisphere stroke group. The mean NAB Naming score in the left hemisphere stroke group, however, was within normal limits (M = 40.1, SD = 15.9).
Turning to BNT short forms, there are multiple 15- and 30-item BNT short forms that do not include the noose item (Katsumata et al., 2015). We recently examined the utility of five 15-item short forms and three 30-item short forms in relation to the standard BNT administration in rehabilitation patients (Attridge et al., 2020). With impairment on the BNT as criterion (≤35 T) all three 30-item short forms showed outstanding classification accuracy (Hosmer & Lemeshow, 2000) and two 15-item short forms demonstrated classification accuracy comparable to the 30-item short forms. Although these results offer promising initial support of short forms, the available normative data on BNT short forms are derived from samples aged 50 years and older. In a rehabilitation setting, older adult normative data are useful for the majority of cases involving stroke, but are not relevant to the majority of patients with a history of traumatic brain injury (TBI), who are typically younger adults. In comparison, the standard BNT has normative data for adults aged 20 to 85 years (Heaton et al., 2004) and for adolescents aged 15 to 18 years (Martielli & Blackburn, 2016). Normative limitations currently preclude use of BNT short forms with patients younger than 50 years.
While this article was under review, one of the testing companies that distributes the BNT, Pro-Ed, released a Stimulus Upgrade Kit that includes an alternative item that can be used in place of the noose item (Pro-Ed, n.d.). The Pro-Ed web site provides the follow explanation in support of the proposed change: While the existing item is psychometrically sound, it is racially and culturally inappropriate. As a result, we are encouraging customers to replace Item 48 with the sticker upgrade kit. . . . The new item is “boomerang.” The rationale for the new items [sic] is as follows:
It is a multisyllabic word. Word length is a lexical factor that positively correlates with word finding difficulties. Multisyllabic words are long words and thus helpful in drawing out word finding difficulties in inaccurate namers who have difficulty retrieving the complete phonological sequence of long words (commonly known as a twist of the tongue).
The target word resides in a low density or sparse neighborhood (few phonological neighbors [words that sound similar to the target word]). Individuals with word finding difficulties have more difficulty with words from sparse neighborhoods.
It is a low frequency word. Individuals with word finding difficulties have more difficulty with low frequency words.
The web site does not include any linguistic data on the proposed replacement item to demonstrate its comparability to the original item. No psychometric data are presented that support the item’s utility in classifying naming impairment. Furthermore, no data are presented that demonstrate equivalence between BNT versions with the noose item and with boomerang item. Given the lack of psychometric data supporting its use, the novel item should undergo additional study prior to widespread clinical adoption.
In addition to these limited alternatives to the standard BNT administration, another solution is to omit the item and prorate the BNT raw score allowing extant BNT normative data to be used. The equivalence of prorating the BNT has not been established, and this study aimed to examine the psychometric properties of such a procedure. We compared prorated BNT scores and scores from the standard administration to assess for equivalence in physical medicine and rehabilitation patients. We hypothesized that patient’s scores on the BNT form with the omission of Item 48 would be equivalent to their scores on the full BNT. A secondary aim involved subgroup analyses comparing performance on other language tasks in participants who passed Item 48 and those who failed it, examining the performance of White and non-White participants on the item, and considering non-White participant BNT performance with and without Item 48.
Method
Participants
This study involved retrospective examination of individuals who underwent outpatient neuropsychological evaluation in a physical medicine and rehabilitation clinic at an academic medical center in the western United States. Inclusion criteria were age of 20 years or older, completion of the BNT, and valid performance. History of English as a second language was an exclusion criterion with 22 cases (4.6%) excluded for this reason. The sample was used in a secondary analysis investigating the utility of BNT short forms (Attridge et al., 2020). Differences in the peer-review process led to the second paper going to press before the current article.
The sample (N = 458) was 34% female. The average age was 45.9 years (SD = 15.2; range 20-86). Average educational level was 13.9 years (SD = 2.5; range 8-20). Regarding race and ethnicity, the sample was 94.1% White, 2.0% Hispanic/Latino, 1.3% African American, 1.3% Asian American, and 1.3% other ethnicities. Ninety-two percentage of participants were right-handed. Presenting diagnoses included TBI (62%), cerebrovascular events (17%), and other neurologic and psychiatric conditions (21%). Sample diagnostic characteristics are presented in Table 1. Published criteria were used to classify cases with history of TBI (Orman et al., 2011) based on available data in medical records. Mild-complicated and moderate cases were combined (Kashluba et al., 2008).
Frequency of Diagnosis.
Measures
Boston Naming Test (Kaplan et al., 1983)
The BNT is a 60-item visual confrontation naming measure. Participants were administered the BNT according to standardized instructions starting at Item 30 (BNT-S) with an eight-item basal rule and six-item discontinuation rule. Scores were entered as correct for items that were not administered (i.e., prior to the start point or reversal point). BNT prorated scores (BNT-P) were calculated by summing correct responses excluding Item 48 and then using cross multiplication and division to estimate the 60-item score equivalent. BNT raw scores were converted to norm-referenced scores (Heaton et al., 2004).
Additional Language Measures
Additional measures of language were administered to different numbers of participants. These measures were included a single-word pronunciation test (Test of Premorbid Functioning [TOPF]; Pearson, 2009) completed by 444 participants. The TOPF standard score was used. Subtests from the Wechsler Adult Intelligence Scale–Fourth edition (Wechsler, 2008) were examined including Information (n = 439) and Similarities (n = 434). Demographically corrected T scores were used for these subtests. Phonemic (FAS; n = 451) and semantic (AN; n = 449) word generation tasks were also examined (Gladsjo et al., 1999). Demographically corrected T scores were used for these tasks (Heaton et al., 2004).
Analyses
BNT-P and BNT-S scores were compared with concordance correlation coefficients (CCC; Lin, 1989), and root mean square differences (RMSD) using an absolute definition of agreement (Barchard, 2012). The CCC was developed to measure departures from a 45° line that would be expected from plotting the results of two tests with perfect agreement, which is not captured using a Pearson correlation coefficient (Lin, 1989). CCC takes into account departures from the 45° line at the level of observations as well as any deviation from the 45° line in the linear relationship between the two tests (King & Chinchilli, 2001). CCC values range from 0 to 1, and values of .90 or higher are considered evidence of moderate or better agreement (McBride, 2005). CCC was calculated using the concord package installed in Stata 13.1. The RMSD indicates the average difference in scores provided by two tests on the same scale as the tests. An absolute definition of agreement indicates that any difference between scores is considered disagreement (Barchard, 2012). RMSD was calculated using an Excel sheet (Barchard, 2011). Violin plots were used to examine the score distributions of BNT-S and BNT-P.
The lack of a significant difference between scores on a paired t test cannot be taken as evidence of agreement (Lin, 1989). Equivalence tests, however, demonstrate that score from different tests are equal within a predetermined boundary. Given the nonnormality of BNT data, BNT scores were reflected and log transformed prior to analyses with equivalence tests. Equivalence tests were conducted using a two one-sided test (TOST) procedure for dependent samples (Lakens et al., 2018). The TOST approach involves a priori determination of an equivalence boundary (± delta) within which score differences can be considered trivial. Two t tests are then conducted: (a) testing the null hypothesis that the lower boundary of a 90% confidence interval around the true score difference is lower than the low boundary of the equivalence interval; and (b) testing the null hypothesis that the upper boundary of a 90% confidence interval around the true score difference is higher than the upper boundary of the equivalence interval. If both of these null hypotheses are rejected, then the scores can be said to be equivalent within the designated equivalence interval. Delta was set a priori at 1.1 raw score points, which log transforms to 0.1 and results in an equivalence interval of −0.1 and 0.1. BNT-P and BNT-S difference and norm-referenced scores were examined descriptively. Subgroup analyses involved descriptive statistics, chi-square tests for categorical data, and t tests for continuous data. BNT-P and BNT-S difference scores were compared in White and non-White groups using a Wilcoxon rank-sum test. A logistic regression model was fitted with BNT Item 48 score (correct or incorrect) as the outcome and age, education, TOPF standard score, verbal ability, and dichotomous race/ethnicity as predictors. A verbal ability summary score was calculated by averaging the demographically adjusted T scores on Information, Similarities, FAS, and AN.
Results
BNT-P and BNT-S scores demonstrated almost perfect agreement (CCC = .999, p < .001; McBride, 2005). The RMSD between BNT-P and BNT-S was 0.27 indicating an average change in raw score between measures less than one-third of a raw score point.
BNT-P (M = 53.2, SD = 5.9, Mdn = 53.9) and BNT-S (M = 53.2, SD = 6.0, Mdn = 54) raw scores were similar. Violin plots of BNT-P and BNT-S scores are shown in Figure 1. The TOST lower boundary test was significant, t(457) = −19.25, p < .001. The TOST upper boundary test was significant, t(457) = 20.85, p < .001. Together, the TOST findings demonstrated that transformed BNT-P and BNT-S scores were within the equivalence boundary (±0.1 log transformed points; ±1.1 raw score points).

Violin plots of raw scores from standardized and prorated administration.
The mean difference between BNT-P and BNT-S scores was −0.01 (SD = 0.3, Mdn = −0.1, range −0.5-1). In 89% of cases, the difference score rounded to 0. The mean difference between BNT-P and BNT-S scaled scores was 0.03 (SD = 0.2, Mdn = 0, range 0-2). Scaled score differences were 0 in 96.9% of cases, 1 in 13 cases, and 2 in one case. In the 14 cases with scaled score differences, T score differences were 4 points in 11 cases, 5 points in 2 cases, and 8 points in 1 case. A slope graph of changes in T scores between administrations is in Figure 2.

Slope plot of T score changes from standardized and prorated administrations in cases (n = 14) with score differences.
Subgroup Analyses
Fifty participants responded incorrectly to Item 48 representing 11% of the sample. These participants were not significantly different from participants who responded correctly to the item in age, t(456) = 0.16, p = .87. Participants who responded incorrectly to Item 48 had completed significantly less education (M = 13.2, SD = 2.0) than participants who responded correctly to the item (M = 14.0, SD = 2.5), t(456) = −2.11, p = .04. Participants who responded incorrectly to item 48 performed significantly worse on additional measures of language including TOPF, Information, Similarities, FAS, and AN. Details are in Table 2.
Boson Naming Test Item 48 and Performance on Other Language Measures.
Note. Scores are demographically corrected T scores except for TOPF, which is a standard score. Data are presented as M (SD). TOPF = Test of Premorbid Functioning; Information = Wechsler Adult Intelligence Scale–Fourth edition information subtest; Similarities = Wechsler Adult Intelligence Scale–Fourth edition similarities subtest; FAS = phonemic word generation; AN = semantic word generation.
Comparing White and non-White participants on Item 48, 9.5% of White participants responded incorrectly to the item, and 32.1% of non-White participants responded incorrectly to the item, χ2(1) = 13.8, p < .0001. With BNT Items 30 through 60 ranked from easiest to most difficult based on percentage correct in this sample, Item 48 was fifteenth among White participants and twenty second among non-White participants. Non-White participants showed slight differences in BNT-S (M = 51.1, SD = 6.5, Mdn = 50.5) and BNT-P (M = 51.3, SD = 6.3, Mdn = 50.8) with an average difference of 0.17 (SD = 0.4, Mdn = −0.05). BNT-S and -P difference scores were not significantly different between White and non-White participants, z = 1.8, p = .08. A logistic regression model predicting Item 48 was preferred over a constant-only model, χ2(5) = 55.7, p < .0001. A Pearson goodness-of-fit test was not significant (p = .51) indicating good model fit. The model correctly classified 89% of cases. While controlling for age and education, TOPF standard score (odds ratio [OR] = 1.05), verbal ability (OR = 1.09) and ethnicity (OR = 5.66) made significant contributions to the model.
Discussion
The BNT has been criticized for including an item that is likely offensive to some examinees and examiners (Horwitz & McCaffey, 2010). Despite this concern, there has not been any empirical study, to our knowledge, of the effect of removing the item on BNT scores. This study is the first examination of the effect of omitting the noose item and prorating BNT scores. Given the limitations of typical null hypothesis testing and correlation coefficients with regard to the study question, our analytic approach focused on methods of evaluating agreement between two measures including CCC, RMSD, and TOST equivalence procedures. These findings were promising in demonstrating nearly perfect agreement between standardized and prorated BNT scores and score equivalence within a fairly narrow equivalence boundary.
Norm-referenced scores also demonstrated evidence supporting the utility of prorated BNT scores. Scaled score differences were 0 in just under 97% of cases. In the 14 cases with scaled score differences, the T score differences were 5 points or less in 13 cases and 8 points in one case. Considering these T score changes in terms of impairment (defined as scores ≤35 T), only one case showed a meaningful change with the score moving from an impaired range to an unimpaired range (i.e., BNT-S = 34 T and BNT-P = 38 T; Figure 2). Although score changes were observed in the other 13 cases, the impact was minimal in terms of score impairment with impaired scores remaining impaired and normal scores remaining normal. In the vast majority of cases in the current sample, prorating BNT scores demonstrated limited impact on norm-referenced scores and resulting interpretation.
Subgroup analyses identified evidence of the differential performance on Item 48 among non-White participants. Non-White participants were more likely to respond incorrectly to the item consistent with prior research (Na & King, 2019; Werry et al., 2019). Participants who responded incorrectly to Item 48 performed significantly worse than participants who responded correctly to it on separate measures of language including word pronunciation, verbal knowledge, verbal reasoning, and word generation tasks indicating that it has some potential to discriminate language ability. The difficulty of Item 48, however, was different between White and non-White participants, which implicates race and ethnicity as contributing to performance on the item in addition to language ability. The contribution of racial and ethnic factors in performance on Item 48 was also demonstrated by the logistic regression model results in which dichotomous race/ethnicity significantly contributed to the model with a much larger OR than those observed on reading and verbal ability while controlling for age and education. Given differential performance on Item 48 in white and non-White groups and its connotations of racism and violence, it appears reasonable to stop using the item.
Despite these positive preliminary findings, it may be helpful to consider the implications of post hoc modification of a widely studied instrument with well-known clinical correlates and documented psychometric properties. It can be argued that such a modification is inappropriate on clinical grounds because the majority of the research conducted using the BNT was based on the standard administration. Reflecting on the possible rationale for not changing the BNT raised several related questions: Does our opinion that Item 48 is offensive constitute sufficient grounds for modifying a test that we did not create? Will such an act serve as inspiration for other groups of clinicians and researchers to undertake similar modifications of other tests? Do we risk the eventual dismantling of our neuropsychological armamentarium by a subsequent wave of ideologically driven test modifications?
We address these questions in order. Regarding the merits of various reasons for altering an established psychometric instrument, it appears that numerous prior BNT modifications (i.e., creating short forms) were undertaken simply for the purpose of saving time. Although the rationale for such a modification might initially appear pragmatic rather than principled, there is most certainly an underlying ideological position that prioritizes efficiency both as an end in itself and as a means of providing beneficent patient care. We agree with these principles and propose that omission of the noose item promotes both efficiency and patient well-being. A more recent modification was proposed to be in recognition of the “racially and culturally inappropriate nature” of Item 48, so concerns with the item reported here do not appear to be uniquely ours (Pro-Ed, n.d.). Regarding the possibility that these proposed changes to the BNT will inspire similar efforts, we are merely the latest in a long line of researchers to explore modifications of the BNT (Katsumata et al., 2015), so it seems inaccurate to characterize our work as pioneering. Finally, we agree that unchecked, idiosyncratic modification of standardized instruments would potentially reduce the field of neuropsychology and its systematic approach to quantifying cognitive and emotional status to the level of astrologic prognostication or, even worse, to some form of pseudoscientific quackery. Protection of our tests, even from ourselves, is an important consideration and propositions regarding such modifications need to be justified.
An additional counterpoint to modification of the BNT by omitting Item 48 arises from evidence that the item has demonstrated utility in differentiating language ability. The current findings indicate that individuals with incorrect responses to item 48 also showed significantly lower scores on other language tasks than participants with correct responses to Item 48 including word reading, fund of knowledge, verbal reasoning, and word generation measures. The decision whether the item’s utility outweighs its potential to cause offense will be ultimately be determined at the level of the individual clinician or researcher. Our goal has been to provide initial psychometric support for an alternative in the event that there is interest in one.
Limitations to the current study include the relative lack of ethnic diversity and the relatively greater proportion of males in the sample. Given the evidence of differential item function in diverse ethnic groups, these findings require replication in more diverse samples. Although the sample included more general neurologic conditions (e.g., dementia), the majority of cases had a history of acquired brain injury. Additional examination of prorated BNT scores in a wider range of neurologic conditions is important. The current study design involved counting items that were not administered as correct because they were prior to the start point or basal point in cases requiring reversal. It cannot be assumed that items prior to a start point would necessarily have been answered correctly (Stålhammar et al., 2016).
Current findings are interpreted as promising, initial evidence that the noose item can be removed from the BNT and the administered items prorated with limited impact on psychometric and clinical utility of the measure. Future research might examine such prorated BNT scores in a more diverse sample and with a different clinical population. Given pressure to work more efficiently, an additional promising area of research would be to investigate optimized BNT administration with computer-adaptive branching based on responses.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
