Abstract
The Woodcock–Johnson III Tests of Achievement grade norms versus age norms were examined in the calculation of discrepancy scores in 202 college students. Difference scores were calculated between the Wechsler Adult Intelligence Scale–3rd Edition Full Scale IQ and the Woodcock–Johnson III Total Achievement, Broad Reading, Math, and Written Language scores. The proportion meeting the substantial discrepancy criterion of two standard deviations plus or minus the standard error of the difference between means was 7.9% using age norms and 37.6% using grade norms. Using mixed analyses of variance, the authors found main effects for type of norm for all difference scores, with grade norms yielding significantly higher difference scores than age norms. A main effect for student status (traditional-age students n = 124, non-traditional-age students n = 78) was found for Total Achievement, Broad Reading, and Math such that traditional-age students had greater discrepancies. Discrepancy scores are contrasted with absolute scores (SS < 85) in the identification of learning disabilities as well. Implications of using grade versus age norms are discussed.
In their review of the literature, Sparks and Lovett (2009) found dozens of methods in the identification of learning disabilities, with various discrepancy criteria being one of the most often cited methods. An illustration of how differing criteria in the evaluation of learning disabilities can affect the numbers of college students identified can be found in the 2005 study conducted by Giovingo, Proctor, and Prevatt, which compared three common methods of identifying learning disabilities: the intraindividual discrepancy, the intellectual ability–achievement discrepancy, and the underachievement methods. These authors found differing numbers of students classified as having learning disabilities depending on the method used, supporting an argument for a more nationally standardized definition. Historically, because of the confusion that comes from varying state-adopted criteria for learning disabilities and the possibility that people who qualify for assistance in one location may not qualify under different criteria established in another locale or by another agency, the criteria of the fourth edition, text revision of the Diagnostic and Statistical Manual of Mental Disorders (DSM-IV-TR; American Psychiatric Association [APA], 2000) have been one source of such a uniform definition (Dunham & Preston, 2008).
The DSM-IV-TR criteria for learning disorders suggest that if a person’s scores on a measure of achievement are “substantially below” (i.e., at least two standard deviations) his or her scores on a measure of cognitive abilities, he or she meets a major criterion of learning disorders (APA, 2000, p.49). The DSM-IV-TR discrepancy is considered a conservative one. Giovingo and colleagues (2005) used a more liberal 1.3 standard deviations criterion in their research, whereas Sparks and Lovett (2009) compared three levels of discrepancies in theirs. Regardless of the discrepancy value used, the discrepancy approach has been widely criticized. These criticisms have included the observation that some children with low achievement may not have sufficiently high intelligence test scores to be included using this approach (Dombrowski, Kamphaus, & Reynolds, 2004). Furthermore, given the correlation between intelligence test scores and achievement test scores (Psychological Corporation, 1997, 2008), the discrepancy may shrink over time, leading to an improper conclusion that an older child no longer has a learning disability on retesting. Finally, given the inadequate floor for younger children on many achievement tests, young children with learning disabilities may not show a severe discrepancy if evaluated before the age of 9 (Sattler, 2002).
These criticisms led to the proposed changes to the upcoming DSM-5, which are consistent with the reauthorized Individuals with Disabilities Education Improvement Act (IDEIA, 2004). Both IDEIA and the DSM-5 state discrepancy formulas must not be used to determine learning disabilities, but neither gives specific guidelines for their identification. Dombrowski and colleagues (2004) offer some guidance in resolving the problem of vague or absent criteria coupled with a new prohibition against the discrepancy approach and suggest that a new nationally accepted standard be set. Their proposal includes two primary considerations in identifying learning disabilities: achievement standard scores on a norm-referenced measure greater than one standard deviation below the mean (i.e., < 85) and impaired educational performance (the latter criterion is DSM-IV-TR Criterion B for learning disorders as well).
Beyond identification method, an additional decision in the classification of learning disabilities is the choice of cognitive and achievement measures to be used in the evaluation process. There are many well-standardized, reliable, and valid measures from which to choose (Sattler, 2008). According to Schrank and Flanagan (2003), the most commonly administered measure of adult cognitive ability at the time of the present study’s data collection was the Wechsler Adult Intelligence Scale–3rd Edition (WAIS-III; Psychological Corporation, 1997), and a commonly used assessment of achievement in adults is the Woodcock–Johnson III Tests of Achievement (WJ-Ach; Mather & Woodcock, 2001).
Although some other major tests of achievement have only age norms, the WJ-Ach is the only achievement test that gives the test administrator a choice of grade norms for adults older than 19 years old, meaning an individual’s performance can be compared to either age-matched or grade-matched peers. The WAIS-III (and now the Wechsler Adult Intelligence Scale–4th Edition, WAIS-IV; Psychological Corporation, 2008), on the other hand, offers only age norms for converting and interpreting scores of cognitive ability. The practice of comparing scores on measures normed on different samples, such as with the WAIS-III and WJ-Ach, although not advisable, is widely practiced (Sattler, 2008). Schrank and Flanagan (2003) suggest that the convention of using the WAIS-III and the WJ-Ach together could be the result of tradition or preference, among other factors. A more questionable practice is the use of an age-normed measure of cognitive ability and a grade-normed measure of achievement, as one is comparing scores derived from different underlying distributions.
When considering which WJ-Ach norms to use, the clinician must be aware that standard scores are affected by the normative choice (McGrew, Woodcock, & Ford, 2002). The developers of the WJ-Ach state that grade norms are the preferred choice for school-based decisions, whereas age norms would be more appropriate for use in clinical settings (Mather & Woodcock, 2001). For example, they recommend that for an adult returning to college after several years, the most appropriate comparison group is other college students (Mather & Woodcock, 2001). In examining the aforementioned example for non-traditional-age college students, or students older than is considered typical for a specific grade, two distributions exist for age versus grade norms. When a non-traditional-age college student is compared to same-age peers, the full normal distribution of ability levels is represented; however, when compared to those in the same grade (or other college students), the curve is truncated (Anastasi & Urbina, 1997). Grade-normed scores are still normally distributed, but the scale has shifted because of the loss of scores at the lower end of the distribution. Therefore, non-traditional-age college students who have average achievement scores compared to same-age individuals may have below-average scores compared to others of the same grade.
Illustrating this point, in a study comparing proportions of college students identified as meeting the criteria for learning disabilities using either age- or grade-based norms, Giovingo and colleagues (2005) found significantly more students were identified using grade norms than age norms for two of the three learning disability identification methods they studied. Using the Woodcock–Johnson III Tests of Cognitive Abilities (WJ-Cog) and the WJ-Ach as the measures of cognitive and academic abilities, respectively, these authors also found that the fewest students were identified using the intellectual ability–achievement method, regardless of the norms used. Although they gathered WAIS-III scores, these scores were not examined to determine whether similar findings would be observed if the WAIS-III were used as the measure of cognitive abilities (Giovingo et al., 2005).
Although the authors of the WJ-III advance an argument in favor of using grade norms in educational settings, there are several arguments against their use beyond the possible overidentification of learning disabilities suggested by Giovingo and colleagues (2005). Schrank and Flanagan (2003) suggest that the use of any grade-based norm system is inherently unsound as norms for grades are of uneven metrics and are incredibly variable across ages. Furthermore, some state guidelines specify that age norms must be used when interpreting ability level and achievement level for special education programs (Schrank, Flanagan, Woodcock, & Mascolo, 2002).
The present study was conducted to determine the impact of using age norms versus grade norms from the WJ-Ach, when the measure of cognitive abilities was the WAIS-III, among adult college students seeking a learning disability evaluation, using the intellectual ability–achievement discrepancy method. Furthermore, the present study specifically examined the more conservative discrepancy criterion of the DSM-IV-TR (APA, 2000), as recommended by Giovingo and colleagues (2005) as an area for further study. Rates of identification using the absolute score method (i.e., SS < 85) proposed by Dombrowski et al. (2004) were also explored. In addition, the present study aimed to examine the possibility of a differential impact of age versus grade norms depending on the age of the college student. It was expected that the effect of using grade norms would be greater for students of nontraditional age than those of traditional age.
Method
Participants
Participant data came from the records of college student clients who requested an evaluation of learning disorders because of academic difficulty or failure in their college classes. Students were enrolled in a midsized, nonselective, public, 4-year university in the Midwest. All students who completed an evaluation and signed consent for their records to be used for research purposes between the years of 2001 and 2008 were used, yielding 202 usable archival records. The sample was 65% female and 35% male, ranging in age from 18 to 63 (M = 27.07, SD = 9.65). Of these participants, the majority were Caucasian (86.9%), followed by African American (10.1%), and the remainder identified themselves as belonging to other ethnic groups.
The records were divided into two groups: Those from traditional-age college students (n = 124) and those from non-traditional-age college students (n = 78). Nontraditional age was operationalized as being greater than 5 years older than the average for one’s grade. For example, a traditional college freshman was considered to be 18 years old; any freshman older than 23 was deemed to be of “nontraditional age,” any sophomore older than 24, and so forth. The groups did not significantly differ in gender, χ2(1) = 0.38, p > .05, or race, χ2(4) = 4.19, p > .05. There were differences in mean Full Scale IQ, t(200) = 2.44, p < .05, such that traditional-age students had higher mean scores than non-traditional-age students (M = 101.28, SD = 11.44; M = 97.45, SD = 9.88). The groups did not differ in their achievement scores, whether age normed or grade normed, with the exception of grade-normed Broad Written Language scores, t(196) = 2.65, p < .01. Means and standard deviations of WJ-III achievement scores are presented in Table 1.
Mean Standard Scores by Student Status.
p < .05.
Materials
WAIS-III
The WAIS-III is designed to measure several specific abilities that, when combined, compose an estimate of the individual’s overall cognitive ability or Full Scale IQ (FSIQ). FSIQ scores are reported using age-normed referencing. FSIQ was used as the measurement for individual ability in this study. The FSIQ score has demonstrated internal consistency reliability, with an alpha coefficient of .98 across all ages that the measure is designed to assess (Psychological Corporation, 1997). FSIQ scores of the WAIS-III and the WAIS-IV are highly correlated (r = .94; Sattler & Ryan, 2009), as are the FSIQ of the WAIS-III and the equivalent score from the Stanford-Binet IV (r = .88; Psychological Corporation, 1997).
WJ-Ach
The WJ-Ach is designed to measure several areas of achievement that, in turn, compose an overall measure of achievement, or Total Achievement. Raw scores can be converted to standard scores using either age or grade norms. A subset of the adult participants (190 in a 2-year college and 975 in a 4-year university) enrolled in a school or university was used to construct the grade norms (McGrew & Woodcock, 2001). The internal consistency of Total Achievement has been reported as .98 (McGrew & Woodcock, 2001). Three subscores that reflect more individualized areas, or clusters, of achievement are the Broad Reading, Broad Math, and Broad Written Language scores. The internal consistency estimates of these cluster scores are .94, .95, and .94, respectively (McGrew & Woodcock, 2001). Correlation coefficients between .65 and .79 have been reported between the WJ-III and other leading measures of achievement (McGrew & Woodcock, 2001; Sattler, 2008).
Procedure
Information including age, gender, race, college classification, and WAIS-III and WJ-Ach scores was gathered from existing client files. WJ-Ach scores were converted into both age-normed scores and grade-normed scores. The participants used to construct norms for the college and university grade brackets came from the 13th grade through the 18th grade, although when using the computer scoring software, the researcher enters the grade and specifies as a comparison group either a 2-year or 4-year college or university. In the present study, the 4-year college or university norms were selected. Each of the WJ-Ach age-normed and grade-normed Total Achievement, Broad Reading, Broad Math, and Broad Written Language scores were subtracted from the WAIS-III FSIQ score to come up with difference scores that were used in the discrepancy analyses. The DSM-IV-TR learning disorder criterion of a substantial discrepancy was used (i.e., a person’s Total Achievement, Broad Reading, Broad Math, and Broad Written Language score(s) on the WJ-Ach being greater than or equal to two standard deviations plus or minus the standard error of the difference between means below his or her FSIQ). Although WAIS-III Verbal IQ (VIQ) and Performance IQ (PIQ) discrepancies were calculated for all clients, FSIQ was universally used to calculate discrepancy scores. Nine participants had a statistically and clinically significant VIQ-PIQ discrepancy (seen in fewer than 5% of the standard sample), with two falling in the traditional-age student group and seven falling in the non-traditional-age student group.
Results
The number of traditional-age students who met the DSM-IV-TR substantial discrepancy criterion using age-normed Total Achievement (ATA), Broad Reading (ABR), Broad Math (ABM), and Broad Written Language (ABWL) scores was 14 (11.29%). When the FSIQ scores were compared to any of the grade-normed Total Achievement (GTA), Broad Reading (GBR), Broad Math (GBM), or Broad Written Language (GBWL), the number of traditional-age students with significant discrepancies increased to 49 (39.52%). Comparatively speaking, when the non-traditional-age students’ FSIQ scores were compared with their ATA, ABR, ABM, and ABWL scores, 2 (2.56%) individuals met the substantial discrepancy criterion, whereas 27 (34.6%) met the discrepancy criterion using the GTA, GBR, GBM, and GBWL scores. Proportions by type of norm and student status can be found in Table 2.
Proportion of Students Meeting Discrepancy Criterion.
For comparison purposes, proportions of students meeting the absolute score criterion proposed by Dombrowski and colleagues (2004) are reported in Table 3. Again, the proportions meeting this criterion using grade norms exceed those found when age norms are used, with total proportions ranging from 5.9% to 13.4% using age norms. When grade norms are used, the range jumps to 33.2% to 55.9% of students identified.
Proportion of Students Meeting Absolute Score Criteria (SS < 85).
Traditional-age students met the DSM-IV-TR discrepancy criterion for reading disability using GBR scores significantly more often than non-traditional-age students, χ2(1) = 5.53, p < .05. Traditional-age students met the absolute score criterion for mathematics disability using ABM scores significantly more often than non-traditional-age students, χ2(1) = 5.31, p < .05. The proportions meeting either the discrepancy or absolute score criterion were not significantly different between traditional-age and non-traditional-age students, using either age or grade norms, for the remaining learning disorder (LD) types.
Data were evaluated with 2 × 2 mixed analyses of variance (ANOVAs), with traditional-age versus non-traditional-age students serving as the between-group factor, whereas FSIQ-WJ-Ach discrepancy scores using age norms versus grade norms served as the within-group factor. When comparing Total Achievement scores, a main effect for type of norm revealed significantly larger difference scores when using GTA scores over ATA scores, F(1, 190) = 1532.79, p < .0001 (means and standard deviations of difference scores can be found in Table 4). Furthermore, there was a main effect for student status such that traditional-age college students had significantly greater difference scores than non-traditional-age students, F(1, 190) = 4.48, p < .05. For reading scores, there were main effects for both type of norm and student, such that all students had significantly greater difference scores when using GBR scores instead of ABR scores, F(1, 196) = 188.23, p < .0001, and traditional-age students had significantly greater discrepancy scores than non-traditional-age students, F(1, 196) = 6.64, p < .05.
Mean Difference Scores by Student Status.
Note: Means in the same row that share subscripts of the same letter differ at p < .05.
Similarly, main effects for type of norm and student status were found for math discrepancy scores. Students had significantly larger discrepancy scores when using GBM scores instead of ABM scores, F(1, 195) = 4,453.23, p < .0001, and traditional-age students had significantly larger discrepancy scores than non-traditional-age students, F(1, 195) = 8.43, p < .001. Furthermore, a significant interaction was found between type of norm and student status, F(1, 195) = 22.50, p < .0001, such that traditional-age students had significantly greater ability–achievement discrepancy scores than non-traditional-age students using GBM scores as opposed to ABM scores. In written language, a main effect for type of norm was found, such that the use of GBWL scores led to significantly greater difference scores than the use of ABWL scores, F(1, 196) = 1214.76, p < .0001. A significant interaction between type of norm and student status was found, F(1, 196) = 31.50, p < .0001, such that traditional-age students had significantly greater discrepancy scores than non-traditional-age students using GBWL scores as opposed to ABWL scores.
The mixed ANOVAs were rerun excluding the nine students who had a significant split between their VIQ and PIQ scores. All significant main effects and interactions remained, except the main effect for student status using Total Achievement scores, F(1, 182) = 3.59, p < .06. However, the removal of these nine students yielded a significant interaction between type of norm and student status for Total Achievement discrepancy scores, F(1, 182) = 5.61, p < .05.
When examining the absolute score method of identification, using a mixed analysis of variance (with WJ-Ach score and type of norms as the within-subjects variable and student status as the between-groups variable), there was a significant main effect of norm type, F(1, 190) = 2261.21, p < .0001, such that grade-normed standard scores were lower than age-normed scores.
Point-biserial correlations were calculated to determine whether the difference between a student’s age and grade was related to the likelihood of meeting the substantial discrepancy or absolute score criteria. Age–grade correspondence was calculated by subtracting grade from age. Regardless of type of norm used. There was no significant correlation found between age–grade correspondence and meeting the DSM-IV-TR substantial discrepancy criterion. Correlations ranged from –.13 to .00, with a median correlation of –.09. The same was true for the absolute score method (correlations ranged from –.15 to .08), with one exception. Having an age-normed score of less than 85 on the Broad Mathematics scale was correlated with greater age–grade discrepancy.
Discussion
The goal of this study was to investigate how the use of grade norms from the WJ-Ach affects the identification of learning disabilities using the DSM-IV-TR substantial discrepancy criterion in college students. Although Mather and Woodcock (2001) suggest that the WJ-Ach grade norms are the most appropriate choice in some cases, the question that was presented in this investigation was the following: What is the effect of using WJ-Ach grade norms versus age norms in an adult population?
Analyses from this study compared two groups (traditional-age and non-traditional-age students) and two types of norms (age based and grade based) when using the substantial discrepancy criterion. The proportion of both traditional-age and non-traditional-age students who met this criterion was greater when using grade norms as opposed to age norms. The proportion was 7.9% using age norms and 37.6% using grade norms, with generally similar increases in proportions for both traditional-age and non-traditional-age students. Furthermore, there were significantly larger difference scores using grade-normed scores than age-normed scores for Total Achievement, Broad Reading, Broad Math, and Broad Written Language. This trend was seen for both traditional-age and non-traditional-age students, and sometimes even more so for the traditional-age college students.
It was predicted that the use of grade norms from the WJ-III would lead to greater difference scores in non-traditional-age students, but the fact that this was found to be the case among both traditional-age and non-traditional-age students is intriguing. The focus on non-traditional-age students stemmed from the test developers’ recommendation that grade norms be used with non-traditional-age students, but not necessarily when assessing traditional-age students. The WJ-III developers stated that a more typical age–grade correspondence, such as that seen among traditional-age college students, should result in similar scores across type of norms. The present results, however, suggest that the downward adjustment of standard scores that occurs when norming to grade with adults in the higher grades significantly affects the chances of finding a substantial discrepancy. The finding that both traditional-age and non-traditional-age college students are similarly affected by the use of grade norms was further supported by the nonsignificant correlations between the age–grade correspondence of a student and substantial discrepancies. It was expected that as age–grade discrepancy increased, so would the likelihood of finding a substantial discrepancy, because increasingly older students in cohorts with lower mean levels of education (age) would be compared to younger cohorts with higher mean levels of education (grade); therefore, the gulf between age-compared IQ scores and grade-compared achievement scores would similarly widen. Another unexpected finding was that traditional-age students had significantly greater difference scores than non-traditional-age students on GBR and GBM. These results can be partially explained by the higher mean scores that were seen on the WAIS-III FSIQ among traditional-age students. Since the traditional-age student sample had higher ability scores but similar achievement scores when compared to non-traditional-age students, difference scores would, in turn, be greater.
As predicted by Dombrowski and colleagues (2004), the use of the absolute score method (i.e., SS < 85) resulted in higher proportions of students being identified than when using the discrepancy approach. They indicated that up to 13% to 15% of students could be identified as having learning disabilities with this approach. In the present data, a sizeable proportion were designated as LD using this method using both age and grade norms. Again, grade norms led to a significantly greater proportion than age norms with this identification method, and certainly more than the projected proportion of 13% to 15%.
This study was conducted with a sample of convenience, such that those who desired to be assessed presented with academic difficulties. This means that this study can be generalized only to college students experiencing academic difficulties. Furthermore, in this sample, students of nontraditional age had lower mean IQ scores than those of traditional age. Although both means still fell within the average range, and the difference perhaps does not amount to a clinically significant finding (being less than 4 IQ score points and therefore within the standard error of measure for an individual), this may be a finding that is unique to this sample and not generalizable beyond the current study. Alternatively, even such a small decrement in cognitive ability may be reflective of one factor that leads some adults to postpone furthering their education. Because of the loss of students of lower cognitive ability at the lower end of the achievement grade norm continuum and the present study’s focus on students experiencing academic difficulty, it is unclear whether the same discrepancies between age and grade norms at the middle and upper ranges of the achievement score distributions would be observed.
The WAIS-III is not normed on the same population as the WJ-III, nor are grade norms available. Perhaps comparing age-normed ability scores to grade-normed achievement scores is akin to comparing apples to oranges, as one is comparing scores derived from different underlying distributions. Because the developers of the WJ-III state it is preferable to use grade norms in assessing non-traditional-age students, it would be instructive to replicate this study using the grade norms from both the WJ-Cog and WJ-Ach. Using the same norming system when assessing a college student, as opposed to mixing and matching the types of norms as is often practiced, is more psychometrically sound.
If opting for the use of both parts of the WJ-III in evaluating adults, one should be aware that Krasa (2007) found that of 52 subtests on the Cognitive and Achievement batteries of the WJ-III, only 18 of them had adequate ceilings and item gradients for ages 16 to 25 and grades 10 to 18. Of the 10 subtests of the WJ-III standard cognitive battery for measuring ability, 3 subtests passed the ceiling and item gradient standards, whereas of the 12 subtests found in the WJ-III standard achievement battery, half passed these standards. This finding raises a question as to the adequacy of these batteries in college students, who may fall at the upper ends of both the age and grade continua. Krasa did report that the item gradients below the mean and floor scores for all of the WJ-Ach subtests are considered adequate for adults and/or college students. This is important because in both the discrepancy and absolute score methods of identification, achievement scores should be below the mean of the sample. However, the influence of individual subtest scores on broader scales, which may artificially inflate or deflate the broader scores, may still be a concern.
It is also important to note that the WAIS-IV (Psychological Corporation, 2008) has been available for several years. Whether the same trends that were seen in the present study would be seen in this new version despite changes and enhancements that have been made in the development of the WAIS-IV is an empirical question. Given the reported correlation (r = .94) between the FSIQ scores of the two versions, similar results are likely. Certainly, the conundrum of using different norming groups will still be present, even with this updated instrument.
The primary hypothesis investigated in this study was found to be only partially supported. The use of grade norms did significantly increase the number of students identified using the substantial discrepancy method. However, the prediction that this would be true more often for non-traditional-age college students was not supported. This suggests that the use of grade norms affects the identification of LD in college students regardless of traditional-age or non-traditional-age student status. Based on the finding that with the use of grade norms 37.6% of the sample qualified for an LD using the substantial discrepancy criterion and even more using the absolute score method, if this were generalized, one third to one half of college students presenting with academic difficulties could be diagnosed with an LD. According to the National Center for Education Statistics (2006), 11.3% of all college students have a disability. Of these, 7.5% report their disability to be a specific learning disability. This means that approximately 0.85% of college students have an LD. As the majority (83.3%) of students being evaluated for learning disabilities at the center where the current study was conducted had not previously been identified as LD, the use of grade norms seems to lead to an overidentification of LD in college students. Both the DSM-IV-TR criteria and those proposed by Dombrowski and colleagues (2004) suggest that learning disabilities are developmental in nature and therefore do not manifest for the first time in the college years. Grade-normed standard scores may be helpful additions to a psychoeducational evaluation of a college student for descriptive purposes, such as in characterizing his or her performance in his or her current academic environment, but should not be used to identify learning disabilities in students who are struggling academically for the first time in their academic careers.
We echo Dombrowski and colleagues’ (2004) recommendation for a uniform method of classifying learning disabilities. Because of the possibility that two college students, with similar cognitive and academic abilities, could be classified differently depending on which normative group is used, we would like to add to their guidelines a call for adopting a single standard that requires use of one type of norm. Based on the data from the present study, age norms appear to be the more appropriate choice when evaluating college students.
Footnotes
Authors’ Note
This article is based on the master’s thesis of the first author.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
