Abstract
We evaluated the validity of using the GRE General Test to assist with graduate school admissions for individuals with disabilities. We studied a sample of 16,239 graduate students from 10 U.S. research universities in three groups: students without any reported disabilities, students who reported disabilities and took the computer-delivered GRE with accommodations, and students who reported disabilities but took the computer-delivered GRE without accommodations. We examined differential prediction using multilevel modeling and residual analyses. The results revealed that the first-year graduate grade point average of students with disabilities was neither over- nor underpredicted by more than one tenth of a point on the 0- to 4-grading scale. However, variations on the magnitude and direction of differential prediction existed among students with different types of disabilities. We discuss data collection needs and research on students with disabilities attending graduate and professional schools.
The GRE General Test (GRE) measures prospective graduate students’ verbal and quantitative reasoning abilities and analytic writing skills. Empirical research suggests that GRE scores predict various measures of success in graduate school, including the widely used first-year grade point average (GPA; see Burton & Wang, 2005; Kuncel, Hezlett, & Ones, 2001; Kuncel, Wee, Serafin, & Hezlett, 2010; Liu, Klieger, Bochenek, Holtzman, & Xu, 2015; Powers, 2004). However, there is limited research on GRE-related predictive validity for students with reported disabilities, herein referred to as students with disabilities for brevity. The most recent study on students with disabilities taking the GRE was completed more than 30 years before this study (Braun, Ragosta, & Kaplan, 1986). Since then, several significant changes have occurred, including the introduction of an adaptive computer-based test format, the addition of new item types, and changes to the constructs being assessed. The examinee population has also changed, in part due to more students with disabilities studying in U.S. graduate schools. Such changes make new validity evidence important for graduate students with disabilities and admissions officials.
The GRE was most recently revised in 2011 and is currently delivered to most students using a multistage adaptive test model (see Robin, Manfred, & Liang, 2014). Students with disabilities must be approved to take the test with specific accommodations, but they may take the test without accommodations. The following types of accommodations are available for the computer-delivered GRE: extended time, extra breaks, screen magnification, selectable colors, screen reader, and refreshable braille compatibility (Educational Testing Service [ETS], 2018). In addition to changes to the test over the past three decades, there have also been changes to disability policy, including the Americans With Disabilities Amendment Act (ADAA) and the decision by ETS and other testing organizations to stop alerting score users when tests were administered under nonstandard conditions (Sireci, 2005). Because of these changes, evaluating whether differential prediction of the GRE exists for examinees with disabilities relative to those without disabilities is important. Periodic evaluation of fairness related to score use (e.g., using differential prediction analyses) helps to ensure that examinees with disabilities who take the GRE with or without accommodations obtain scores that support valid inferences about future success in graduate school (see American Educational Research Association, American Psychological Association, & National Council on Measurement in Education [AERA, APA, & NCME], 2014).
Our general goal in this study was to examine the predictive validity of GRE section scores for students with one or more types of disabilities and those without disabilities. Our research question is “Do GRE scores, alone or in combination with undergraduate GPA (UGPA), predict first-year graduate GPA (FYGPA) differently for graduate students with disabilities relative to those without disabilities?” We specified two focal groups of students with disabilities based on how they took the GRE: (a) examinees with disabilities who requested the use of accommodations and were both approved for and took the GRE with one or more testing accommodation (group D1) and (b) examinees who took the GRE test with no accommodations but reported on a background questionnaire administered during the test registration process that they have a disability (group D2). Individuals in group D2 either did not submit a request for accommodations or requested them but were not approved to receive accommodations. Examinees without a reported disability (group ND) constituted the reference group.
Prior Research on Postsecondary Differential Predictive Studies
In a meta-analysis of GRE predictive validity studies, Kuncel et al. (2001) described the theoretical and practical arguments for the design of such studies, including the choice of criteria. In most research studies of using GRE scores to predict graduate school performance, the criterion is either first-year or cumulative GPA. Other variables, for example, degree attainment and time to degree completion (e.g., Kuncel et al., 2001, 2010), are often hard to obtain and are affected by factors other than the skills that GRE scores are intended to predict (e.g., Burton & Wang, 2005). Predictive validity evidence for the GRE has been gathered in studies spanning various disciplines (see, for example, Wendler & Bridgeman, 2014). Adjusted correlations between GRE scores and graduate GPA have generally been estimated in the .30s, on average.
Evidence on the fairness of using GRE scores for subgroups is typically studied via differential prediction analyses (AERA, APA, & NCME, 2014). Differential prediction for students with disabilities has been studied for applicants and enrolled students in medical (Searcy, Dowd, Hughes, Baldwin, & Pigg, 2015), law (Amodeo, Marcus, Thorton, & Pashley, 2009), and graduate schools (Braun et al., 1986). Performance differences, which are often related to differential prediction, have also been studied for prospective business school applicants who took the Graduate Management Admission Test (GMAT) with accommodations (Johnson, Rudner, & Sibert, 2008). A key question in differential prediction studies is whether one model predicts the outcome of interest similarly for individuals regardless of group membership. Differences in prediction can be caused by either difference in the means of the predictors, the means of criterion variable, the relationship between predictors and the criterion, or all three (Linn, 1978). We summarize relevant studies in the areas of medical, law, and business schools before detailing a previous study on differential prediction for students with disabilities taking the GRE.
In a study of students taking the Medical College Admissions Test (MCAT), Searcy et al. (2015) reported that students with disabilities who took the test with extended time accommodations had no significant differences in average MCAT scores compared to students who reported no disabilities and had no extended time accommodations, either among the medical school applicants (d = .00) or among the enrolled students in medical schools (d = .05). Medical students with disabilities in their sample had lower UGPAs, on average, than medical students without disabilities. Using logistic regression models estimated separately by medical school, they found no significant difference in medical school acceptance rates after controlling for MCAT scores and UGPA. However, they found that MCAT scores significantly overpredicted measures of intermediate and long-term success in medical school for students with disabilities, especially medical licensing exam pass rates and graduation rates.
Amodeo et al. (2009) reported that applicants and students with disabilities enrolled in law school were comparable to those with no disability, a trend similar to that in the MCAT study (Searcy et al., 2015). In addition, in a study using the GMAT, Johnson et al. (2008) found no significant performance differences on admissions test performance, on average, between students taking the test with and without accommodations. Although the study by Johnson et al. did not evaluate differential prediction, we include it here to summarize that the literature on professional school applicants taking exams other than the GRE shows no performance differences on the assessments used as predictors, but there is evidence of differential prediction for the two tests studied, MCAT and Law School Admission Test (LSAT).
The most recent predictive validity study for GRE examinees with disabilities was conducted by Braun and colleagues in 1986. Braun et al. (1986) studied scores of examinees with learning disabilities (n = 19), visual impairment (n = 105), or a physical disability (n = 48) who took an accommodated administration of the GRE and examinees with disabilities who were administered the GRE under standard conditions (n = 184). Specific types of disabilities were not indicated for the latter group. The criterion variables were first-year and overall graduate school GPA. Of the 432 graduate schools that were contacted, 78% responded, and of them, 151 institutions reported information on students with disabilities who took the GRE with or without accommodations. Graduate school information was obtained for students with disabilities in 41 graduate departments. Due to limitations with their data collection, students in their reference group, students with no disclosed disabilities, were not from the same institutions as students in their focal groups with disabilities. However, mean GRE scores of all score senders to each of the graduate school departments were available from ETS records and were used in the analysis (N = 2,025). Information on specific testing accommodations was not reported.
In the context of these findings, we wanted to revisit the relationships between GRE scores and performance in graduate school. As developed previously, we examined whether the GRE scores of students with disabilities, alone or in combination with their UGPA, predict first-year graduate students’ GPA differently for individuals with disabilities relative to those without disabilities.
Method
Given the limitation that students with disabilities were not in the same departments as students without disabilities in Braun et al. (1986) and several decades have passed since their study, we revisited the question of the relationship between GRE scores and FYGPA. In this section, we describe the data sources, data collection procedures, and analytic methods.
Sample
Our first focal group (D1) included examinees with disabilities who were approved to use one or more testing accommodations and took the 2011 version of the computer-based GRE. The test score database included the approved accommodation(s) on a given test date and linked them to individual examinees’ GRE section scores, but disability subtypes were not included in the database. This is because individuals approved for accommodations did not go through the regular test registration process and consequently did not complete a voluntary Background Information Questionnaire that included an item about disability types for research purposes. To link their types of disabilities, we used an electronic file that listed GRE examinees approved for accommodations along with their type(s) of disabilities from the ETS Disability Office. This file and the test record file do not share a common identification number, so we used combinations of first name, middle initial, last name, and Social Security number to match them. Our procedures were approved by the institutional review board, and we followed standard protocols for privacy and security. For the population of examinees between August 2011 and June 2015 who were approved for accommodations, we were able to match 82% of their test records to the file providing disability type. Those unmatched cases included missing values for students’ personal identifiers, duplicate names, changes to last names, and spelling errors in the test records. Among individuals listed in the file of disability types, we matched 54% to test records. Only 5% of unmatched individuals in the disability type file were approved to take a paper-based form of the test. The high percentage of unmatchable records in the disability-type file may have been due to examinees who submitted documentation for accommodations but then did not end up testing or tested after the dates from which we drew our sample.
Our second focal group (D2) included examinees who used the regular test registration process and reported a disability type from an item on the Background Information Questionnaire administered during regular registration for the GRE. These examinees either did not request accommodations or were not approved for accommodations. There were no students in our sample who were classified in both D1 and D2. There were a few students in group D2 who received accommodations on at least one other test administration but did not send those scores to our studied institutions. We assigned these students to group D2, assuming our studied institutions used only the nonaccommodated test scores they had on record. We assigned all other students not in groups D1 or D2 to group ND and assumed they had no reported disabilities. Note that the ND group may include students with disabilities who did not identify themselves. We did not have access to independent information to check the extent to which students with disabilities were included in the ND group.
Institution Recruitment
With two focal groups of examinees identified, we began to recruit graduate institutions in order to obtain FYGPA for examinees who subsequently enrolled in and completed their first year of graduate school. We attempted to increase the sample size per institution relative to Braun et al. (1986) by focusing on approximately 45 institutions with large volumes of examinees with disabilities. Note that similar to Braun et al., we had to use GRE records for score senders, not enrolled students, to identify institutions that might have data to contribute to this study, given that students with disabilities are not typically identifiable by their graduate institutions.
We used two main approaches to determine which institutions would be invited to participate in this study. In the initial approach, we selected the top 43 institutions with the highest counts (out of 1,172 institutions), each with at least 153 score senders (D1 students) who used accommodations.
In our second approach, we invited D2 examinees through e-mail to participate in a two-question online survey. In the survey, we asked them whether they had been enrolled in a U.S. graduate program between 2012 and 2016, and if so, we asked them to provide the name of the institution in which they were enrolled. Participants in the short survey were entered into a gift card raffle with a chance to win either a $50 or $25 gift card. Similar to the first approach, we tabulated counts by attending institution of GRE examinees and who responded to our survey (n = 283). We identified two institutions that we had not previously contacted in the initial approach but that had five or more respondents to our survey. Thus, our final recruitment pool included a total of 45 institutions. In most cases, the disability support services staff helped us connect with the institutional research office or with the registrar’s office to obtain the data. Of the 45 institutions we contacted, 10 agreed to participate.
Data matching
We provided institutions that agreed to participate with three data-sharing options. In the first option, the institution provided ETS with a file that listed student name, gender, birthdate, UGPA (if available), master’s or doctoral program, major Classification of Instructional Programs code, GRE scores used in the admissions decision, and FYGPA for all graduate students enrolled in their institution beginning with the 2012–2013 cohort and ending with either the 2014–2015 cohort (one school), the 2015–2016 cohort (seven schools), or the 2016–2017 cohort (two schools), depending on when the data were collected from the institution. This information was matched to ETS data sources. The second data-sharing option was similar to the first, except that files were exchanged such that the institutions kept the personally identifiable information (i.e., student name, birthdate, etc.) separate from any file containing test scores and achievement data (i.e., GRE scores, UGPA). In the third data-sharing option, we provided to the participating institutions a list of names of examinees (with and without disabilities) who submitted their GRE scores to the institution paired with nonidentifying ID numbers. We asked the institutions to provide us with the same variables listed under the first option for those individuals who enrolled and completed their first year and the ID numbers (again, institutions did not send back to us identifiable student information with student grades). We did not share information with any institution about examinees’ information related to disability type or approved accommodations. Note that no matter which data-sharing option an institution chose, we did not know the counts of students with disabilities enrolled in each institution until after contracts were executed and the data were shared and fully matched. Details of the sample characteristics are available in the Results section.
Institutional sample characteristics
We contacted a total of 45 institutions for participation in the study. Of these, 23 had minimal-to-no response to the initial recruitment e-mails, 11 declined to participate, one institution was interested but did not have the data we needed, and 10 institutions agreed to participate in the study. Of the 10, six belonged to American Association of Universities only, one belonged to American Association of Universities and the Consortium on Financing Higher Education, and two did not belong to either organization. See Table 1 for institutional sample characteristics, region, and the number of students from each institution.
Institutional Sample Characteristics.
Note. Region is based on the United States Census classification. Source for institutional characteristics is the Carnegie Classification of Institutions of Higher Education. AAU = American Association of Universities; COFHE = Consortium on Financing Higher Education; D1 = students with at least one type of disability who took the GRE with accommodations; D2 = students who reported having at least one type of disability but took the regular GRE with no accommodations; ND = students without disabilities.
Refers to R1 doctoral degree–granting institutions that have a very high level of research activity, as compared to R2 (high research activity) and R3 (doctoral or professional universities).
Comprehensive graduate programs, with or without medical or veterinary school.
There did not appear to be major differences between the 10 institutions that elected to participate in the study versus the 35 that did not. The 10 participating institutions were large (n = 10; 100%), were mostly public (n = 8; 80%), generally had very high research activity (n = 9; 90%), tended to be part of the American Association of Universities (n = 7; 70%), and were mostly undergraduate (n = 9; 90%), although most had comprehensive graduate programs with medical or veterinary schools (n = 8; 80%). The 35 institutions that did not elect to participate were also large (n = 33; 94%), generally had very high research activity (n = 33; 94%), tended to be part of the American Association of Universities (n = 27; 77%), and were also mostly undergraduate (n = 25; 71%) and had comprehensive graduate programs with medical or veterinary schools (n = 27; 77%). However, only about half were public institutions (n = 17; 50%), and the other half (n = 18; 50%) were under private control. It is reasonable to assume that differences between the participating and nonparticipating institutions are not significantly related to differences in FYGPA between students with and without disabilities.
Analysis
To prepare for the exploration of differential prediction for students with disabilities relative to students without disabilities, we first computed descriptive statistics on the three sections of the GRE: Quantitative Reasoning (GRE-Q), Verbal Reasoning (GRE-V), and Analytical Writing (GRE-AW). The descriptive statistics are based on three samples: (a) the population of GRE examinees between 2011 and 2015, (b) the examinees who sent scores to one or more of the 10 institutions where we collected data, and (c) the students (n = 21,742) who enrolled in the 10 institutions. For students in the last sample (collected in this study), we analyzed two additional variables: school-provided UGPA and school-provided FYGPA. Within each sample, we compared the two focal groups to the reference group with effect size estimates for each variable where data are available. For the effect size estimates, we computed mean differences in the unit of the standard deviation of each variable based on the reference group (group ND). We chose the reference group standard deviation in our effect size computation because of the small size of the focal groups and the small differences between the combined and the reference group in their standard deviations (Glass, McGaw, & Smith, 1981; Hedges, 1981).
We prepared the data for the differential prediction analysis in the following ways. Each participating institution provided records of students’ official UGPA. However, a small percentage of the institution-provided UGPA were not in the 4-point scale between 0 and 4, with some between 4 and 20 and a few from 30 to 100. These cases were international students from three institutions (D, G, and J), about 7.5% at school D, and less than 1.4% at both schools G and J. The cases with a UGPA above 4 or below 1 were removed from related analyses. We used a similar approach for FYGPA. As a result, we included only students with UGPA or FYGPA between 1 and 4. For GRE-AW scores, scores between 0 and 6 were all considered valid, and all other institution-supplied values were treated as missing. We removed cases with invalid values on UGPA, FYGPA, and GRE scores and obtained a restricted sample that included examinees with a valid score on all five variables (n = 16,239).
Considering the data structure, with students nested in majors, and majors nested in schools, we deemed a multilevel model appropriate for predicting FYGPA from GRE section scores and UGPA. A multilevel model could explicitly account for the potential variation of regression slopes and residual variances among majors while also providing a mechanism to test hypotheses about the variation in these parameters, which makes it more flexible than traditional ANOVA method (e.g., Raudenbush, Bryk, Cheong, & Congdon, 2004; Raudenbush, Bryk, & Congdon, 2013). Although our data with only 10 institutions are likely to result in downward biased estimates of the between-institution variance (Snijders, 2005), our interest was more in the point estimates of the prediction slopes, which have been shown to be adequately accurate even with a small number of institutions (Maas & Hox, 2005).
To explore the possibility of differential prediction, we analyzed the data using four different three-level hierarchical linear models, including the baseline model with no predictors, a model using UGPA as the only predictor of FYGPA (Model 1), a model using GRE scores as the only predictor of FYGPA (Model 2), and a model using both GRE scores and UGPA as the predictors of FYGPA (Model 3). In each model, we chose to account for the nesting structure at three levels, that is, students within each major or department (31 majors in total) and major nested in each institution (10 institutions in total). The model specification is in the online appendix.
We chose to estimate the parameters using the total sample (with both focal groups—students with disabilities—and the reference group—students without a disability), an approach that has been used in research at the graduate level (e.g., gender and ethnic groups for GMAT, Sireci & Talento-Miller, 2006) and at the undergraduate level (e.g., Young, 2001). The prediction residual mean and standard deviation were computed for each student and then compared among the students without and with disabilities. A positive residual indicates that the predicted FYGPA is lower than what we observed, (i.e., underprediction); a negative residual indicates a predicted FYGPA that is higher than what we observed (i.e., overprediction).
Results
Descriptive Statistics
In total, data from 21,742 students were obtained from 10 institutions that were geographically scattered around the United States. Of the sample, 0.5% were in group D1 and this varied from 0.2% to 1.2% across institutions. About 1.3% of the total sample were in group D2 and varied from 0.8% to 1.9% across institutions. See Table 1 for counts across institutions. We combined master’s and doctoral students together because our focus was on grades in first-year courses, which are often common between the two groups.
We classified students’ majors into five broad categories: business, education, humanities, social sciences, and STEM (science, technology, engineering, and mathematics) (Bridgeman, Burton, & Cline, 2008; National Center for Education Statistics, 2014). Table 2 describes the broad majors for enrolled students in our sample institutions and compares them to both estimates from the GRE examinee population and a national sample of students with conferred degrees. The trends were generally consistent among students in our three groups, ND, D1, and D2. The distributions of students in ND, D1, or D2 appear to vary significantly among the five broad major categories, (χ2 = 89.09, df = 8, p < .005).
Distribution of Students in Five Broad Majors Within a National Population, the GRE Examinee Population, and Each Subgroup Within Our Sample (in percentages).
Note. ND = students without disabilities; D1 = students with at least one type of disability who took the GRE with accommodations; D2 = students who reported having at least one type of disability but took the regular GRE with no accommodations; STEM = science, technology, engineering, and mathematics.
GRE population estimate based on Table 3.2 in Educational Testing Service (2016).
The national estimate was based on National Center for Education Statistics’ (2014) data of graduate degrees conferred.
The sample consisted of 47.6% female overall, with 3.4% not identifying gender. This distribution was not significantly different from the GRE examinee population, with 49.5% females and 5.9% not identified, respectively (χ2 = 0.91, df = 2, p = .339). The association between gender and the three student groups was not significant, either (χ2 = 42.71, df = 4, p < .001). White students accounted for 46.5% of the total sample, followed by Asian or Asian American students (5.6%, including Hawaiian and Pacific Islander), Hispanic (4.0%, including Mexican, Puerto Rican, and other Hispanic and Latino American students), students identifying in other groups (3.2%), Black or African American students (3.0%), and students who did not identify their ethnicity information (37.2%). We did not have race-ethnicity information for group D1. For group D2, there were many more White (67.0%) and American Indian or Alaskan (2.1%) students and slightly more Hispanic (5.7%) and African American or Black (5.7%) students.
Table 3 shows the means and standard deviations of the three GRE section scores for the studied sample (N = 21,742; ND, n = 21,356; D1, n = 103; D2, n =283). They are labeled “enrolled students” in the table. To compare, we included descriptive statistics for individuals who sent scores to the 10 institutions in the past 5 years (D1, n = 2,551; D2, n = 6,807) and the GRE examinee population. Because of the very large sample sizes, all mean differences were statistically significant based on planned comparisons between each of the focal groups and group ND; consequently, we report effect sizes to compare groups D1, D2, and ND. Of particular note, students in group D1 consistently scored higher, around a third to half of a standard deviation, on GRE-V and GRE-AW compared to students in the ND group. These mean differences were also significant based on planned comparisons (GRE-V, t = 3.32, p < .001; GRE-AW, t = 3.86, p < .001). Students in group D2 scored about the same as group ND on GRE-V and GRE-AW, and the differences were not statistically significant (GRE-V, t = 0.49, p = .63; GRE-AW, t = 1.37, p = .17). Among enrolled students and score senders, both groups D1 and D2 scored lower than group ND on GRE-Q (enrolled D1, t = 2.07, p = .04; enrolled D2, t = 8.80, p < .001; senders D1, t = 11.53, p < .001; senders D2, t = 42.49, p < .001). In general, trends across groups for GRE section scores were similar for the population, score senders, and enrolled students. Note that average GRE section scores were consistently lowest for the GRE examinee population and highest for enrolled students. Such a difference may be related to selection from the admission process, as admitted students tend to have higher mean scores than the examinee population.
Descriptive Statistics.
Note. Standard deviations are in parentheses. Data are from most recent test date and include examinees with a valid score on any of the variables. Test dates for score senders and the GRE examinee population ranged from August 2011 to July 2015. ND = students without disabilities; D1 = students with disabilities who received accommodations; D2 = students with disabilities who did not test with accommodations; UGPA = undergraduate grade point average; FYGPA = first-year graduate grade point average.
Standardized mean difference = (Focal mean – reference mean)/reference standard deviation. The reference group is group ND.
For enrolled students only.
p < .05. **p < .01. ***p < .001.
Also shown in Table 3 are means and standard deviations for UGPA and FYGPA for enrolled students. Students in group ND had a lower mean UGPA than group D1 but about the same as group D2. Mean FYGPA was lower for groups D1 and D2 relative to group ND by about one tenth to one fifth of a standard deviation.
To answer the differential prediction research question, we used only records with complete data on all GRE section scores, UGPA, and FYGPA for the enrolled sample. This required removing about 23% to 25% of examinees in each subgroup and resulted in a total of 16,239 students, including 15,945 in ND, 79 in D1, and 215 in D2. Figure 1 shows the deviation of mean scores for students in groups D1 and D2 from that of group ND on UGPA, GRE section scores, and FYGPA in the unit of standard deviation for each variable based on group ND, following a similar approach as in Braun et al. (1986). Note that the trends are similar to those reported in Table 3 for enrolled students with at least one valid data point across GRE sections, UGPA, and FYGPA.

Mean scores of students in the studied sample (N = 16, 239) in groups D1 and D2 on UGPA, GRE sections, and FYGPA in relation to students in group ND. ND = students without disabilities; D1 = students with disabilities who received accommodations; D2 = students with disabilities who did not test with accommodations; UGPA = undergraduate grade point average; FYGPA = first-year graduate grade point average.
Table 4 shows the average performance for students in the studied sample within groups D1 and D2 by disability type. Given the small sample sizes, we created three categories for students in D1: learning disabilities only (LD), attention deficit hyperactivity disorder or attention deficit disorder only (ADHD), and students with a combination of disabilities, which could include LD and ADHD or other physical or cognitive disabilities (multiple). Students with ADHD had the lowest UGPA and FYGPA, by approximately one third and two thirds of a standard deviation of group D1, respectively, but the highest GRE scores (0.10 to 0.30 standard deviations higher than the total group D1). Table 4 shows variability both across and within groups.
Mean Undergraduate Grade Point Average (UGPA), First-Year Graduate Grade Point Average (FYGPA), and GRE Score by Disability Type.
Note. Standard deviations in parentheses. The sample size of D1 was based on the restricted sample (n = 16,239). D1 = students with disabilities who received accommodations; D2 = students with disabilities who did not test with accommodations; ADHD = attention deficit hyperactivity disorder or attention deficit disorder only; LD = learning disabilities only.
Students in this group chose other types of disabilities on the test registration questionnaire.
For group D2, disability subtypes include all response categories available to examinees on the background questionnaire, including other. Of particular note, students who identified as having physical disabilities had relatively higher UGPA, GRE-AW, and GRE-V (about a third of a standard deviation higher) but the lowest FYGPA (about 0.40 of a standard deviation lower), on average. Students who identified as having multiple disabilities had lower GRE scores relative to the total group D2, on average, but their UGPA and FYGPA were higher than the group average. Note that the sample sizes within groups, particularly for multiple disabilities, were very small.
Analysis of Differential Prediction
In this section, we describe the analyses we used to answer our research question, “Do GRE scores, alone or in combination with UGPA, predict FYGPA differently for students with disabilities relative to students without disabilities?” The results in this section were based on the restricted sample with complete information (n = 16,239). As noted previously, the sample included 15,945 students in ND, 79 in D1, and 215 in D2. Note that the slope estimates, residuals, and variance estimates at each level differed only trivially from those we estimated based on all the data for models not using UGPA (not reported here). Recall that relative mean values for this restricted sample are reported in Figure 1, using effect sizes with the ND group as reference. Before estimating the multilevel models, we estimated Pearson product correlation coefficients within each institution among institution-reported UGPA, FYGPA, and GRE section scores (Table 5). Overall, GRE-Q appears to have a low-level association with FYGPA at most institutions. The correlation coefficients in Table 5 are lower than those between FYGPA and UGPA, GRE-V, and GRE-AW at each institution. The correlations between GRE-Q and FYGPA were generally lower than in other studies, but the pattern of lower correlations for GRE-Q relative to GRE-V is a consistent finding (Kuncel et al., 2001, 2010). The pattern of correlations for the total sample did not hold when broken out by the five broad major categories.
Correlations Between First-Year Graduate Grade Point Average and Each Predictor, by Institution.
Note. Values in parentheses have been corrected for range restriction and reliability on each GRE section score. UGPA = undergraduate grade point average; STEM = science, technology, engineering, and mathematics.
In the online appendix, in Table A1, we report model estimates and fit statistics for the baseline model with no predictors, Model 1 (UGPA only), Model 2 (GRE only), and Model 3 (GRE and UGPA). Evaluation of the model fit based on the likelihood estimates (Table A2) suggested that Models 2 and 3 fit better than the baseline model. We used Model 3 to compute residuals and average them over the three groups.
A comparison of residuals (Table 6) among students grouped by disability type revealed that students in groups D1 and D2 had negative residuals, on average, regardless of whether the model included GRE scores alone, UGPA alone, or both. But the magnitude of the residuals aggregated at the subgroup level is small. This is evidence of slight overprediction of FYGPA, on average, for students in groups D1 and D2. Note that these residuals are on the standard 4-point grading scale. For example, a residual of –.07 suggests that a student who had an FYGPA equal to 3.40 would be predicted to earn a 3.47 GPA. Notice the variation in average residuals when aggregated by groups of students with similar primary disabilities within groups D1 and D2. We did not find evidence of differential prediction for students with LD or students with multiple disabilities. In contrast, we would predict that students with ADHD would score more than .3 points higher on the 4-point GPA scale (e.g., 3.7 for an observed FYGPA of 3.4). Within group D2, examinees with physical disabilities have negative residuals of .28, on average, and examinees who reported having visual impairments have negative residuals of about .15, on average. Examinees who reported having hearing-based disabilities, LD, or other types of disabilities did not show evidence of over- or underprediction.
Residual Means and Standard Deviations, Overall and by Disability Subtype.
Note. UGPA = undergraduate grade point average; ND = students without disabilities; D1 = students with disabilities who received accommodations; D2 = students with disabilities who did not test with accommodations; ADHD = attention deficit hyperactivity disorder or attention deficit disorder only; LD = learning disabilities only.
Discussion
Our research question was, “Do GRE scores, alone or in combination with UGPA, predict FYGPA differently for students with disabilities relative to students without disabilities?” We were able to collect data from 10 U.S. graduate schools on 386 GRE examinees who reported disabilities, among whom 103 were approved to receive accommodations on the GRE. We were able to apply statistical models to data from 294 students who reported disabilities, which included 79 students who were approved to use accommodations on the GRE. We found minimal differential prediction for students with disabilities in both focal groups overall, with some variability across disability subtypes. Differential prediction was minimal despite large differences in scores on some sections of the GRE. The minimal level of differential prediction is consistent with findings in Braun et al. (1986), who found that GRE scores overpredicted first-year graduate school grades for students with disabilities by about one tenth of the standard deviation of first-year graduate grades for students without disabilities. The findings in this study also converge with trends observed in other predictive validity studies for postsecondary students with disabilities (Searcy et al., 2015; Sireci & Talento-Miller, 2006).
We found minimal differential prediction for students with disabilities in both focal groups overall, with some variability across disability subtypes.
The results also suggest that in general, GRE scores could positively predict FYGPA, either by themselves or jointly with UGPA. Most of the institutions in our sample had a predictive relationship for verbal and analytical writing scores toward the lower end of the 90% confidence intervals based on previous studies in similar contexts (Briel & Michel, 2014; Kuncel et al., 2001, 2010). But the relationship between quantitative scores and FYGPA was much lower in our samples than what was found in meta-analyses. We attribute the lower correlations in part to a combination of sensitivity to our data-processing decisions and the negative skew in the distribution of FYGPA.
The results also suggest that in general, GRE scores could positively predict FYGPA, either by themselves or jointly with UGPA.
For admissions officers, policy makers, and other users of GRE test scores, the results from this one study must be interpreted cautiously, primarily given that the samples were nonrandom and the numbers of students with disabilities were small. That is, we cannot draw broad conclusions about the possibility that the GRE is predicting graduate school success measures differently for examinees with disabilities. Although we cannot generalize to all institutions using the GRE, our results provide some evidence that the GRE is working similarly for predicting FYGPA for students with disabilities, with and without accommodations, and students without disabilities. The results corroborate other evidence, namely, minimal overprediction. In other words, we have not found evidence that students with disabilities would have performed better in their first year of graduate school than their GRE score predicted. But neither this study nor the previous study (e.g., Braun et al., 1986) provides direct evidence that using GRE scores in the graduate admissions process would tend to keep out otherwise qualified graduate students or let in students who tend to leave without graduating. Consequently, further research on this population is warranted.
Although we cannot generalize to all institutions using the GRE, our results provide some evidence that the GRE is working similarly for predicting FYGPA for students with disabilities, with and without accommodations, and students without disabilities.
Limitations
That this study shared similar data collection limitations as other studies points to the need for the field of higher education to put in place mechanisms to collect high-quality data on postsecondary students with disabilities (cf. Green & Rabiner, 2012). In the following, we describe the limitations of this study in detail to assist researchers and policy makers alike in setting up conditions to facilitate a stronger data collection. The study required us to collect data on the criterion measure, graduate school success, from graduate institutions in which students with disabilities and their peers enrolled. In the absence of an existing database with criterion measures and one that identified graduate students with disabilities, we embarked on a complex data collection procedure to find enrolled graduate students with disabilities and obtain student data from institutions. It is estimated that about 7% to 8% of graduate students are identified as having a disability (Council of Graduate Schools, 2011), and less than 2% of the GRE test-taking population is identified as having a disability. Given the small number of students with disabilities in graduate school and the fact that we did not know a priori in which institutions the examinees with disabilities were enrolled, randomly sampling graduate institutions was not feasible. For the differential prediction analysis, the studied sample appeared to differ from the GRE population slightly, including in the distributions of gender, ethnicity, and disability status. Our studied sample represented a slightly higher-performing group than the GRE population, and the differences between students with and without disabilities in GRE scores were narrower in our studied sample than the GRE score senders. This likely is due to a selection effect through the admission process.
Despite several years of targeted recruitment, we could not obtain large-enough samples to reliably disaggregate by disability type or by disability type within programs and majors for all types of disabilities. Our analyses were limited to examinees who took the GRE in the computerized multistage test format, which excluded those who were administered the test in paper format or those who took the test prior to August 2011. Like the MCAT and LSAT predictive validity studies we reviewed earlier (Amodeo et al., 2009; Searcy et al., 2015), we were unable to collect information on whether GRE examinees received accommodations for their disabilities in graduate school. If it were the case that our sample comprised institutions that offer more resources and better services to students with disabilities, then our results would not generalize to institutions that do not provide sufficient resources. Further, we were not able to control for potential selection bias; the applicant pool may differ between large and small schools, or schools that strongly support students with disabilities relative to those who do not. These potential differences in the applicant pool or availability of resources in graduate school could have systematically influenced our differential prediction results in ways we could not control for. These limitations underscore a need in the field to study the extent to which students with disabilities are receiving supports, including accommodations and other services that would help their learning activities in graduate and professional schools and how these resources and services vary across institutions. This would not only help provide information about the validity of criterion measures (e.g., FYGPA) but would also shed light on learning opportunities for graduate students with disabilities.
Implications
Despite the limitations, the findings in this study converge with trends observed in other predictive validity studies for postsecondary students with disabilities (e.g., Braun et al., 1986; Searcy et al., 2015; Sireci & Talento-Miller, 2006). These findings provide empirical evidence that is useful to practitioners who work as part of the graduate admissions process and could ease their concerns regarding differential prediction when using GRE scores to assist with admission decisions, especially in relation to students with disabilities.
In addition, the current findings also revealed that though there appear to be significant differences in GRE section scores between students with and without disabilities, such differences in the UGPA or graduate GPA appear to be minimal and not significant, which could be related to the fact the college and graduate school are continuously challenged by the issues associated with grade inflation, as both UGPA and FYGPA have a mean close to or above 3.50, with many close to 4 (the maximum of the scale). Such issues may have contributed to the limited prediction power of graduate GPA using GRE scores and would be hard to adjust using statistical procedures that are currently available.
Supplemental Material
EC872688_Appendix – Supplemental material for The GRE and Students With Disabilities: A Validity Study at 10 Universities
Supplemental material, EC872688_Appendix for The GRE and Students With Disabilities: A Validity Study at 10 Universities by Guangming Ling, Heather Buzick and Vinetha Belur in Exceptional Children
Footnotes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
