Abstract
This study investigates how the college readiness of participants in a compensatory program designed to facilitate interest in science and engineering was determined. Archival data were used to qualitatively analyze the performance reports of 205 student participants during the compensatory program’s first 5 years. Findings indicate participants were evaluated favorably for maintaining a positive disposition toward coursework regardless of their actual numeric scores. Consequently, many participants, whose numeric scores made them less viable candidates, were recommended for admission to college. Because 95% of the students participating in the program were African American, this article highlights how context-specific evaluation can reduce biases encountered by this population when colleges rely on traditional measures of achievement to determine college readiness.
Keywords
Although regularly used in college admissions, critiques of standardized tests cast doubt upon their fairness in this process. Of particular concern is whether and how these instruments are biased toward African American test takers. When discussing possible sources of bias, Gould (1995) differentiated between cultural and statistical ones. Statistical bias focuses on those instances when two individuals receive the same test score, yet perform differently in some other area, such as college grade point average (GPA). Buttressing such concerns about statistical validity are data indicating that, despite the import attributed to them, standardized test scores account for a relatively small proportion of the variance in freshman grades (Culpepper & Davenport, 2009; Kuh, Cruce, Shoup, Kinzie, & Gonyea, 2008; Morgan, 1989; Ramist, Lewis, & McCamley-Jenkins, 1994). Defined as the tendency for one group to perform consistently lower than some reference group, cultural bias in standardized testing leads African American test takers to score worse than their White counterparts (Gould, 1995). This cultural bias exists because standardized tests privilege the knowledge of the White middle class and do not present items equally familiar or unfamiliar to African American and White test takers (Cushman, 2009; Grodsky, Warren, & Felts, 2008; Johnson, 2000).
While concerns of bias apply to all standardized tests, because of the weight given to it in the college admissions process, many scholars have tried to determine whether and how bias operates specifically within the Scholastic Aptitude Test (SAT). A series of studies has linked differences in verbal performance among African American and White test takers to cultural bias. A consistent theme across this research is that Whites perform better on easy verbal test items, while African Americans perform better on the difficult items (Freedle, 2003; Freedle & Kostin, 1990, 1997; Freedle, Kostin, & Schwartz, 1987). Freedle and Kostin (1990) argued that cultural groups assign their own connotations to common or easy words while relying on the dictionary definition or denotation for less common, more difficult words. Scoring well on easy verbal test items, the authors argue, reflects familiarity with the dominant culture, whereas scoring well on the difficult verbal items reflects academic skill and engagement.
In addition to pointing out bias within the test, scholars have also suggested ways to remedy it. Echoing the concerns raised by others, Carlton and Harris (1992) argued that African American test takers would benefit from the inclusion of questions that relied less on cultural familiarity and more on textbook-like abstractions. Other scholars have demonstrated the potential benefits of computer adaptive tests for improving African American standardized test performance (Gallagher, Bridgeman, & Cahalan, 2002; Schaeffer et al., 1998). Instead of focusing on the type of questions or the mode of administration, Freedle (2003) and Santelices and Wilson (2010) suggested that African American test takers would benefit most from revising the way the SAT is scored, giving more weight to the more difficult test items.
While each method aims to reduce the inequities African American test takers face, they continue to value standardized testing as the primary method to assess student readiness for higher education. This is curious for two reasons. First, Santelices and Wilson (2010) replicated Freedle’s (2003) finding that the SAT functions differently for African American and White test takers. This confirmed that Freedle’s finding represented a consistent pattern of bias against African Americans within the SAT and was not a fluke or the result of a statistical anomaly within his 2003 study. Second, the Educational Testing Service continues to dispute claims of bias within the SAT, and attempts to redesign the test have not addressed the core issue of bias against African American test takers (Santelices & Wilson, 2010). This led Santelices and Wilson (2010) to argue “the confirmation of unfair test results throws into question the validity of the test and, consequently, all decisions based on its results” (p. 126).
Standardization in and of itself is a problematic ideal. Even if the tests were redesigned so as to eliminate all statistical and cultural biases, the very nature of standardization suggests that these tests will be useful in assessing student readiness for higher education in vastly different contexts. According to the College Board, the SAT measures the skills students need for success in the 21st century (College Board, 2012). Perhaps this sort of instrument would be sufficient if, after testing, students did not then opt for different academic settings. Presumably, each setting requires the mastery of context-specific skills if students hope to meet with success.
In addition to the cultural and statistical critiques of standardized tests, once we consider the differences across higher education settings, it becomes easier to understand why such instruments would offer little in the way of predictive value for college performance (Strauss & Volkwein, 2002). There are more than 2,700 four-year colleges in the United States (National Center for Education Statistics, 2011a) whose missions range from training students in the liberal arts to professional disciplines. Some of the schools are small, while others have more than 40,000 students on campus each year. Some are located in the middle of bustling urban metropolises and others are situated within isolated rural regions. One test could hardly predict student success in such vastly different contexts. Yet what alternatives to determine student readiness for post-secondary education exist? How do these alternatives account for the specificity of higher education settings? And, how might African American students benefit from their use?
Alternative Means of Assessing College Readiness
There are more than 2,800 programs designed to aid students in the high school to college transition (Trio Quick Facts, 2011). Known as compensatory transitional or bridge programs, each strives to enable students from disadvantaged educational backgrounds to gain the skills necessary to seek a higher education and to succeed in college (Domina, 2009; Domina & Ruzek, 2012; McElroy & Armesto, 1998). The programs are typically crafted around a set of academic skills (e.g., English or algebra) and provide students with access to information on the college application process, obtaining financial aid and academic counseling (Tsui, 2007).
Scholars evaluate compensatory transition programs in relation to their participants’ enrollment and future success in college (Domina, 2009; Pascarella & Terenzini, 2005). The best-designed compensatory programs, particularly those focused on science and engineering, take an integrated approach by encouraging participants to build a peer support network, providing participants with personal attention from faculty members and offering participants a bridge to a post-secondary setting (Building Engineering and Science Talent, 2004). Although limited, the available research suggests that compensatory transitional programs are positively related to college attendance and persistence (Bergin, Cooks, & Bergin, 2007; Dynarski, Hyman, & Whitmore Schazenbach, 2011; Pascarella & Terenzini, 2005). For instance, Tierney and Jun (2001) found that, compared with a rate of 20% college attendance for students from Los Angeles’ minority neighborhoods, 60% of those who participated in a transitional program sponsored by the University of Southern California enrolled in a 4-year college. While the current research aims to understand whether participants enroll in college and how program participants fare once there (Pascarella & Terenzini, 2005; Seftor, Mamum, & Schirm, 2009), little is known about how performance in these transitional settings can serve as a basis for evaluating student readiness for higher education.
The current study argues that compensatory transitional programs are an ideal alternative to standardized testing for evaluating student readiness for college. This is particularly so for African American students, because compensatory transition programs have become an established part of the effort to recruit minority students to colleges and universities (Ackerman, 1991; Bergin et al., 2007; Tierney & Jun, 2001). These programs provide faculty and staff with an opportunity to observe student behavior within a particular college or university context. Whereas standardized testing yields information regarding test takers’ mastery of generic skills, a compensatory transitional program has the potential to provide information about participants’ mastery of the context-specific skills expected and required of those attending the host university. This type of evaluation can benefit African American students because it would make possible the observation of skills linked to this particular population’s readiness for higher education.
When assessing African American students, non-cognitive factors that reflect experiential or contextual intelligence have been shown to better predict academic performance than do standardized test scores. Scholars have produced a body of research indicating that characteristics such as confidence, realistic self-appraisal, preference for long-range goals, and leadership experience are generally more useful in predicting GPAs for African American students than traditional measures of achievement (Sedlacek, 2003, 2011; Sedlacek & Sheu, 2008; Tracey & Sedlacek, 1984, 1989). Other scholars have highlighted cultural background and contexts (Jenkins, Harburg, Weissberg, & Donnelly, 2004; Nasim, Roberts, Harrell, & Young, 2005), psychosocial factors (Powell & Jacob Arriola, 2003), compatibility between teaching and learning styles (Rovai, Gallien, & Wighting, 2005), racial identity (Chavous et al., 2003), attachment and engagement (Kirkpatrick Johnson, Crosnoe, & Elder, 2001), gender (Arlin Mickelson & Greene, 2006; Thompson, Gorin, Obeidat, & Chen, 2006), and a sense of belonging (Hurtado et al., 2007) as playing key roles in predicting African American student performance and retention.
While the instruments and procedures used to assess non-cognitive measures of readiness include questionnaires, interviews, and essays (Sedlacek, 2003), observing a student’s non-cognitive ability is also important. Because they provide faculty and staff with an opportunity to observe participant behavior, compensatory transitional programs are uniquely situated to evaluate non-cognitive measures of readiness. In particular, interactions among students and faculty within these programs provide a means to assess a student’s approach toward their studies, another non-cognitive factor known to influence academic performance and achievement.
Carbonaro (2005) found that, independent of academic ability and prior achievement, the amount of time and energy or effort a student expended on schoolwork positively influenced academic performance. Similarly, Rua and Durand (2000) used the concept of “academic ethic” to articulate the process by which students who devote long, regular hours to their studies achieve success. Despite often feeling alienated or less attached to school, African American students consistently value education (Harris, 2006) and put forth the types of effort related to doing well academically. For instance, Kirkpatrick Johnson et al. (2001) found that African American students felt less connected to their schools but reported higher levels of engagement with school (e.g., skipped class fewer times, had less trouble getting homework done) than their White counterparts. And Harris and Robinson (2007) found it was prior skills, not differences in behaviors related to effort (e.g., doing homework, being attentive, working hard), that led African American students to experience lower levels of academic achievement than their White peers. In light of such findings, this article investigates whether and how context-specific evaluation within a compensatory transitional program benefits African American students.
Empirical Setting and Methods
The names of the organizations and individuals referenced herein have been altered to preserve anonymity. Summer Program (SP) is a compensatory transitional effort sponsored by Technical College (TC) since 1984. TC is located in Midwestern, a town with approximately 103,000 residents (U.S. Census Bureau, 2010). Students attending TC major in engineering (e.g., mechanical, electrical), pure and applied sciences (e.g., mathematics, chemistry), and science-based management disciplines (e.g., operations, information systems). Since its inception, SP (1984) has sought to “aid participating [minority] students in developing basic academic skills and knowledge required to be successful in [college] engineering and management programs, and to help them change affective characteristics which may have hindered their education previously,” by inviting them to attend a 6-week residential experience on TC’s campus during the summer between their junior and senior year of high school (p. 1).
SP is a local effort established to encourage minority students to enter science and engineering based fields, because Midwestern’s home state lags behind the national numbers for minority Bachelor of Science degree recipients (National Center for Education Statistics, 2011b). Furthermore, although they are just as likely as their White peers to enroll in science programs, racial and ethnic minority group members are less likely to persist in these programs (Anderson & Kim, 2006; Bonous-Hammarth, 2000; Chubin & Babco, 2003; Goodchild, 2004; Hurtado et al., 2008). As a result, nationwide, African, Hispanic, and Native American students account for a small proportion of science, technology, and engineering degree recipients. In 2009, each group earned 7.5%, 7%, and 0.6% of science, technology, and engineering degrees, respectively (National Center for Education Statistics, 2011b). The percentage of Bachelor of Science degree recipients actually decreased, from 8.2% to 7.5%, for African Americans between 2001 and 2009 (National Center for Education Statistics, 2011b).
To be eligible to participate in SP, a student must have completed their junior year of high school; maintained a 2.8 GPA or better in math, science, and English courses; taken 2 years of high school algebra and English and 1 year of geometry and chemistry; and expressed an interest in engineering or management (SP, 1984). After submitting an application to the program, eligible students are then interviewed by SP’s corporate sponsors for selection into the program. An average of 36 high school students is selected to participate in the tuition free program each summer. Although the program has expanded to include Hispanic and Native American students, 95% of the students classify themselves as Black or African American. All students take the same college-level courses in calculus, chemistry, computer programming, and communications during SP (SP, 1984). In addition to these topical courses, students are also required to attend a study skills course focused on adolescent development, building support groups, time management, and study techniques to facilitate a transition to the TC community and SP program (SP, 1984).
Participants who do well in SP are offered admission to TC, some even receiving scholarship aid based not on standardized test scores but on performance during the transitional program. Therefore, given that 95% of the program’s participants identify as African American, how SP evaluates participants is critical to understanding African American student recruitment and retention at TC. According to one staff member, since 2000, SP has become TC’s primary, if not the only, recruitment channel for African American and other racial/ethnic minority students, making an investigation of assessment all the more important (SP staff member, personal communication, July 21, 2010).
Although SP does not track the number of its participants who apply to TC, the program does track those students who go on to attend TC. When data collection for this project began in 2010, the most recent figures available on SP participant matriculation and graduation were from the year 2003. Between 1984 and 2003, 442 of the 711 SP participants (62%) were admitted to TC. Of these 442 participants, 238, or roughly half, enrolled. Thus, between 1984 and 2003, on average one third of all SP participants enrolled at TC. Even more striking, among the 238 SP participants who enrolled at TC, 68% persisted and graduated with Bachelor of Science degrees in engineering, science, or management. Among 4-year not-for-profit, private colleges like TC, 65% of all students earned a bachelor’s degree within 6 years (National Center for Education Statistics, 2011c). Among African American students, 17% earn a bachelor’s degree within 6 years (National Center for Education Statistics, 2011c). Thus, the SP participant figures outpace national retention averages among private schools and African American students.
This study uses archival program data to investigate context-specific evaluation. SP compiled a summary report of each student’s performance during the summer session. At the conclusion of each session, this report was forwarded to TC’s admission office. In a one-paragraph statement, based on their opinion and experience interacting with them, faculty documented each SP participant’s strengths and weaknesses in each subject area and final grades, and their overall thoughts on the student’s likelihood of succeeding at TC. Although free to comment on whatever element of a participant’s performance they preferred, as the analyses ahead elaborate, faculty comments focused on a participant’s ability to grasp key concepts and participant attitude, including those times when faculty indicated that many participants needed to work on their “notation.” In the case of mathematics, the meaning of an equation changes drastically depending upon the accompanying notation, and the notational differences between one function and another are often quiet subtle. Alternatively, faculty chose to comment on participant attitude, noting participants’ positive or negative disposition toward coursework.
Between 1984 and 1989, the report included both quantitative and qualitative evaluations in the form of numeric course scores and written comments for 205 students. These evaluations served as the basis for the analyses. Empirical and theoretical reasons exist to restrict the analyses to the 1984-1989 time period. Empirically, while the written evaluations were consistently included in the reports through 2010, the numeric scores were reported with less consistency following the 1989 session. Focusing on the performance reports from the first 5 years helps ensure the comparability of the data (Singleton & Straits, 2009). Theoretically, the first years in an organization’s existence are a critical time in which the norms and routines that will have a lasting effect emerge (Baum, 1996; Lawrence & Suddaby, 2006; Stinchcombe, 1965; Zucker, 1977). In the context of SP, the evaluation of student participants in the first 5 years established the dimensions of performance faculty and staff would care about and use to determine a participant’s readiness for college. Thus, investigating how participants were evaluated between 1984 and 1989 is paramount to understanding whether and how SP developed a mechanism to assess student readiness that stood in contrast to those that relied primarily on traditional achievement measures.
When data collection began in 2010, three of the four faculty members who taught SP courses between 1984 and 1989 continued to teach and to advise participants. Of these three, two identify as Black or African American. Throughout their careers, SP faculty have been acknowledged for their research, teaching, and service to the TC community. Their accolades include outstanding teacher awards, recognition for excellence in pedagogical research and scholarship, endowed chairs, and campus leadership and service awards. Likewise, the three staff responsible for overseeing SP’s day-to-day operations, all African American and with the program since the 1990s, have received national recognition and awards for their efforts to diversify the engineering pipeline. The stability of both SP faculty and staff provided an opportunity to utilize multiple data collection techniques, all geared toward understanding how SP evaluated its participants. One week of classroom sessions were observed in both July 2010 and 2011. Classroom observations took place with the three faculty members who had been with the program since the 1980s. In addition, the three faculty members and the three staff members who had been with the program since the early 1990s were interviewed in July 2010. Combining several data collection techniques that did not share the same weaknesses (i.e., triangulation) increased the confidence in the research findings presented here (Singleton & Straits, 2009).
The Dimensions of Participant Performance
Table 1 details SP participant performance between 1984 and 1989. Based on a 100-point scale, student performance across the four subject areas averaged a score of 81. Looking at the individual course scores provides further information. Compared with the communications course where participants averaged a score of 91, the technical course averages of 72, 81, and 77, for chemistry, computers, and math, respectively, were much lower. Given that these were high school students attempting college-level coursework, one could easily argue that these scores suggest promise on the part of the SP participants. Yet, given TC’s rigorous scoring standards, an admissions officer could develop a different interpretation. According to TC’s academic policies at the time, students had to maintain a combined average of 77 to remain in good standing (TC, 1984). If converting to a 4.0 scale, a score of 77 was the lowest score equivalent to a 2.0 or a C and a score of 76 was equivalent to a D+ (TC, 1984). Therefore, the average SP chemistry, computers, and math scores of 72, 81, and 77 translated to 1.0 (D), 2.5 (C+), and 2.0 (C) when converted to the 4.0 scale used by TC. Thus, the average SP participant’s numeric scores would not stand out as particularly impressive to an admissions officer reviewing performance reports.
Average SP Course Scores, 1984-1989.
Note. SP = Summer Program.
Assessing the Whole Participant
Fortunately for SP participants, they were evaluated not only on their ability to grasp the challenging academic material presented to them during the session but also with respect to their ethic, particularly their willingness to put forth effort toward doing well, regardless of how well they actually did. The analyses revealed that when evaluating, faculty assess participant performance on two dimensions—coursework and ethic.
Figure 1 depicts the relationship between these two dimensions. Dependent on the interaction of their coursework and ethic, SP faculty assessments of participants can be classified according to one of the four quadrants. Faculty often evaluated participants in relation to the average, writing comments like “ranked far below average” or “performed average.” In view of this, the top row of Figure 1 was reserved for those participants who performed at or above the session average. Alternatively, participants who fared less well, scoring below the session average, were placed in the bottom row of Figure 1. Unlike the numeric scores that ranked participants in relation to the session average, the ethic evaluations were more nuanced character assessments in which faculty dissected participant behavior based not merely on numeric performance but on the participant’s disposition toward coursework. For this reason, the left side of Figure 1 was reserved for those participants whose ethic went unquestioned, while the right side was for those participants whose ethic was questioned by one or more faculty.

SP evaluation quadrants and sample quotes for each.
Figure 1 also provides a sense of the type of comments faculty made about a participant’s ethic or commitment to studies within each evaluation quadrant. When focusing on participants who did well (top row), faculty comments indicated participants in Quadrant 2 had superior academic skills but lacked the engagement of their peers in Quadrant 1. Comparatively, when focusing on those who did poorly (bottom row), faculty comments indicated participants in Quadrant 3 worked hard and remained engaged despite not achieving high marks, whereas those in Quadrant 4 lacked the expected engagement.
Focusing on a few participants will highlight the differential evaluations accorded to individuals within each quadrant. Figure 2 contains the names of 10 students, the year each participated within SP, and their corresponding evaluation quadrant. Participant performance was easiest to interpret when it fell in either Quadrant 1 or Quadrant 4. In both scenarios, a participant’s course performance aligned with the professors’ assessment of their ethic.

Sample SP participants and their performance quadrants.
Quadrant 4
Marsha’s scores indicate that she struggled throughout the 1987 SP session. Yet, her poor course performance was not the only concern noted in her final evaluation. Faculty perceive students who do not participate in class, seek extra help, or avoid distractions as less engaged (Kirkpatrick Johnson et al., 2001). This appears to have been the case with Marsha. One professor felt she “did not apply herself” and that “although she was performing poorly, she did not seek any help from me.” Another commented that her “work tended to be hurried,” while yet another felt she “was the most talkative student in class.” Because of these assessments, Marsha ranked squarely in Quadrant 4. Faculty felt that her poor academic performance was matched and perhaps exceeded by her poor ethic, leading each to feel that “she did not have any interest in what was going on.”
Quadrant 1
Likewise, a student’s performance was easy to assess when it fell into Quadrant 1. While the faculty remained unimpressed with Marsha’s abilities during the 1987 session, they held a far more favorable view of Anna. Her overall average of 93 indicates superior academic capabilities. Yet, getting good grades alone does not indicate a good ethic. Students who view grades as a means to an end often lack the behaviors associated with a strong academic ethic (Rua & Durand, 2000). However, faculty assessments of Anna indicate a strong commitment to her studies. After noting that she had a solid background, one professor stated, “This is a very good student. She is a little bit shy but her work is extremely good.” Another professor claimed she “was one of the most serious students” and that she sought “assistance when necessary.” One professor was even more direct in assessing Anna, stating outright, “I would recommend her for admission [to TC].”
Quadrant 3
Because their coursework and ethic were not in synch, not all participants were as easy to classify. This was perhaps most true in the case of those participants who clearly struggled with coursework but displayed a greater capacity in the ethical skills faculty believed important for success at TC. When evaluating these cases, the faculty adopted a developmental approach (Cuban, 1993; Jarvey, McKeough, & Pyryt, 2008; Metz, 1978; Votruba-Drazl, Li-Grining, & Maldonado-Carreno, 2008) and sought to reward those students who were less academically prepared but maintained a positive attitude toward school and learning (Kelly, 2008; Lortie, 1975). In the case of SP, this is even more important, because Manswell Butty (2001) did find that, among African American students, a positive attitude toward math influenced achievement.
In the case of SP, if faculty classified participant performance based only on coursework, many SP participants would not be seen as viable TC students. As indicated by Table 2, despite being characterized by faculty as having put forth their best effort, 22% of the participants (Quadrant 3) performed below the numeric course average between 1984 and 1989. If past performance is the best predictor of future performance, then such participants are likely to struggle should they matriculate to TC. Despite this assumption, unlike the 34% of participants in Quadrant 4, if a participant performed poorly in coursework but displayed a good ethic, all was not lost in the professors’ eyes.
Percentage of SP Participants Within Each Evaluation Quadrant.
Note. Due to rounding, percentages do not total to 100 each year. SP = Summer Program.
When evaluating SP participants, professors did not automatically equate poor performance with poor ethic. Teachers’ inability to disentangle student skill and ethic from academic performance is a problem that plagues African American students (Harris, 2006; Harris & Robinson, 2007). A benefit of context-specific evaluation is that it provides enough information that professors can separate behaviors from outcomes. If a student demonstrated a willingness to work hard despite encountering difficulties, professors looked upon this favorably and took care to highlight the ways in which a student’s ethic would serve them well going forward at TC. For instance, in 1984, Patricia’s overall score of 76 fell below the mean. Despite this, the computer professor characterized her as an “industrious student and serious about her coursework.” The professor even developed a rationale for her below average performance: “She seemed to have little background.” Nevertheless, Patricia “always worked hard,” and even referenced “her future goals in becoming a computer programmer.” In another case, despite a score of 62, Lesley, another 1984 participant, was described as a “very hard worker” and perceived as wanting “to do better.”
When assessing Kathie, a 1987 SP participant, the math, chemistry, and computer professors each noted her struggles, stating that she was a “weak” student and “had difficulty grasping concepts.” Yet, each also noted that she “had the desire” and “asked many questions.” Her ethic led the chemistry professor to conclude that, despite her struggles, “she was learning a lot of chemistry and should do well next time.” In another case, Jane, also a 1987 SP participant, “started the session slowly” and finished below the average. But according to faculty, Jane also “studied hard” and sought “help whenever she needed it.” Because of this, one professor felt “she will be successful in college,” and another “recommend[ed] her for admission.”
Quadrant 2
Although beneficial, particularly to those participants in Quadrant 3, context-specific evaluation can be problematic. Teachers reward students whom they perceive as being engaged and interested in class (Ainsworth-Darnell & Downey, 1998; Downey & Pribesh, 2004; Kelly, 2008; Yair, 2000). Because of this tendency, context-specific evaluation can shed a negative light on individuals who would have otherwise appeared exemplary had focus remained on their numeric scores. SP participants in Quadrant 2 did well in their classes but nevertheless faculty perceived these students as lacking the necessary engagement. As the following cases illustrate, doing well in a course must be accompanied by complementary displays of ethic.
When discussing Stanley, whose score was higher than the 1984 mean, the computer professor simply stated, “Did well in class but did not hear much from him.” In Marcus’ case, also a 1984 participant, despite having one of the highest overall averages, one professor summarized, “Excellent student. I feel he did not put a great effort in the class. In any event, he ranked third in the class.” At other times, professors remarked that participants like Stanley and Marcus adopted a “know-it-all mentality” and did not understand that “seeing something before does not necessarily mean one has the necessary understanding.” Like Stanley and Marcus, Eli also scored above the mean during the 1984 session. However, Eli received praise because he “asked questions” and “became discouraged at times but never turned back.” This suggests that participants perceived as academically gifted may be penalized if they do not display the ethic valued by SP faculty.
The 1988 SP Session
Context-specific evaluation can also be problematic in those cases where faculty opinions differ, as illustrated by evaluations from the 1988 SP session. Table 2 suggests that the faculty as a whole found most participants from this session deficient with regard to ethic. In some cases this was true. For instance, one faculty member felt Ernie had “an attitude problem.” Another complained that Ernie had potential, “but he did not take advantage of the sources available,” while another felt Ernie was simply “disinterested in the class.” Although each complained about a different dimension of Ernie’s ethic, collectively the faculty felt this participant lacked the ethic necessary to do well at TC. Yet, during the 1988 session, most of the participants in Quadrants 2 and 4 were classified based on one professor’s assessment alone. Time and again, the professor noted that participants “began the term with some motivation that disappeared around midterm.” To other faculty members, these same participants appeared “very serious” and “very attentive” throughout the session. If ethical assessments are going to be a part of the admission process, this disparity suggests the need to consider how majority and minority opinions will be weighted.
Discussion and Conclusion
While on campus, SP participants are immersed in a science-based curriculum. Furthermore, through the use of laboratory experience and site visits to local corporations, SP provides participants with the opportunity to understand the real-world applications of material presented in the classroom (SP, 1984). This is precisely the type of program structure scholarly research has shown to foster student engagement and motivate students to become involved in the learning process (Weissman & Boning, 2003). Consequently, when students do not appear to be engaged, SP faculty take note. In terms of programming, SP has established an instructional context that should motivate students to take on a learning-focused orientation (Young, 2003), yet not all students adopt this stance. SP faculty place value on those behaviors that indicate a participant’s commitment to utilizing a variety of strategies—discussing ideas with others, speaking up in class, seeking help when needed, and so on—that reflects an academic ethic of deep engagement with learning (Laird, Shoup, Kuh, & Schwarz, 2008).
As faculty reflect on whether SP participants possess the behavioral skills necessary to do well at TC, they perform a similar role as the Education Testing Service and other organizations sponsoring standardized tests. As gatekeepers, each has the power to decide who will and will not have access to higher education. However, the way in which SP goes about differentiating students is significant. The analyses indicate that faculty assess participant performance in courses, a more traditional measure of achievement, as well ethic, a less traditional measure of readiness for higher education.
Doing well academically is important. Indeed, of the 46 SP participants from the 1984-1989 sessions who went on to graduate from TC, 41% were from Quadrant 1. While important, this study suggests that numeric course performance is not the only valid means of determining whether a student has the capacity to do well in a post-secondary setting. Despite struggling with coursework in the transitional program, 22% of SP participants who went on to graduate from TC were from Quadrant 3. These students ranked below the numeric average while in SP but displayed exceptional academic ethic. Interestingly, a comparable percentage (24%) of SP participants who went on to graduate from TC came from Quadrant 2. Although further study is required, this hints at the possibility that, under the right circumstances, academic ethic is a skill that may substitute for academic preparedness.
Based on their experience, SP faculty and staff know the characteristics necessary to do well at TC. They look for participants whose academic skill is accompanied by a strong academic ethic. Their experience also enables them to identify students who may be less academically prepared for TC, but nevertheless possess other, more difficult to quantify, characteristics that will enable them to do well in college. In the words of one staff member,
If a student ranks low in the program, has low test scores and GPA, our office will not recommend that they come, but each student is looked at on an individual basis. If they really want to come, and we think they can make it, we will work with admissions to get them accepted. (SP staff member, personal communication, June 6, 2012)
This sentiment echoes Seymour and Hewitt’s (1997) finding that those students who switch from a science-based to a non-science-based major are no different in terms of their ability. Instead, when compared with those who remained, those who left a science-based major failed to make effective use of the academic resources available to them and could not find ways to tolerate the challenges of a science-based curriculum (Seymour & Hewitt, 1997). The most widely used standardized testing instruments (e.g., ACT, SAT) do not allow for this type of nuanced assessment of the student as a whole. Such instruments claim to measure a student’s academic readiness for college, yet by design, these instruments are incapable of assessing characteristics like persistence, inquisitiveness, and emotional maturity, even though these factors are just as likely to bear upon a student’s readiness for college (Sedlacek, 2003, 2011). Redesigned standardized tests (Freedle, 2003; Santelices & Wilson, 2010) will never yield this type of information. SP enables decision makers to obtain more relevant and context-specific information about students hoping to attend TC.
This is not to suggest that context-specific evaluation is not without its own problems. As the analyses revealed, those participants in Quadrant 2 who performed well in their courses were evaluated negatively because the faculty felt these students lacked a strong ethic. Furthermore, the fact that 13% of the SP participants between 1984 and 1989 who went on to graduate from TC were from Quadrant 4 is perplexing and suggests that further study is needed to identify the factors that enable such students to persist in college despite poor course and ethical performance in the transitional program.
The question remains as to whether and how context-specific evaluation can be used by other institutions. There are many reasons why colleges and universities would prefer to use standardized testing, compared with context-specific evaluation, as the main form of evaluation in the admissions process. The costs associated with the program and faculty commitment are just two of the factors that may give many institutions pause before implementing a context-specific form of evaluation. In SP’s first year, sponsorship entailed a US$1,200 fee to cover the cost for one student, and by the program’s fifth year, the cost had reached US$2,700. In 2010, sponsors paid US$4,500 to cover the cost for one student to attend SP. Relying on private sponsors is not the only alternative, because the federal government funds transitional programs, awarding more than 800 grants to programs in the 2012 fiscal year (Upward Bound Program, 2013). Yet, faculty and staff must commit time and resources to navigating the federal granting process and model the transitional program so that it fits within the government’s program requirements (Federal Register, 2011). Alternatively, colleges and universities are largely removed from the cost associated with standardized testing. The burden of paying for these tests is shouldered by students and their families. Registering to take the SAT costs US$50 (College Board, 2012) and preparation courses for this test range from US$300 to US$1,000 (Kaplan Test Prep, 2012).
Even if an institution is able to secure the funding from public or private sources, finding faculty to oversee the process associated with context-specific evaluation presents another challenge. In the case of SP, having faculty who teach at TC is critical in discerning whether or not program participants have the skills necessary to succeed in this particular context. This may be the most prohibitive factor of all. Given the demands on faculty to produce research (Greene et al., 2008; Kain, 2006), colleges and universities may find it difficult to find tenured or tenure-track faculty members willing to dedicate their summers to teaching within a transitional program.
Despite the growing trend toward hiring adjunct faculty on college campuses (Jacoby, 2006), using adjuncts to staff transitional programs presents its own set of problems. Their part-time status and lack of job security encourage adjunct faculty to utilize less challenging methods of instruction (Benjamin, 2002). More worrisome are the findings indicating that students taking a high percentage of courses with adjunct faculty are less likely to persist in attaining a degree (Ehrenberg & Zhang, 2004; Harrington & Schibik, 2001). The type of student advising, developmental education, and counseling that makes SP a success is not likely to occur if adjunct faculty are primarily responsible for teaching within the program (Jacoby, 2006).
Scholars and practitioners alike should devote more attention to the staffing and funding of transitional programs. In the 2012 fiscal year, the federal government allocated roughly US$270,000,000 for Upward Bound, perhaps the most well-known transitional program in use, and estimated the average cost per student at US$4,302 (Upward Bound Program, 2013). Government reports on Upward Bound devote more attention to student outcomes than to how successful programs allocate their budgets or staff their offices (Calahan & Curtin, 2004). The focus on students is understandable, yet it is essential to expand the scope of our interest to the mechanics of operating transitional programs so as to further explicate how schools might implement a successful program like SP.
As stated earlier, the college completion rates of African American SP participants far exceed the national averages for this population. Accordingly, context-specific evaluation geared toward assessing academic ethic belongs in the repertoire of tools used to select and retain African American college students. Because they are more likely to attend high schools that do not offer advanced courses (Adelman, 1998), African American students are often less academically prepared than others, particularly in subject areas such as mathematics (Baranchik & Cherkas, 2002; Hallinan, 2001; Kelly, 2009) that are essential to doing well on standardized tests. Further complicating matters, several studies suggest that, while college attendance and completion rates appear to be positively influenced by participating in compensatory programs, these programs have little effect on participants’ standardized test scores (Dynarski et al., 2011; McLure & Child, 1998; Seftor et al., 2009).
Such findings make it all the more essential to utilize selection mechanisms that will yield relevant information regarding the likelihood that African American students can persist in college. The statistical and cultural biases inherent in standardized tests make them inappropriate tools for determining the college readiness of African American students (Freedle & Kostin, 1990, 1997; Gould, 1995). Incorporating mechanisms to evaluate African American students’ ethic is crucial because it provides decision makers with relevant data on the non-cognitive factors known to better predict African American academic success (Sedlacek, 2011; Tracey & Sedlacek, 1989). In the case of SP, having access to such information encourages decision makers to overlook academic shortcomings resulting from inadequate preparation. Other colleges and universities should follow suit.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
