Abstract
One of the most critical problems that hospitality firms face in selecting employees is to ensure that any employment tests the employer uses are valid and do not screen out minorities. For example, the use of cognitive ability tests often leads to subgroup differences between majority and minority group members. Such a discrepancy opens an employer to charges of adverse impact (against minorities), and employers often have adjusted (or otherwise disregarded) test scores to avoid potential adverse impact and give minorities an even-handed opportunity for employment or promotion. The practice of adjusting test scores in this way was set aside in 2009 by the U.S. Supreme Court, in Ricci v. DeStefano, in which the city of New Haven, Connecticut, attempted to avoid adverse impact by disregarding test results. The court said this amounted to discrimination against the majority group members who did well on the test. The court’s holding means that if a test creates apparent adverse impact, in the absence of strong basis in evidence for disregarding the test scores, the employer may face the awkward choice of being sued for adverse impact or disparate treatment, depending on how it treats the test. The implications of Ricci v. DeStefano for hospitality employers include ensuring that jobs are correctly analyzed before any test is given and that multiple forms of various valid types of test are used to select job candidates.
To hire the best employees, many organizations use selection tools (e.g., tests of cognitive ability) that are intended to predict future job performance. Although many selection tools have been found to be good predictors of job performance, these tools also produce varying degrees of subgroup differences (Pyburn, Ployhart, and Kravitz 2008). Tests of cognitive ability are one of the best predictors of job performance, but also have been found to produce large subgroup differences between racial or ethnic majority and minority group members (Outtz 2002; Sackett et al. 2001). Such an outcome, when minority group members tend to score lower than majority group members, is potentially in violation of the provisions of the U.S. Civil Rights Act of 1964. However, as we explain in this article, a holding of the U.S. Supreme Court has indicated that employers seeking to avoid such a violation generally may not disregard test outcomes without good reason.
The law notwithstanding, most organizations have set a goal of hiring a diverse workforce. As a result of group differences in test scores, however, organizations often face a trade-off between the validity of the selection procedure and the goals of hiring a diverse workforce. This has been termed the diversity–validity dilemma (Ployhart and Holtz 2008). Although not intentional, these subgroup differences in test performance can lead to discrimination claims, which can lead to legal action, with all its negative outcomes (Goldman et al. 2006).
The recent U.S. Supreme Court’s decision in Ricci v. DeStefano (2009) was the first decision in decades to deal with the issues relating to employment testing and discrimination claims under Title VII of the Civil Rights Act of 1964. The case involved eighteen firefighters who passed two exams for promotion to lieutenant and captain positions. This group sued the City of New Haven, Connecticut, for refusing to use the test results and make the promotions, when the city determined that the tests had adverse impact on minority employees (white candidates scored higher than African American candidates). The U.S. Supreme Court decided in favor of the white firefighters. The five–four Ricci decision means that the results of a valid, job-related test cannot be thrown out simply because the test may result in adverse impact. Instead, the court set a standard of a “strong basis in the evidence” for an employer’s rejection of test scores. In this article, we suggest two outcomes of the court’s decision. First, we see a more stringent standard of “strong basis in evidence.” This applies when an employer decides to not certify test results because it believes the statistical differential in pass and fail rates for racial or ethnic minorities may lead to claims of adverse impact under Title VII. Second, the decision constrains an employer’s actions with regard to certifying a test once the test has been given, even in light of evidence that the test produced adverse impact.
Our purpose here is twofold: first, to examine the implications of Ricci v. DeStefano for hospitality employers and examine various alternative ways to score the test, and, second, to provide employers with recommendations for testing, based on the large body of applicable research. Our recommendations for testing derive from a series of meta-analyses and research. Although some of the recommendations might seem elementary, we focus on them in large part because the testing company that developed the test used in Ricci failed to follow these basic, but important, steps. This failure to translate human resources management research findings into practice is not uncommon (Rynes, Giluk, and Brown 2007).
Although the Ricci v. DeStefano case is not from the hospitality industry, it has strong implications for hospitality operators, particularly given that the industry employs a wide diversity of people. Black and Latino employees make up 52.3 percent of hourly employees and 48.5 percent of line-level supervisors in lodging (Jackson and DeFranco 2005). The more stringent standard of “strong basis in evidence” means that hospitality employers would have to find reasons other than adverse impact if they want to set aside a test. To avoid this, hospitality firms will have to ensure in advance that any test or requirement related to employment decisions will be free from adverse impact. As we explain below, the weighting and scoring schemes used to make the decisions will be set in concrete once the tests are distributed and completed.
For these reasons, in line with the goals set out by Rynes, Giluk, and Brown (2007) to disseminate research findings to practitioners, we offer recommendations for hospitality management practice based on our review of the Ricci case. We are unaware of a published paper in the hospitality management literature that has examined the Ricci case together with the implications for hospitality employers.
Summary of the Facts in Ricci v. DeStefano
In Ricci v. DeStefano, the city of New Haven hired a company to design and administer a job-related, statistically valid test that would be given to firefighters who wanted to be promoted to lieutenant or to captain. The test was composed of multiple-choice and oral-interview items. Under the contract between the city and the firefighter’s union, the multiple-choice portion was given a 60 percent weight and the oral interview was given a 40 percent weight. Also by contract, the cutoff score for promotion eligibility was set at 70 percent. Despite attempts to ensure that the test would not favor white candidates, it in fact produced a statistically significant disparity between the white candidates and their African American and Hispanic coworkers, as shown in Exhibit 1. Under Equal Employment Opportunity Commission (EEOC) rules, the disparity alone created the legal presumption of adverse impact. With both sides threatening suit, the city first thought to adjust the test results and then discarded them, citing the race-based statistics as its justification (to avoid adverse impact). The white firefighters then sued the city on the basis of disparate treatment discrimination.
Test Outcomes, Demonstrating the “Four-Fifth” Rule Results Using 60 Percent for the Written Component and 40 Percent for the Oral Component
Note: Adverse impact is generally triggered by what is called the “four-fifths rule.” When the selection rate for any protected group is less than four-fifths (or 80%) of the rate for the majority, the Equal Employment Opportunity Commission (EEOC) presumes adverse impact. This occurred at all levels in the New Haven test. For the Lieutenant test, the pass rate for whites was 58.1%, while the rate for Hispanic firefighters was 20% and for black applicants 31.6%. The 80% cutoff rate is calculated as follows. For black firefighters: 31.6%/58.1% = 54%, which is less than the 58% for whites. The calculation for Hispanic applicants is 20%/58.1% = 34.4%. The same was true for the Captain test. The 80% calculation of the 64% white pass rate exceeded the 37.5% tallied by Hispanic and black applicants. For black firefighters, the calculation was 37.5%/64% = 58.6%, and it was 20%/58.1% = 34.4% for Hispanic applicants. Data obtained from Miao (2010).
In a disparate treatment case, liability depends on whether the protected trait actually motivated the employer’s decision. In respect to Ricci, disparate treatment occurred when the city did not certify the results of the test and consequently did not promote the seventeen incumbent white firefighters and one Hispanic firefighter who passed the promotion test. The reason the city did not certify the result of the test was that none of the incumbent African American firefighters’ scores qualified for promotion—which would indicate adverse impact.
The city’s defense involved this presumption of adverse impact. It argued that the statistically significant disparity between the white candidates and the ethnic minority candidates was a valid reason to not use the test results. By that logic, there would not be disparate treatment involved in deciding not to promote the white candidates. The U.S. Supreme Court rejected this logic—contrary to lower court rulings in similar cases.
New Testing Standards from Ricci v. DeStefano
We see two implications for employers. First, as we suggested above, this case has set a higher standard for “strong basis in evidence.” The court held that statistical disparity alone cannot constitute a strong basis in the evidence for an employer’s rejection of test scores. This is a departure from past rulings, which suggested that a prima facie case was enough for a “strong basis in evidence” (Dickson 2009). For example, in Hayden v. Nassau County (1999), the county developed a test that comprised twenty-five components to fill entry-level police positions. After giving the test to the applicants and before employment decisions were made, impact analyses showed adverse impact on members of protected groups. As a consequence, the county hired experts who limited the test to results from just nine components and thereby reduced the adverse impact. The county made its hiring decisions using those nine components, and the unsuccessful white candidates claimed discrimination. The federal district court sided with the county, stating that employers should find alternative testing methods with less adverse impact.
In Bradley v. City of Lynn (2006), a cognitive-loaded test was used for entry-level firefighter positions, with an outcome of adverse impact. Expert witnesses suggested that the criterion validity of the test was questionable, and the judge concluded that cognitive ability is not the only indicator for ranking firefighter candidates. The court concluded that alternative selection tests with less or no adverse impact should have been used. Similarly, in Johnson v. City of Memphis (2006), a judge ruled that alternative, valid tests with less adverse impact should replace a cognitive-loaded test that was used for police sergeant promotions, despite the finding that the promotion test was valid and job related.
On their face, the Ricci facts appear similar to those three cases. The Ricci test validity and the weighting scheme was questioned by expert witnesses. An expert asserted that the racial disparity in the test results would have been reduced if the multiple-choice exam had been given less weight than the oral exam. Firefighter experts pointed to multiple-choice items that were not related to this particular department’s job, such as asking about equipment that was not used by the department, or asking about “uptown” and the “Second Battalion,” which do not apply to New Haven (Brodin 2011). In an amicus brief, the Society for Industrial and Organizational Psychology argued that “the lack of evidence supporting the validity of the New Haven Fire Department tests undermine their value as a selection tool” (p. 17). Despite these concerns, the U.S. Supreme Court ruled against the city.
That holding brings the second implication into focus. Once a test has been given to applicants the employers must use the tests results, even if there is evidence that the test produced adverse impact. The Supreme Court essentially held that an employer cannot use the threat of being sued for adverse impact as a legal defense for intentional discrimination directed at successful test takers. Contrary to the findings of the lower court cases, not only will employers be forced to use the results of tests in the face of adverse impact, but employers will not be able to make changes to the scoring of the tests to make them less discriminatory. In that regard, the Supreme Court explicitly rejected the idea that the city could have used the test but with different weighting of the multiple-choice and oral components, stating that changing the method of interpreting the test would have been discriminatory.
Employers with tests that do result in adverse impact will effectively face the following dilemma: they can choose to be sued for disparate treatment (i.e., intentional discrimination) or for adverse impact. If they do not certify a test to avoid a potential adverse impact case, they may face a disparate treatment or intentional discrimination suit. On the other hand, if they use the test results (and thereby avoid the disparate treatment case), they may face a potential adverse impact action. This is the diversity–validity dilemma.
Given that the Supreme Court has held that test results generally cannot be set aside or altered once the test results are available, we see two areas of research that provide alternative solutions for the diversity–validity dilemma that we just posed. The first area of research involves the use of alternative predictor constructs, such as using measures of personality (e.g., measures of conscientiousness, extraversion, or agreeableness). The second area of research examines the use of alternative predictor measurement methods, such as using interviews or assessment centers, which measure multiple constructs simultaneously (e.g., cognitive ability and personality). We explain those potential approaches in the following section.
Alternative Predictor Constructs
Alternative predictor constructs tap a single latent construct, such as conscientiousness or general cognitive ability, that is related to job performance. Considerable research offers alternative metrics that are valid predictors of job performance and produce smaller differences across ethnic groups. However, we note that these measures are not better predictors than cognitive ability. Instead, using these alternative tests along with measures of cognitive ability might be a better approach for employers than using cognitive ability measures alone. See Exhibit 2 for a summary of ethnic differences and predictive validity on job performance of the alternative predictors that we review here.
Alternative Predictors and Methods: Ethnic Differences and Predictive Validity of Job Performance
Racial-ethnic differences are based on the d statistic. Cohen’s thresholds for small, moderate, and large d statistics are 0.20, 0.50, and 0.80, respectively (Cohen 1988).
Based on Cohen’s effect sizes (Cohen 1988), validities greater than .50 are considered large, validities between .30 and .50 are medium, and validities below .30 are small.
Differences favor ethnic minorities.
Meta-analyses by Barrick and Mount (1991) and by Ones, Viswesvaran, and Schmidt (1993) of personality constructs and job performance showed that measures of conscientiousness, extraversion, agreeableness, openness to experience, and emotional stability are useful for selecting and promoting employees, particularly for customer service occupations. Not only are personality measures valid predictors of job performance, but these measures also produce smaller differences across ethnic groups (for a review, see Ployhart and Holtz 2008).
This is also true of emotional intelligence—the ability to perceive one’s emotions and the emotions of others, regulate emotions in the self and others, and use emotions to facilitate performance (Cote and Miners 2006; Mayer, Caruso, and Salovey 1999; Van Rooy and Viswesvaran 2004). Meta-analyses showed that tests of emotional intelligence are valid predictors of job performance (Van Rooy and Viswesvaran 2004) and can minimize differences in cognitive ability scores across ethnic groups (Van Rooy, Alexander, and Chockalingam 2005).
Integrity tests are also significant predictors of job performance and counterproductive behaviors, such as violent acts on the job, tardiness, dishonesty, and absenteeism, as shown in a meta-analysis by Ones, Viswesvaran, and Schmidt (1993). In a recent study, Sturman and Sherwyn (2009) found that the integrity tests were effective in detecting high-risk applicants and did not lead to adverse impact. Similarly, job knowledge tests—multiple-choice questions to evaluate technical or professional expertise and knowledge—also provide alternative predictors of job performance that produce less adverse impact than tests of cognitive ability (for a review, see the meta-analysis by Roth, Huffcutt, and Bobko 2003).
Alternative Predictor Measurement Methods
Alternative predictor measurement methods, such as using interviews, measure multiple constructs simultaneously. Also as shown in Exhibit 2, the alternative predictor measurement methods we discuss below provide employers with tools that have useful levels of validity for predicting job performance and have lower levels of ethnic group differences. Unlike tests of cognitive ability, predictor measurement methods such as structured interviews, assessment centers, and work sample tests provide a holistic perspective of applicants.
Structured interviews (i.e., using predetermined questions for every applicant) are one of the most effective predictor measurement methods. Research shows that structured interviews can be as valid and reliable predictors of job performance as are cognitive ability tests (Campion, Campion, and Hudson 1994; Campion, Palmer, and Campion 1997; Huffcutt and Arthur 1994). For example, using meta-analytic methods, Schmidt and Hunter (1998) reported a high predictive validity of .52, which was as large as the validity for cognitive ability (.51). More important, structured interviews produce an assessment of job candidates that is less open to interviewer bias (Campion, Palmer, and Campion 1997; Huffcutt and Roth 1998) and produce smaller ethnic subgroup differences in outcomes than do cognitive ability tests (Bobko, Roth, and Potosky 1999; Ployhart and Holtz 2008). Despite the positive results of structured interviews, a major limitation of structured interviews is that ethnic subgroup differences increase as the cognitive loading of the interview increases.
Another significant predictor measurement method that provides a solution for the diversity–validity dilemma is the assessment center, which is an array of standardized tests (e.g., job-related simulations, interviews, and psychological tests). Like the structured interview, meta-analyses show that assessment centers display high predictive validity, but also produce little adverse impact (Ployhart and Holtz 2008; Schmidt and Hunter 1998). Again, however, ethnic group differences increase as the cognitive load of the tests increases (Goldstein, Yusko, and Nicolopoulos 2001).
Situational judgment tests measure applicants’ judgment regarding situations that are likely to occur in the workplace (McDaniel et al. 2001). Situational judgment tests have useful levels of validity for predicting job performance, are often perceived as having face and content validity by applicants, and lead to less adverse impact than tests of cognitive ability (Chan and Schmitt 1997; McDaniel et al. 2001; Roth, Huffcutt, and Bobko 2003; Whetzel, McDaniel, and Nguyen 2008).
Work sample tests require applicants to complete a portion of the work they will be doing; that is, these tests simulate on-the-job situations (e.g., delivering a presentation, simulations, or role-plays). As such, work sample tests are valid predictors of job performance and are perceived by applicants as being highly related to the job. Meta-analytic research shows that work sample tests lead to less adverse impact than tests of cognitive ability (Roth, Bobko, and McFarland 2005).
What’s a Hospitality Employer To Do?
The U.S. Supreme Court’s decision in Ricci v. DeStefano has spawned considerable discussion and research on the diversity–validity phenomenon. As we explained above, a potential implication of the Ricci decision for all employers including those in the hospitality industry is that disparate treatment cannot be used to void a test in an attempt to avoid adverse impact. That is, a hospitality firm cannot avoid a lawsuit by intentionally throwing out test results when a disproportionate percentage of minorities do not pass the test cut-off scores. Beyond adverse impact, employers now face a more stringent requirement to show a strong basis in evidence for not certifying a test, and once a test has been given the employer must use the tests results. Consequently, employers must be proactive in employee testing, since they can no longer make changes after the test is given. We offer six proactive strategies below. Whatever strategies are used, employment tests should include valid alternative predictor constructs and predictor methods, as discussed above. For the hospitality industry, personality tests, emotional intelligence tests, structured interviews, and assessment centers seem particularly suited to employment testing, given the common belief that the best employees for hospitality jobs are those that have the right personality and attitude. Data from empirical research substantiate this line of reasoning (Tracey, Sturman, and Tews 2007).
In light of the Ricci case, it appears that a hospitality employer should take the following six steps to avoid liability in the selection and promotion contexts:
Conduct a thorough job analysis. Although the job analysis is an elementary first step for most practitioners, the failure to properly conduct an appropriate job analysis in the Ricci case underscores its importance. Although New Haven did conduct a job analysis in creating the test used in the Ricci case, it was flawed because the respondents to the job analysis questionnaire were mostly white—about 67 percent of respondents, potentially biasing the results (Dickson 2009). The job analysis should be the fundamental starting point of every human resource system (Bloom 2010). Position descriptions, selection and assessment, organizational structure, performance management, training and development, performance appraisals, career progression, and compensation should all be linked to the job analysis. Although there is not one tried and true method on how to conduct a job analysis, at the very least the employer needs to gather as much information about the job as possible. This may draw on existing information from job descriptions, training materials, and subject matter experts. However, employers who decide to conduct the job analysis need to make sure they are capturing the entire job and a representative sample of the employees, demonstrated in the Ricci decision.
Use alternative test techniques. While the EEOC clearly recognizes the role of classical test theory in the employee testing arena, the Ricci decision may make a strong case for instead using item response theory (IRT), which is a newer approach. IRT is the basis for computer adaptive testing (CAT). With CAT, the candidates are exposed to a small subset of the questions. The level of difficulty of each question received depends on the accuracy of the previous answer. Thus, the process adapts to the individual test taker. CATs are fairly accurate at measuring a candidate’s actual abilities in relation to the test item and give an employer a much better assessment of potential performance than a traditional test. CAT also reduces bias in the questions and has proven effective in measuring well-defined constructs. Consequently, CATs are better suited for measuring cognitive abilities or knowledge areas. With the advent of the internet and unproctored tests, test providers are suggesting the use of CAT with increasing frequency.
Profiling is also a relatively new testing method that identifies people who have significant potential to be successful in a particular industry or profession and in a particular workplace or culture. To create the test questions for a profile, a test provider will interview successful incumbents. This is another way of assessing individuals’ aptitudes for performing a job they’ve possibly never done before. It can also assess the individuals’ ability to adapt to an unfamiliar corporate culture. Given the hospitality industry’s service culture, the ability to identify potentially successful individuals up front will be beneficial. Several companies have proven track records of providing profiling services for the hospitality industry.
Combine types of assessments. Another step that hospitality firms can take to minimize liability for adverse impact is combining assessment tools. The city of New Haven took a step in this direction by combining the oral interview with its cognitive test, but in the end weighted its determination too heavily on the cognitive test. The resulting adverse impact against racial-ethnic minorities has also been found in research studies (e.g., Bobko, Roth, and Potosky 1999; Ployhart and Holtz 2008). This is the reason that we have emphasized the importance of combining different assessment tools, including cognitive and noncognitive tests, job performance, and structured interviews. Combining the assessment tools will help hospitality firms avoid litigation as well as gain a competitive advantage by addressing the diversity–validity dilemma.
Assume responsibility for the validation process. Hospitality firms also need to make sure that they are measuring and testing the important knowledge, skills, abilities, and other characteristics for the job, based on the job analysis. Before hospitality firms use a test, they must ask the test developer for validation. New Haven failed to do this, and the results demonstrate that the employer bears the burden of requesting evidence of validity or a validation test. This is true whether the test is objective, as in the Ricci case, or when using “subjective” tests. In certain cases, an employer does not need to validate, as follows: (1) when a test has no known adverse impact, (2) when they are simply transporting validity from another job or location, or (3) under very limited circumstances where the validity is generalized from studies of similar jobs.
Set appropriate cutoff scores. In accord with the contract between the city and the firefighter’s union, the cutoff score for promotion eligibility was set at 70 percent. While we do not question the cutoff score, we do argue that cutoff scores need to be consistent with normal expectations of proficiency in the job category. The 70 percent cutoff score in the Ricci case was set by the union rather than the city. A better approach would be to develop the cutoff score using a job-related process. For example, subject matter experts (SMEs) could be used to assign minimum passing rates for each portion of the test. The same SMEs could be used to develop an appropriate weighting scheme for the test (Biddle 2009). Once again, this cannot be done after the test is administered. Ideally, a validated test allows for qualified candidates to rise to the top of the selection process. However, this can only be accomplished if the test accurately reflects the key success factors for the job. Thus, the weighting scheme should be an accurate assessment of the key skills necessary to do the job, and the cutoff score should accurately reflect the minimum competency level for each skill (Biddle 2009). In any event, it is advisable that the organization document the rationale for the cutoff score (Goldstein, Scherbaum, and Yusko 2010).
Monitor for adverse impact. Finally, the organization can protect itself by vigilantly monitoring the job, looking at the content of the job, and making sure changes in the job are documented. A hospitality employer must pay careful attention to any adverse impact that may occur because of these changes (Lundquist and Ashe 2010).
The Hospitality Connection
Although the Ricci case involved public employees, we focused our recommendations on the hospitality industry for several reasons. First, the hospitality industry’s high turnover means that hospitality operations are continually selecting new employees. This often involves some type of candidate testing. One lesson of Ricci is that combining assessment tools to minimize liability for adverse impact (recommendation 3) and setting appropriate cutoff scores (recommendation 5) should be part of a consistent selection plan for all levels of employment in the hospitality industry.
Given the hospitality industry’s vast variety of jobs, tests used for selection and promotion should also be varied to fit the variety of positions and managers’ perceptions of the requirements for those positions. A recent study by Tews, Stafford, and Tracey (2011) demonstrated that hospitality managers tend to focus more on agreeableness and conscientiousness than general intelligence when making hiring decisions, despite the fact that general intelligence has been shown to be a significant predictor of performance for entry-level employees (Tracey, Sturman, and Tews 2007). So, in addition to seeking agreeable employees, managers could use alternative test techniques (recommendation 2) and combining assessment tools (recommendation 3) to fill the industry’s many diverse positions.
Managers’ preference for agreeable, mature people makes sense, given that the hospitality industry relies so heavily on exceptional customer service as its product. Employee personality and emotional intelligence is often perceived as one of the most important attributes for employees (Kusluvan et al. 2010; Tracey, Sturman, and Tews 2007). Once again, a combination of test techniques will create a more complete selection process. Using alternative test techniques, like profiling (recommendation 2), and combining assessment tools, like personality tests (recommendation 3), seem to be a logical fit for hospitality employers to use as part of their selection process. We remind employers, however, to validate their test (recommendation 4) and monitor for adverse impact (recommendation 6), even though personality tests produce smaller ethnic differences than other tests do.
Finally, because of changes in technology, most notably self-service technology, the hospitality industry environment is constantly changing (Lema 2009), as are many jobs. For example, computer and technology skills and knowledge have become essential for many jobs that traditionally did not include such skills. New skills are added and requirements are changed, but the job titles often remain the same, meaning that thorough job analyses (recommendation 1) will be vital to keep up with the changes. This will ensure that employers are selecting for the appropriate skills and abilities, and that they will make use of the best, most appropriate selection tests and methods.
Footnotes
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
