Abstract
Prior research indicates mixed findings regarding the consistency of adjudicators’ ratings at large ensemble festivals, yet the results of these festivals have strong impacts on the perceived success of instrumental music programs and the perceived effectiveness of their directors. In this study, Rasch modeling was used to investigate the potential influence of adjudicators on performance ratings at a live large ensemble festival. Evaluation forms from a junior high school concert band festival adjudicated by a panel of three expert judges were analyzed using the Many-Facets Rasch Model. Analyses revealed several trends. First, the use of assigning “half points” between adjacent response options on the 5-point rating scale resulted in redundancy and measurement noise. Second, adjudicators provided relatively similar ratings for conceptually distinct criteria, which could be evidence of a halo effect. Third, although all judges demonstrated relatively lenient ratings overall, one judge provided more severe ratings as compared to peers. Finally, an exploratory interaction analysis among the facets of judges and bands indicated the presence of rater-mediated bias. Implications for music researchers and ensemble adjudicators are discussed in the context of ensemble performance evaluations, and a measurement framework that can be applied to other aspects of music performance evaluations is introduced.
Keywords
Introduction
Large ensemble festivals, adjudicated performances in which a panel of expert adjudicators assign performance ratings and provide written feedback to performing ensembles, are prevalent in secondary school music programs in the United States (Fiske, 1983). Instrumental ensembles often participate in festivals on a regular basis to receive written feedback and ratings, which can be used diagnostically to improve future performances. At festivals, adjudicators typically provide ratings of a variety of specific performance criteria (e.g. articulation, intonation, and tone quality), as well as written comments designed to inform performers about areas of accomplishment and areas needing improvement (Boyle & Radocy, 1987; Forbes, 1994). Results of large ensemble festivals are often used to determine the relative success of instrumental programs (Forbes, 1994) and their directors (Burnsed, Hinkle, & King, 1985). In fact, Burnsed et al. stated that “[e]nsemble directors often place the success of their programs and indeed, the success of their own careers, in the hands of a panel of judges” (Burnsed et al., 1985, p. 22).
One of the common criticisms of large ensemble festivals is that adjudicators do not always agree on the ratings assigned to evaluate musical performances (Cooksey, 1982; Fiske, 1983; Forbes, 1994). For this reason, numerous studies have been conducted to evaluate the consistency of adjudicators’ performance ratings. Levels of interjudge reliability have been reported to be notably varied at large ensemble festivals (Brakel, 2006; Burnsed et al., 1985; Garman, Boyle, & DeCarbo, 1991). Boyle and Radocy (1987) explained that even though adjudicators sometimes demonstrate agreement in terms of overall, global ratings of performances, their ratings of specific performance criteria tend to vary. Research findings are not unanimous, however. Some researchers (Burnsed et al., 1985; Burnsed & King, 1987) reported that interjudge reliabilities were higher for global performance ratings than for ratings of specific performance criteria, although others (Garman et al., 1991) reported low reliabilities for both overall ratings and specific ratings. Based on these mixed findings, the relationship between overall and specific performance ratings in the context of ensemble evaluations is not clear.
Some researchers (Brakel, 2006; Heath, 1976; Winter, 1993) have reported that training resulted in increased rating consistency among adjudicators in music performance evaluations. Fiske (1983), however, explained that training does not always result in greater consistency but suggested in contrast that these disparate findings could be related to the particular training procedure used. Kinney (2009) also introduced the potential impact that varied statistical methods used to measure reliability could have on these different results across studies.
In addition to varied levels of reliability, researchers have also noted a predisposition for leniency among adjudication panels. Specifically, Hash (2013) and Lehman (1968) claimed that many adjudicators are reluctant to assign overall ratings lower than “excellent,” and Boeckman (2002) reported an increase in assigned performance ratings at a statewide band festival over a 50-year period. This leniency could be a result of generosity error, a phenomenon associated with a tendency of adjudicators to empathize with those individuals they are rating to such an extent that it modifies their ratings of specific performance tasks/criteria (Thorndike, 2005).
Given this predisposition for leniency and varied levels of reliability among adjudication panels (Brakel, 2006; Hash, 2013), music researchers have sought to understand what factors increase adjudicators’ rating consistency. Results of prior research have indicated that rating consistency can be improved with increases in adjudicators’ expertise (Kinney, 2009), size of adjudication panel (Brakel, 2006; Fiske, 1983), and familiarity with music performed (Kinney, 2009). The focus of the present study was to explore similar adjudication effects at a live music performance, but in this case, these effects were investigated within the framework of the Rasch measurement model.
Rasch modeling
Initially conceived by Danish mathematician Georg Rasch (1960/1980), the term Rasch model actually refers to a family of probabilistic measurement models that are used to model the facets of item difficulty (or endorsability) and person ability (or agreeableness) on an equal-interval linear scale (Engelhard, 2013). The Rasch model is related to a one-parameter item response theory model and is based on similar measurement theory (Bond & Fox, 2015; DeVellis, 2003). Logits, or log-odds units, are the fundamental units of measurement in Rasch analyses (Engelhard, 2013). By modeling the facets of item difficulty (i.e. the difficulty of one item in comparison to other items in the same instrument) and person ability (i.e. the ability/proficiency of each respondent/examinee) on an equal-interval logit scale, researchers are able to evaluate both of these measurement facets independently on the same linear scale. Bond and Fox (2015) explained that this use of true interval-level measurement is one of the greatest strengths of Rasch measurement.
When conducting Rasch analyses, researchers are able to view each measurement facet on a variable map, an important measurement tool that is similar to a ruler or yardstick. A variable map provides a visual representation of the underlying construct of interest (Engelhard, 2013), allowing researchers to view and make comparisons between different measurement facets (e.g. item difficulty and person ability) on the same linear logit scale (Bond & Fox, 2015; Myford & Wolfe, 2004).
Rasch analyses let researchers overcome many of the limitations often associated with classical test theory (CTT), including: (1) the erroneous treatment of ordinal data as interval data (Bond & Fox, 2015); (2) the treatment of raw scores as fundamental measures of psychological constructs (Royal, 2010; Wright & Mok, 2004); and (3) the interpretation of all survey/scale items to be of equal weight (Bond & Fox, 2015; DeVellis, 2003; Engelhard, 2013). Furthermore, by applying the Rasch measurement model to calibrate measurement instruments, researchers are able to demonstrate the rigorous property of measurement invariance, which refers to the stability of item difficulty estimates and person ability estimates across other distributions of items and persons – a prime indicator of psychometric quality (Bond & Fox, 2015; Engelhard, 2013).
Facets other than person abilities and item difficulties can systematically impact measurements, such as the influence of a judge or rater. The Many-Facets Rasch Model (MFRM) was initially developed by Linacre (1989), based on a realization that scoring examinees or performers based on traditional raw-score methods was unfair to those who encounter raters who are more severe or lenient than others. Linacre developed the MFRM in response to rater severity because he considered the role of a rater or judge to be an additional measurement facet, rather than an innocuous, passive aspect of the measurement scenario. In a discussion of the utility of the MFRM in rater-mediated performance evaluations, Bond and Fox elaborated:
Why then, in important evaluation situations, do we continue to act as though the judge, rater, or examiner has merely a benign role? On a personal level, we might try to avoid the tough marker, complain that some judges are biased against us, or avoid the examiner’s specialist topic, but we might as well face it; in high-stakes testing, we often have the suspicion that the marker, not the candidate or the test, might mean the difference between pass and fail, that the scorer rather than the performance determines silver, not gold. (Bond & Fox, 2015, p. 167)
Using the MFRM, one or more facets (e.g. rater/judge severity) can be modeled in addition to person ability and item difficulty (Bond & Fox, 2015; Linacre, 1989; Myford & Wolfe, 2004). Like other Rasch measurement approaches, all measurement facets are modeled concurrently on an equal-interval logit scale, and all facets are visually illustrated on a variable map. Due to the joint calibration of all measurement facets, researchers are able to evaluate the severity of the rater/judge on the same linear scale as the ability of the performer/examinee and the difficulty of the items/performance criteria to be rated (Myford & Wolfe, 2004). Thus, in a music performance evaluation context, the MFRM allows researchers to not only consider the proficiency of the performer(s); they can also examine the difficulty of the rating scale criteria and the severity of the adjudicator(s).
Researchers have used the MFRM in a variety of situations in which fundamental measurement is mediated by a rater or judge. Linacre (2009) identified systematic bias among the judges of the 2002 Winter Olympics Paired Figure Skating Competition. Other authors have studied the influential role of the rater/judge among writing assessments (Eckes, 2008), PowerPoint presentations of first-year college students (Basturk, 2008), “creative inventions” in an undergraduate design class (Hung, Chen, & Chen, 2012), second-language oral assessments (Bonk & Ockey, 2003), and grades assigned by elementary, middle, and high school teachers (Randall & Engelhard, 2009). In total, these findings give credence to the notion that judges may be one of the more influential facets of measurement in performance-based evaluations.
Need for the study
Music researchers have explored a number of issues related to performance evaluations, including the role of adjudicators. Adjudicators are not always in agreement when evaluating musical performances (Burnsed et al., 1985), and previous research (Brakel, 2006; Burnsed et al., 1985; Hash, 2013) has indicated notably varied levels of interjudge reliability among adjudicators at formal ensemble performance evaluations. It should be noted that these traditional approaches to interjudge reliability are most often based on composite statistics (i.e. means or sums), rather than specific ratings of individual items or performance criteria. When using these composite-score approaches, researchers may overlook the fine-grained and subtle differences in adjudicators’ ratings of specific performance criteria that meaningfully impact measurement scenarios (DeVellis, 2003), especially since such criteria are often differentially weighted in “real life” practice (Bond & Fox, 2015). Thus, it seems wise for music researchers to consider measurement models that allow for item-level or criterion-level inspection, rather than global ratings only. Additionally, traditional statistical approaches to interjudge reliability are dependent upon the particular set of raters used, whereas Rasch analyses permit researchers to assess whether or not an instrument is invariant across samples of raters (Engelhard, 2013). These reasons support the importance of examining music performance evaluations using measurement models that overcome such limitations, such as the MFRM (Bond & Fox, 2015; Linacre, 1989, 2012).
Wesolowski, Wind, and Engelhard (2015) applied the MFRM to the evaluation of recorded jazz ensemble performances and found that ratings were not invariant across a panel of expert judges. They reported that judges demonstrated systematically-different levels of severity across ensembles of different school levels (middle school, high school, collegiate, and professional). Although rater-mediated evaluations are quite common in various other music performance evaluation settings (Boyle & Radocy, 1987), no studies have used many-facet Rasch modeling to examine the potential influence of adjudicators at live large ensemble festivals. Thus, the central purpose of this study was to investigate the influence of adjudicators on performance ratings at a live concert band festival using the MFRM. This purpose was achieved through two common Rasch analysis procedures, specifically: (1) a descriptive examination of three measurement facets involved in the evaluation context (bands, performance criteria, and judges); and (2) an exploratory interaction analyses to investigate potential sources of bias among the rating behavior of adjudicators. Furthermore, because the adjudicators in the present study used a system of “half points” between adjacent rating scale options, a secondary purpose of the study was to evaluate diagnostics of the rating scale, particularly the psychometric results of using “half points” because anecdotal evidence suggests that this practice may be common among solo, chamber ensemble, and large ensemble evaluations.
Method
Procedure
In an effort to sample from an authentic large-group adjudication setting, data were collected for the present study from a regional concert band festival, which was held on two consecutive days in a metropolitan area in the Pacific Northwest region of the United States. Concert bands from all middle schools (grades 6–8) and junior high schools (grades 7–9) in the statewide district area participated in the festival, resulting in a listing of 31 participating bands. The bands ranged in size from 11 to 135 performers (M = 42.7, Med = 39.0) and represented a wide range of performing proficiencies. Per festival guidelines, each ensemble performed three contrasting pieces of concert band music, which they were permitted to select for their performance. Each ensemble was rated on their overall performance, rather than separate ratings for each piece (i.e. ratings were aggregated across all three pieces).
The festival was adjudicated by a panel of three expert judges. These judges were considered experts for this study because they had completed graduate degrees in music (two with doctoral degrees; one with a master’s degree) and because they were experienced adjudicators of concert bands. Furthermore, all of the judges had achieved success as band directors as deemed by the festival chairperson and the festival leadership committee, who selected these judges through majority vote.
Judges were mailed a training packet two weeks before the date of the festival, which included a blank adjudication form and an adjudication instructions document. Judges were asked to read the instructional document and study the sample adjudication form before arriving at the festival. On the first day of the festival, the chairperson led a 30-minute adjudication training session with the judges before the first ensemble performed. As described by the festival chairperson (to the first author), the training session was used to review the material on the instructions document, discuss the adjudication form, and to give the judges an opportunity to ask questions about the adjudication procedures. There were no other training criteria or training procedures used.
Evaluation form
The instrument used to evaluate the ensembles was the National Federation of State High School Associations Large Group Music Adjudication Form (NFHS, n.d.). On this instrument, judges rated each ensemble on eight performance criteria similar to those used in previous studies (e.g. Johnson & Geringer, 2007; Springer, 2016) – tone quality, intonation, rhythm, balance/blend, technique, interpretation/musicianship, articulation, and “other performance factors,” which were described as follows: “choice of literature, appropriate appearance, poise, posture, general conduct, [and] mannerisms” (NFHS, n.d.). Judges rated each criterion on a five-point scale with the following anchors:
5 = A superior performance – outstanding in nearly every detail 4 = An excellent performance – minor defects 3 = A good performance – lacking finesse and/or interpretation 2 = A fair performance – basic weaknesses 1 = A poor performance – unsatisfactory. (NFHS, n.d.)
Although a 5-point scale was printed on the adjudication form, the festival chairperson encouraged the judges to utilize “half points” between adjacent categories to enable more refined response options. This instruction was communicated to the judges in written form (on the adjudication instructions document) and verbally (during the training session on the first day of the festival). As instructed, all judges adhered to the practice of assigning half points, which resulted in a scale comprising nine discrete response options – 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, and 5. All of these response options were assigned by judges, with the exception of 1.5.
Data analyses
Data analyses were conducted using Facets computer software (Linacre, 2014), which uses maximum likelihood estimation procedures to conduct MFRM analyses. Analyses for the present study were based on a three-facet model (bands, performance criteria, and judges) using the following formula (Bond & Fox, 2015; Linacre, 1989; Myford & Wolfe, 2004):
As outlined in this MFRM formula, the probability of a particular rating was modeled as a function of the proficiency of the band (B n ), the difficulty of the performance criterion (D i ), the difficulty of the category threshold between adjacent response options in the rating scale structure (F k ), and the severity of the judge (C j ; Bond & Fox, 2015). Results of the MFRM analysis allowed for the evaluation of three primary facets of interest (bands, performance criteria, and judges) independently and to model those facets on an equal-interval logit scale (Myford & Wolfe, 2004).
Results
Rating scale diagnostics
Although a 5-point scale was printed on the adjudication form, as indicated previously, all judges utilized a modified version of the scale (using “half points”) composed of nine discrete response options. For this reason, diagnostics of the rating scale were evaluated to determine whether this adjusted scoring procedure functioned as intended before conducting primary analyses. As shown in the upper half of Figure 1, probability curves for the modified 9-point scale indicate much redundancy and dependence among response probabilities because the curves do not exhibit independent probability “peaks,” and also evident in the upper half of Figure 1 is a lack of clear thresholds between adjacent response options, which provides further evidence of measurement noise (Bond & Fox, 2015). Based on common practice used in Rasch scale calibration (Bond & Fox, 2015; Wright & Linacre, 1992), the nine response options were recoded to a 5-point scale as originally intended using the following scheme: 1, 2, 2, 3, 3, 4, 4, 5, and 5. The recoding procedure was conducted following a visual inspection of the initial probability curves using procedures recommended by Wright and Linacre (1992). This recoding scheme provided evidence of improved rating scale performance as demonstrated in the lower half of Figure 1, where all five response options illustrate independent probabilities with clear thresholds visible between adjacent response options.

Probability curves for scale with nine response options (upper graphic) and five response options (lower graphic).
Analysis of the three-facet model
After examining the diagnostics of the rating scale, a MFRM analysis was conducted to evaluate the three facets included in the measurement model (proficiency of the bands, difficulty of the performance criteria on the rating scale, and severity of the judges). The majority of variance (66.81%) was explained by the Rasch model, with the remainder (33.19%) being explained due to residuals. Summary statistics of the three-facet analysis are provided in Table 1. As shown in Table 1, bands and criteria were centered with a mean element measure of zero to study the potential influence of judges. Mean square infit and outfit statistics indicated a good degree of model fit (ranging from 0.95 to 0.99 across all facets). Significant differences were found for all three facets: bands, χ2(30) = 929.80, p < .001; criteria, χ2(7) = 207.40, p < .001; and judges, χ2(2) = 17.20, p < .001. These results indicate differences among the elements of each facet that are beyond the scope of measurement error. Thus, the hypothesis that the elements of each facet have the same measure (allowing for measurement error) was rejected due to the significant chi-square values for each facet (Linacre, 2012).
Summary statistics for the three-facet Rasch analysis.
Note. Summary statistics for the performance measure are reported in logits. To examine the influence of judges, the bands and criteria facets were centered with a mean element measure of zero. SD refers to the true standard deviation, adjusted for measurement error; RMSE = root mean-square error, a statistical average of standard errors for the measures; separation ratios (G) are calculated as true SD/RMSE and are interpreted as a ratio of the spread of performance measures relative to their precision; strata indices (H) are calculated as (4G+1)/3 and are used as an index of the number of different measurement levels (strata) detected within each facet. Rasch reliability statistics are calculated as the ratio of true variance to observed variance and provide an indication of the reproducibility of the measures; chi-square statistics (with associated degrees of freedom) test the null hypothesis that the measures of facet elements are all the same, apart from measurement error (Linacre, 2012; Myford & Wolfe, 2004).
p < .001.
Rasch reliability statistics indicated a high degree of reproducibility for the facets of bands (0.97), criteria (0.98), and judges (0.88). Separation ratios 1 (G), which function as a measure of the spread of measures relative to their precision, and strata indices 2 (H), which refer to the number of different measurement levels (strata) observed in each facet, are also displayed in Table 1. Based on these results, multiple measurement strata were observed for each facet – bands (7.64), criteria (9.45), and judges (3.96). Statistically, these values indicate that there were over seven strata (or grouped clusters) of bands based on proficiency, as well as over nine strata for performance criteria and over three strata for judges. These reliability, separation, and strata values were found to be “good” to “excellent” according to Fisher’s (2007) quality control criteria, indicating adequate measurement qualities for a MFRM analysis.
As shown on the variable map in Figure 2, each facet is visually illustrated along a logit scale. The bands represented a wide range of performance proficiencies (approximately ten logits), with Band 11 and Band 3 serving as the highest-performing and lowest-performing ensembles, respectively. Regarding the performance criteria facet, intonation functioned as the most difficult criterion, as noted by its highest placement on the variable map. “Other performance factors” functioned as the easiest criterion. A visual inspection of the map indicates that all but one of the performance criteria clustered near the midpoint, meaning that these categories did not effectively distinguish among the highest-performing ensembles or the lowest-performing ensembles. For this reason, it may be beneficial to include additional criteria on this evaluation form for future use – particularly more difficult criteria that could discriminate among the highest-performing ensembles at the top of the variable map. This observation could also be indicative of a halo effect, a phenomenon in which adjudicators rate the performance of a person/group similarly on criteria that are conceptually different (Eckes, 2011; Myford & Wolfe, 2004).

Variable map of the three facets modeled in the Rasch analysis (bands, criteria, and judges).
Judges were also mapped on the same logit scale, and their placement suggests some amount of leniency or generosity error (Myford & Wolfe, 2004; Thorndike, 2005). As shown on the map, Judges A and B were roughly equivalent in terms of severity/leniency, but Judge C demonstrated more severity by approximately one logit. Further interaction analyses were conducted to examine patterns of rating behavior across facets more closely (Linacre, 2012).
Exploratory interaction analysis
An exploratory interaction analysis (Eckes, 2011) was conducted using an alpha level of 0.05 (with a Bonferroni adjustment for inflation of Type I error) to investigate whether or not there were differential patterns of severity in the rating behavior of judges across bands and across criteria. Summary statistics are presented in Table 2: a statistically significant judges × bands interaction was observed, χ2(93) = 138.80, p < .001 which indicates that judges demonstrated systematically varied levels of severity across ensembles (Myford & Wolfe, 2004). Thus, performance ratings were not invariant across the panel of three judges. Thirteen out of the 93 total interactions between judges and bands (13.98%) were significant. Neither the two-way interaction among judges and criteria nor the three-way interaction among judges, criteria, and bands were statistically significant (p > 0.05), however. Although four statistically significant observations for the judges × criteria interaction are noted in Table 2, the overall interaction between these two facets was not statistically significant, χ2(24) = 29.50, p = .20, so these four observations could simply be due to error.
Summary statistics for the exploratory interaction analysis.
Note. Summary statistics include the number of facet elements considered in the interaction analysis (n observations), the number of elements with statistically significant t-values (p < .05, two tailed), the minimum and maximum t-values (with their associated degrees of freedom), and the means and standard deviations of the t-values (Eckes, 2011).
p < .05, **p < .01.
Figure 3 is a display of the bias control chart, with quality control lines, for the significant judges × bands interaction. Judges’ ratings that fall within the lines are typical, expected aberrations due to measurement error, but ratings that exceed these boundaries are systematic deviations and are evidence of bias (Eckes, 2011; Linacre, 2012). On the bias control chart, observations that occur above the quality control line are indicative of a systematically lower (i.e. more severe) rating, and observations that occur below the line provide evidence of a systematically higher (i.e. more lenient) rating. In this MFRM analysis, the logit scores for bands were adjusted for bias in other facets. This type of adjudicator bias can be seen regarding the haphazard ratings of Band 26. Judge C provided an unexpectedly low rating of this band’s performance (more severe), while Judge B rated their performance systematically higher (more lenient). The opposite trend was found for Band 1. Judge C’s rating was systematically higher than expected, while Judge B rated the performance unexpectedly lower. These were only two examples of adjudicator bias that were revealed based on the MFRM analysis, which provided empirical data indicating that the judges did play an influential role on the measurement of music performance. Based on the observed bias found in this study, future efforts to reduce adjudicator bias through training or other procedures are needed.

Bias control chart illustrating judges × bands interaction.
Discussion
The purpose of this study was to investigate the influence of adjudicators on performance ratings at a live concert band festival using an application of the MFRM, and results highlighted the influential role of the judges in an authentic large ensemble adjudication setting. By using a three-facet latent trait model (specifically, the MFRM), a critical inspection of performance ratings was possible – one that incorporated three measurement facets impacting the evaluation of performance quality (i.e. the proficiency of the bands, the difficulty of the performance criteria being rated, and the severity/leniency of the judges).
As a result of this MFRM analysis, certain trends were observed that could not have been noticed using CTT approaches. For instance, the practice of using “half points” between the five response options resulted in measurement noise due to redundancy and a lack of independence among response options. This finding supports adjudicators utilizing the original five-point scale as instructed on the form for best practice. It is important to note, however, that these recoding conclusions are based on a verbally-communicated half-point strategy. For this reason, it is not surprising that the half points had lower probabilities because they were used with less frequency. Although there was evidence of improved rating scale performance when the data were recoded to a five-point scale, it is not possible to know how these results would apply to a printed nine-point scale. It is also worth noting that the re-categorization from a nine-point scale to a five-point scale could have also resulted in more leniency among certain judges because those who used more half points would have had more ratings recoded to the higher adjacent category.
Results of the Rasch analyses indicated that all judges demonstrated relatively lenient rating behavior – a finding that is consonant with earlier reports (Boeckman, 2002; Cooksey, 1982; Hash, 2013; Lehman, 1968). Perhaps this result could be due to the relatively young, lesser-experienced age groups associated with the middle school and junior high school performing ensembles who were evaluated in this study. It could be surmised that the judges viewed this opportunity to evaluate these young ensembles with leniency in order to provide a more supportive, motivating evaluation of their performance, instead of a more critical, strict evaluation. This generosity error could also be explained in part due to “a widespread unwillingness, at least in the United States, to damn a fellow human with a low rating” (Thorndike, 2005, p. 372).
Despite this observed leniency, however, one judge (Judge C) demonstrated a tendency to rate more severely overall. This additional severity could simply reflect that judge’s honest evaluation of the performances, which may have been more demanding than that of the other judges. It could also be the result of inadequate adjudication training and calibration. Although prior research (Brakel, 2006; Fiske, 1983; Heath, 1976; Winter, 1993) is mixed regarding the effects of training on adjudicators’ reliability, Fiske (1983) explains that different training methods used in previous studies could account for these inconsistent results. Thus, in the present study, the particular strategies used to train adjudicators during the 30-minute training session (i.e. discuss adjudicator instructions document, discuss evaluation form, and allow time for adjudicators to ask questions) could have been insufficient to successfully calibrate the adjudication panel. The results of this study are based on data collected from three raters, though. Thus, the results are specific to a particular performance evaluation and a small panel of judges, and the generalizability of these findings is limited.
Judges demonstrated some amount of homogeneity in their ratings of performance criteria because all of the criteria (with the exception of “other performance factors”) occurred near the midpoint of the variable map. These criteria were restricted to a difficulty range of approximately one logit (log odds unit), which could signify a halo effect. There is ample evidence that judges tend to rate specific performance criteria based on general, overall impressions (Boyle & Radocy, 1987; Myford & Wolfe, 2004), and some authors have pointed out that this effect is common with adjudicators of music performances (Forbes, 1994). One possible explanation for this halo effect could be due to the fact that the judges who were hired for this concert band festival were from the same geographic region as the performing ensembles. It is possible that they knew some of the local directors personally or could have been familiar with past performance histories of the ensembles, which could have resulted in more of an emphasis on generalized impressions rather than specific performance criteria (Boyle & Radocy, 1987; Forbes, 1994; Thorndike, 2005).
Of all performance criteria listed on the evaluation form, intonation functioned as the most challenging performance criterion, and “other performance factors” functioned as least challenging. It is not surprising that intonation served as the most challenging criterion due to the level of experience of the ensembles being rated. Previous researchers (Johnson & Geringer, 2007; Springer, 2016) have indicated that intonation is often judged to be the element that is most in need of improvement with younger concert bands. The “other performance factors” criterion, which was defined on the evaluation form as “choice of literature, appropriate appearance, poise, posture, general conduct, [and] mannerisms” (NFHS, n.d.), offered the judges opportunities to rate a variety of performer characteristics and behaviors based on their stage presentation. This criterion was notably different than the other criteria on the evaluation form, which were focused on specific performance outcomes. Judges may have used this category to reward these young musicians on their perceived effort by rating them highly on this criterion.
The most compelling results describing the influence of the adjudicators, however, were illustrated by the significant judges × bands interaction. The bias control chart (see Figure 3) indicates that judges did not demonstrate uniform severity from one band to another. This interaction between judges and bands provides evidence of adjudicator bias (Eckes, 2011; Engelhard, 2008; Linacre, 2012) resulting from the presence of unwanted “construct-irrelevant variation” (Myford & Wolfe, 2004, p. 41) in performance ratings, which confounds the idea of measurement precision. In the case of the present study, 13 out of 93 possible interactions between judges and bands (13.98%) were statistically significant, indicating the presence of bias among those ratings. This finding is highly consistent with results of a previous study (Wesolowski et al., 2015), whose authors reported that 15% of bias interactions were significant among a panel of expert jazz ensemble judges. Because there was an observed tendency for leniency among the judges in the present study, nine of the 13 observed deviances were below the lower quality control line in Figure 3, indicating that many ratings were systematically higher than expected. Ratings were systematically lower than expected for the four additional deviances. Because performance ratings assigned by festival adjudicators have such an impact on instrumental programs (Forbes, 1994) and their directors (Burnsed et al., 1985), it is important that continued efforts be undertaken to identify and reduce these forms of bias in adjudication settings.
Implications for adjudication practice
Even though these adjudicators were identified as experts who were experienced in ensemble adjudication, their rating behavior provided evidence of bias. Despite efforts made at increasing the consistency among adjudicators’ ratings through training, albeit limited to 30 minutes prior to judging, results suggested that further training procedures and judge calibrations were needed. This finding is consistent with results of prior studies (Boeckman, 2002; Fiske, 1983) that indicated no improvements as a result of training but inconsistent with others (Brakel, 2006; Heath, 1976; Winter, 1993).
To improve the training of adjudicators, future efforts might include the use of an anchoring technique, whereby the adjudicators rate sample recorded performances from previous festivals to identify, through consensus, the aural qualities necessary for rating each performance criterion on the scales provided on the evaluation form (Boyle & Radocy, 1987; Heath, 1976). This practice could improve the consistency of adjudicators’ ratings because it would likely help calibrate their level of leniency/severity a priori and establish conceptual boundaries between rating scale categories. Furthermore, it may also be possible to show adjudicators the results of a MFRM analysis (e.g. the variable map or bias control charts) after adjudicating music performances to provide a visual indication of their own severity, or lack thereof, in comparison with other adjudicators. This exercise may result in greater consistency in future adjudication sessions – an assumption that is based only on anecdotal evidence and should be examined empirically in future studies.
All of the bands performed different pieces of music at this festival, and the difficulty levels of each piece varied as well. Unfortunately, the specific pieces that were performed (and their respective difficulty levels) were not included on the evaluation forms for this study, so it was not possible to model the musical selection (or difficulty level). In future studies, it would be useful to collect a rating of perceived music difficulty by the judges and include this facet in the model. Since each band was allowed to select three pieces of music to perform, it is certainly possible that judges could have responded differently based on the music selected by the bands. This potential confound should be investigated in future studies to determine potential sources of bias that could be the result of music selection (i.e. a possible music × rater bias interaction).
Some authors (Boyle & Radocy, 1987; Forbes, 1994) have recommended that adjudicators be hired who have no prior knowledge of performers or directors as a means of reducing halo effects and generosity error, so festival chairpersons might consider hiring adjudicators from different geographic regions. Myford and Wolfe (2004) outlined other possible solutions, which included: (1) training adjudicators to be aware of halo effects and generosity error so they can make efforts to avoid this tendency; and (2) using a forced distribution method of assigning ratings in which adjudicators must place a predetermined number of ensembles in each rating category. Potential halo effects could also be detected by modeling a differentiated rating for the three pieces of music that were performed (as suggested above). Researchers may also consider using measurement instruments that include more performance criteria. Based on the results found in this study, there was a need for more performance criteria, especially more difficult criteria that could discriminate among the highest-performing ensembles at the top end of the variable map (Figure 2).
Implications for music research
Although studies using Rasch modeling have been disseminated for many years in other fields, there is scant evidence that these procedures have been used in the field of music research, with a few notable exceptions (e.g. Bergee & Antonetti, 2010; Pascoe & Waugh, 2001; Springer, Rojas, & Bradley, 2014; Wesolowski et al., 2015; Yim, Abd-El-Fattah, & Lee, 2007). Applications of the Rasch model offer strong benefits when compared to CTT analyses and traditional statistical approaches. Rasch analyses allow researchers, and subsequently practitioners, to work with true interval-level measurement, thus avoiding the inaccurate treatment of ordinal data as interval data that is common in previous studies (Bond & Fox, 2015; DeVellis, 2003; Engelhard, 2013; Royal, 2010). Although traditional approaches to interjudge reliability are useful in gaining a preliminary idea of how well adjudicators agree on performance evaluations, those approaches are limited. This is due to the fact that traditional statistical approaches are not based on item- or criterion-level measurement; rather, they are based on means or other summary statistics, which are not as fine-grained and informative as their item-level counterparts (Bond & Fox, 2015). Furthermore, because the Rasch model permits analyses of dichotomous data, polytomous data with ordered categories, and even partial credit data, its application is not constrained by type of data collected. Thus, music researchers would be wise to consider using this versatile measurement model in the future to develop and validate high-quality measurement tools that are invariant among distributions of items and persons (Bond & Fox, 2015; Engelhard, 2013).
Performance evaluations in music are nearly always mediated by one or more raters/judges (Boyle & Radocy, 1987; Fiske, 1983; Forbes, 1994). For this reason, the many-facet approach is needed for musicians to understand the significant role that judges play in the evaluation of musical performances. The use of the MFRM in particular will be useful in future studies because it allows researchers to consider both the raters (i.e. adjudicators) and ratees (i.e. performers), in addition to the performance criteria on the rating scales. Although results of this study provided empirical evidence of bias among a panel of expert adjudicators, it is important to note that it is not possible to identify the source of the bias based on these data alone (Wesolowski et al., 2015).
Given the results of this study, it may be possible to share the results of a MFRM analysis with adjudicators in graphic form to provide a visual description of their rating behavior. Such visual aids could potentially be used to identify sources of bias and to reduce those biases in future performance evaluations. For multi-day ensemble festivals, sharing data with judges in this format at the end of each day could serve as a way of calibrating their ratings – an idea that should be examined in a future study. Another option for remediating bias would be to use the mean scores from a Rasch-based measurement analysis that have been adjusted for judge severity.
Future studies are needed to investigate the source of adjudicator bias in other performance contexts (e.g. solo performance evaluations and chamber ensemble competitions) to gain a more comprehensive understanding of how adjudicators impact the evaluation of musical performance. Additionally, it would be useful to examine the effect of certain adjudicator training procedures (e.g. the anchoring technique described above) to determine whether or not those procedures can be used to calibrate adjudicators in music performance contexts. Finally, a comparison of traditional raw-score approaches and Rasch-adjusted person ability estimates in a music audition context will be particularly useful for music researchers to better understand the benefits of invariant measurement approaches (Engelhard, 2013). With the knowledge that adjudicators do exert an influence on performance measurement, a notably subjective area for evaluation (Boyle & Radocy, 1987), music researchers should use measurement models that incorporate the judge as an active, prominent facet to result in the most objective measurement possible (Linacre, 1989). Doing so will help improve the consistency of adjudicators’ ratings, which could ultimately result in greater equity/fairness in ensemble performance evaluations.
Footnotes
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
