Abstract
The purpose of this study was to develop a valid and reliable rubric to be used for the evaluation of large ensemble wind band performances. The guiding questions for this study were: (a) what are the psychometric qualities (i.e., reliability and validity) of the scale developed to assess wind band ensemble performance at the high school level? (b) how do the items fit the model and vary in difficulty? (c) how does the structure of the rating scale vary across individual items? and (d) how can the rating scale be transferred into an informative rubric? The primary data analysis tool used in this study was the Multifaceted Rasch Partial Credit Measurement Model. Music content experts (N = 20) were solicited to evaluate 40 wind band performances, each evaluator listening to four. A 4-point Likert-type rating scale (e.g., Strongly Agree, Agree, Disagree, and Strongly Disagree) was used to evaluate each recorded performance. Results indicated good model data fit and resulted in a final rubric containing 24 items ranging from two to four performance categories. Implications for classroom teaching and consequential validity are discussed.
Decisions made by administrators, stakeholders, and policy makers regarding school programs are strongly guided by empirical student achievement data (Swan & Mazur, 2011; Wayman, 2005). Empirical evidence of student achievement is often at the foundation of policy-based decisions (O’Neal, 2012) and is traditionally derived from standardized testing results and stakeholders’ contextualization of the results within specified content and performance standards (U.S. Department of Education, 2009). In the context of performance-based practices such as music and related performing arts, teachers are now held accountable for their role in student learning by providing evidence of student achievement using Student Learning Objective (SLOs) frameworks and other similar methods (Wesolowski, 2014; Wesolowski, Wind, & Engelhard, 2015). Many of these systems are created by either groups of local teachers without adequate training in measuring development processes or district leaders that do not have expertise in the content under evaluation (Wesolowski, 2012). When subject matter experts (e.g., practicing teachers) solely design these assessment systems, validity and reliability concerns exist due to the lack of psychometric analysis of the measurement instruments (Buckley & Marion, 2011). Conversely, when non-subject matter experts (e.g., district leaders or administrators) design these assessment systems, content and construct validity concerns exist due to their lack of content knowledge. Yet, administrators and other stakeholders are increasingly accepting these types of tests as a solidly defended option for providing empirical evidence of teacher effectiveness and student achievement, particularly in the area of music where the subject matter is often conceived of as being highly subjective. As a result, the gathered data often provides an inaccurate description of the actual teaching and learning that is occurring in the classroom. In these instances, decision-making bodies may potentially make decisions for educational programs based upon invalid data, posing unintended consequences stemming from the assessment process and related student achievement outcomes.
As noted by Wesolowski, (2012) the use of rubrics is a fruitful method for collecting empirical evidence of student achievement for music performance assessments. A rubric can be defined as “a set of scoring criteria used to determine the value of a student’s performance on assigned tasks; the criteria are written so students are able to learn what must be done to improve their performances in the future” (Asmus, 1999, p. 6). Rubrics are not only important for collecting student achievement data as a summative method of evaluation, but they are an essential formative assessment method to improve communication and expectations between teachers, students, and parents (Pellegrino, Conway, & Russell, 2015). When a student receives feedback in the form of a rubric, they not only have evidence of their performance achievement as a marked outcome, but they also have tangible information of what specific knowledge, skills, and abilities they need to improve in order to demonstrate achievement at a higher level. The use of rubrics improves the subjectivity of music assessment by using objective markers to evaluate student achievement in a clear and quantifiable way.
Professional development commonly focuses on the development of rubrics for the assessment of an individual student (Kan & Bulut, 2014; Sherman, 2006). However, little research has been done in the field of music on the implementation of rubrics for large ensemble performance evaluation; yet, ensemble performance is often the core activity for many music classrooms in secondary education (Wesolowski, Wind, & Engelhard, 2015). Without valid and reliable assessments being used to evaluate these performances, music students are not receiving the feedback necessary for improved learning. One common method for evaluating large group music performance assessments relies on a rating scale system of four to eight categories that are sum totaled for an overall achievement score (National Association for Music Education, 2016). There is commonly an attached set of qualitative comments from raters in either spoken or narrative form; however, this system lacks clear performance descriptors that specify the interpretation of each of the categories of the rating scale. These performance descriptors are the key pieces of evaluative information needed by judges to reduce construct-irrelevant variability and key pieces of diagnostic information needed by teachers to better evaluate performance quality and set future instructional goals (DeLuca & Benjamin, 2014).
The purpose of this study was to develop a valid and reliable rubric for the evaluation of large ensemble wind band performance using psychometric principles of invariant measurement. This study was guided by the following research questions:
What are the psychometric qualities (i.e., reliability and validity) of the scale developed to assess wind band ensemble performance at the high school level?
How do the items fit the model and vary in difficulty?
How does the structure of the rating scale vary across individual items?
How can the rating scale be transferred into an informative rubric?
Background
A valid and reliable performance evaluation carries value far beyond informing policy-based decisions by administrators, stakeholders, and politicians. A measurement instrument that provides valid, reliable, and fair evaluative data is an important component to improving music teaching and learning throughout the school year. A formal, large ensemble performance evaluation traditionally occurs one time per year for most music students (Colwell, 1970). Outcomes of formal evaluations often guide teachers’ decision-making processes for the next year by allowing the teacher insight into which pedagogical techniques that they have recently used were effective or ineffective (Banister, 1992). The use of an evaluation system that is valid and reliable from judge to judge and year to year can allow teachers to gauge the effectiveness of new classroom techniques, objectives, and repertoire (Abeles, Hoffer, & Klottman, 1994; Howard, 2002).
Colwell (1970) states that the evaluation of ensemble performances “can be the most meaningful evaluation of performance the student receives” (p. 105). Students are motivated by learning contexts that optimize success in learning (Austin, 1988). When evaluations are implemented in a valid, reliable, fair, and meaningful manner, they have the potential to hold great power in improving student motivation (Banister, 1992; Franklin, 1979; K. K. Howard, 1994; Hurst, 1994; Sweeney, 1998). Additionally, when formal evaluations are adjudicated using a statistically validated rubric, diagnostic feedback for improvement becomes more clear and meaningful. As such, students are then able to specifically identify how to improve their performances (Asmus, 1999).
Psychometric considerations
The primary methodology used for the development of measurement instruments in the field of music education has been factor analysis. Factor analysis studies for individual instrument scales include clarinet (Abeles, 1971), euphonium and tuba (Bergee, 1987), voice (Jones, 1986), snare drum (Nichols, 2005), auditioning vocalists (Pazitka-Munroe, 2003), guitar (Russell, 2010), and string instruments (Zdzinski & Barnes, 2002), among others. Musical ensemble performance scales have also been developed using factor analysis methods for choral ensembles (Cooksey, 1977), wind band ensembles (DeCamp, 1980), and string ensembles (Smith & Barnes, 2007). As outlined by Wesolowski (2017), however, the use of factor analysis can be limiting when raters mediate the measure development process.
In the context of performance assessments, data is derived through rater-mediation. More specifically, performances are estimated via a rater’s interaction with evaluative cues set forth in a measurement instrument (Engelhard, 2013). Variability in rater behavior can stem from each rater’s unique experiences, background, and interaction with the measurement instrument (Wilson, 2005). In music, because raters are not trained to evaluate performances with machine-like consistency as in other high stakes performance assessment contexts such as writing (Wesolowski, Wind, & Engelhard, 2015), divergence of rater response is to be expected. In fact, divergence in judges’ responses is often welcomed in music assessment as it provides an opportunity to improve performances from multiple expert perspectives (Wesolowski, 2012). However, when performances are being evaluated empirically, managing the quality of rater behavior is of upmost importance, as construct-irrelevant variability can skew the results (Wesolowski, Wind, & Engelhard, 2016). In this study, the family of Rasch Measurement Models was used to evaluate the quality of rater behavior in the scale development process. The benefit of the Rasch model when specifically using a rater parameter within the model is the five requirements for rater-invariant measurement (Engelhard, 2013). The five requirements for rater-invariant measurement include: (a) rater-invariant measurement of persons (i.e., the measurement of persons must be independent of the particular raters that happen to be used for the measuring); (b) non-crossing person response functions (i.e., a more able person must always have a better chance of obtaining higher ratings from raters than a less able person; (c) person-invariant calibration of raters (i.e., the calibration of the raters must be independent of the particular persons used for calibration); (d) non-crossing rater response functions (i.e., any person must have a better chance of obtaining a higher rating from lenient raters than from more severe raters; and (e) variable map (i.e., persons and raters must be simultaneously located on a single underlying latent variable). When adequate fit of the Rasch model is observed, then the five requirements for invariant measurement as outlined by Engelhard (2013) are met. When the requirements of rater-invariant measurement are met within reasonable empirical boundaries, it is assumed the characteristics of performances, raters, and items do not create construct-irrelevant interference between the data and the model.
The Partial Credit version of the Rasch model (Masters, 1982) provides an important additional interaction parameter that allows for the analysis of the individual rating scale categories between each separate item in the measure. This provides the basis for challenging the basic assumption that rating scale categories are equidistantly spaced across all items. This also allows for an improvement in optimization and precision of the item scales being analyzed due to the review of the unique step difficulties for each rating scale category for each respective item. The statistical output of the Rasch model analysis additionally allows for the review of category distribution within items and appropriate discrimination between performances (Linacre, 2002). The Rasch family of models provides not only a meaningful method for the development of valid and reliable measures, but provides a more consistent method for evaluating rater quality within the context of performance-based assessments. Analysis of the rating data in this study was performed using the computer program FACETS (Linacre, 2014).
Method
Rater content experts
Raters (N = 20) solicited for this study were music educators with an average of 14.2 years (SD = 10.3) of secondary music teaching experience in one southern state in the United States. The raters all had successful experiences directing ensembles that participate annually in large ensemble or festival-style performances. None of the raters had any influence on the development of the items in the rating scale or knowledge of the ensembles for which they were providing ratings.
Development of a priori item pool
The items used in the initial rating scale were developed from both the initial item pool used in DeCamp’s (1980) study as well as through a collaboration of three subject matter experts. DeCamp’s (1980) study employed a factor analysis approach to high school wind band ratings. This study resulted in 117 potential item stems that were then combined and aggregated into 39 initial items used in this study. The items were divided into four domains based on the National Association for Music Education (NAfME) Model Cornerstone Assessments (MCAs) for ensemble performance: (a) tone production (n = 8); (b) rhythm and pulse (n = 10); (c) pitch and intonation (n = 7); and (d) expressive qualities (n = 14) (National Association for Music Education, 2015). Eighteen of the item stems were presented as a positive statement and 21 were presented as negative statements as confirmed with 100% agreement by three subject matter experts not involved in the evaluation process. The items were presented in randomized order for each performance in order to defend against rater fatigue and to attempt to minimize rater errors such as central tendency, halo effect, and/or response sets during the evaluation process (Berk, 2010). A Likert-type rating scale structure was developed for each item. Each item stem was followed by four Likert-type response categories: Strongly Agree, Agree, Strongly Disagree, and Disagree. The rater responses were collected using an online Google form for final analysis preparation. The original rating scale is provided in Figure 1.

Original rating scale.
Performance stimuli
All performance recordings were collected from one formal large group performance evaluation during the 2012 concert festival season. The performances were all professionally recorded. More specifically, the same sound engineer, using the same equipment, and in the same concert hall recorded all of the recordings over a three-day festival event. Seventy-four performances were originally collected from 25 ensembles. The names of the musical selections and performers were made anonymous. An online random number generator was then used in order to assign the performance recordings to 20 distinct raters. Use of the performance recordings and participation of the raters in this study were permissible through the respective research institution’s ethics and review board.
Evaluation collection and rater assessment network
Performance recordings were distributed to the raters using the online Dropbox service. Raters had access only to their four a priori assigned performances. The raters then used a Google Form to complete the evaluation for each individual performance. The evaluation process was designed as an incomplete assessment network (Engelhard, 1997). In this process each rater evaluated four performances, with two overlapping performances per rater (e.g., Rater 1 evaluated performances 1, 2, 3, and 4, Rater 2 evaluated performances 3, 4, 5, and 6, etc.). The final rater was linked to the first rater (i.e., final rater evaluated performances 39, 40, 1, 2.) whereby connectivity is achieved in the assessment network, allowing for the direct comparison of raters across all performances. This specific network design was found to provide the best model fit and smallest standard errors for an incomplete assessment network (Wind, Engelhard, & Wesolowski, 2016). When the evaluations were completed and all data was collected, the negatively worded items were reverse coded before undergoing empirical analysis in order to establish similar directionality in data analysis.
Results
Variable map
The variable map (see Figure 2) is a graphical representation of the analyzed data that operationally defines the latent construct of this study (i.e., performance achievement of high school wind bands). The variable map displays the three facets that were used in the model: (a) item difficulty; (b) performance ability; and (c) rater severity. The logit scale, found in the first column, can be considered the “ruler” for comparative measurement of each of the facets. The placement on the logit scale represents data that is interval level in nature, enabling one to make direct comparisons between data across the separate facets. The second column displays each of the performance measures in order from highest achieving at the top of the column to lowest achieving at the bottom of the column. Each individual performance is represented by a single asterisk on the variable map. The measures ranged from 2.87 to –2.43 logits (M =– .03, SD = 1.22, N = 40). The third column of the variable map displays the severity measurements of the raters from most severe at the top of the column to least severe at the bottom. The measures ranged from 2.05 to −1.76 logits (M = .00, SD = .97, N = 20). The fourth column displays the item difficulty measurement from most difficult near the top of the column to least difficult near the bottom. The measures ranged from 1.24 to −1.43 (M = .00, SD = .75, N = 39).

Variable map.
Summary statistics
Summary statistics are provided in Table 1. The data represented in this table indicates the overall significant difference between performances, χ2 = 1516.6, p < .01, raters, χ2 = 976.7, p < .01, and items, χ2 = 606.2, p < .01. The reliability of separation statistic can be conceptualized as the sensitivity of the measurement instrument to distinguish, or separate, individual elements within a particular facet and its ability to reproduce the logit locations. The reliability of separation for performances (Rel = .98), raters (Rel = .98), and items (Rel = .94) provides confidence to imply that there is enough separation to confirm the construct validity of the measurement instrument. The mean square fit (MSE) statistics (Infit MSE = .99 and Outfit MSE = 1.01) are close to the expected value of 1.00, demonstrating reasonably good data fit to the model based upon Wright and Linacre’s (1994) acceptable range for parameter-level mean square statistics (0.50 to 1.50).
Summary statistics from the PC-MFR model.
p < .01.
Calibration of performances
Table 2 provides a detailed report of the analysis of performances. The measure statistic indicates the logit scale location of the performance’s achievement. A higher measure number indicates greater performance achievement and a lower measure indicates a lower performance achievement. The highest achieving performance was performance 20 (2.87 logits), and the lowest achieving performance was performance 11 (−2.43 logits). The infit MSE statistic was used to determine evidence of performance misfit. Wright and Linacre (1994) and Engelhard (2009) indicate that the best range for data misfit at the element level is between .80 and 1.20 logits for high stakes assessments. Performances demonstrating misfit due to over-fit (fit indices greater than 1.20 logits) include performances 4, 6, 7, 14, 18, 25, 35, 39, and 40. Performances demonstrating misfit due to under-fit (fit indices less than .80 logits) include performances 3, 9, 12, 22, 23, 24, 30, 31, 32, 33, 34, and 38.
Calibration of performance facet.
Note. Presented in measure order from highest achievement to lowest achievement.
Calibration of raters
Table 3 provides a detailed data analysis of raters. The measure indicates the logit scale location of the rater severity. A higher measure number indicates greater rater severity and a lower number demonstrates lower rater severity (i.e., leniency). The most severe rater was rater 8 (observed average = 2.13, logit measure = 2.05) and the least severe rater was rater 2 (observed average = 3.19, logit measure = −1.76). Raters 6, 12, 14, and 20 demonstrated misfit as evidenced by Infit MSE scores greater than 1.20. This substantively indicates ratings that were too sporadic (i.e., too unpredictable) to fit the model. Raters 3, 4, 17, and 19 demonstrated misfit due to an Infit MSE score less than .80. This substantively indicates ratings that were too muted (i.e., too predictable) to fit the model.
Calibration of rater facet.
Note. Presented in measure order from most severe to least severe.
Calibration of items
Table 4 provides a detailed data analysis of rating scale items. The measure indicates the logit scale location of the difficulty of each item. A higher logit measure number indicates a more difficult item and a lower logit measure indicates a less difficulty (i.e., easier) item. The most difficult item was item 24 (Intonation inaccuracy throughout the performance; observed average = 1.92, logit measure = 1.24). The easiest item was item 11 (Tempo is appropriate for selection; observed average = 3.05, logit measure = −1.43). Items 1, 2, 5, 10, 12, 28, 29, 33, 34, and 38 demonstrate misfit due to an Infit MSE statistic less than .80, which demonstrates items that are too predictable when rating performances. Items 6, 11, 19, 23, and 25 demonstrate misfit due to an Infit MSE statistic above 1.20, which demonstrates items that are too unpredictable at distinguishing between higher achieving and lower achieving performances. Items demonstrating misfit were removed from the final rubric.
Calibration of item facet.
Note. Presented in measure order from most difficult to least difficult.
Rating scale category diagnostics
As the data from the facet calibrations began to help shape the final rating scale, a detailed analysis of the function of each item was performed as a result of employing the partial credit (PC) version of the model. A rating scale structure cannot be complete until, as Linacre (2002) suggests, careful steps toward optimization of the rating scale category structures are taken. The four-category Likert-type rating scale construction was carefully modified based on individual item diagnostic statistics in order to address and improve issues of construct validity that are associated with the rating scale as a whole. Analysis of the rating scale structure is what Bond and Fox (2007) claim is vital in order to clarify the meaning of and improve the usability of the measure. The item behavior statistics can be viewed in Table 5. This table defines the empirical data that was used to determine final construction of the scale in its reverse coded form to preserve directional consistency for positively and negatively worded items.
Item behavior of category usage, average observed and expected measures, and Outfit MSE.
Note. Category 1 = “strongly disagree”; Category 2 = “disagree”; Category 3 = “agree”; Category 4 = “strongly agree”.
Category 1 = “strongly agree”; Category 2 = “agree”; Category 3 = “disagree”; Category 4 = “strongly disagree”.
Violation of monotonicity.
The analysis of the rating scale structure included four stages. The first stage of analysis included category eliminations based upon the analysis of category frequency counts for each item of the original scale. Linacre’s (2002) recommendation for this process is 10 uses per category. However, due to the smaller scale of the study, 10% of total responses per item (8 uses) was used as an acceptable cut point for each category. This resulted in the collapsing of adjacent categories in 18 items in order to guard against irregular skew and distributions. These categories include: item 3 (category 4), item 7 (category 1), item 8 (category 4), item 9 (category 1), item 13 (categories 1), item 15 (category 1), item 16 (category 1), item 17 (categories 1 and 4), item 20 (category 4), item 21 (category 1), item 24 (category 4), item 26 (category 1), item 30 (categories 1 and 4), item 31 (categories 1 and 4), item 32 (category 4), item 35 (category 1), item 36 (category 1), and item 37 (category 4).
The second stage of analysis included analysis of the Outfit MSE statistics for each of the categories. Values greater than or equal to 2.00 indicate excessive unpredictability in the rating categories that were being used unexpectedly by raters (Linacre, 2002). After eliminating items from misfit and categorical frequency the only item found exceeding the value of 2.00 was item 18 (category 2). This category was therefore collapsed with its adjacent category, category 1.
The third stage of analysis included analysis of monotonicity. Monotonicity refers to the unidirectional movement up the logit scale as the categories of an item increase in difficulty (i.e., proper step ordering). Category 2 of items 15 and 26 was collapsed with its adjacent categories as a result of violations of monotonicity (e.g.. Item 15, category 1 shows a logit scale rating of .77 while category 2 shows a logit rating of .21 which is less than .77, violating the unidirectional ordering of categories). This continuous advancement of step calibration within the measurement tool is vital to strong construct validity (Andrich, 1996).
The final phase of category diagnostics involved reviewing the remaining categories in order to show large enough measure difference between categories. The threshold of .70 logits was used as the difference in observed average measures to ensure a proper separation of categories. Categories 1 and 2 of item 14 revealed a difference of .17 logits. Categories 1 and 2 of item 39 revealed a difference of .50 logits. As a result, it made substantive sense to collapse the two categories due to a lack of clear definition and separability between categories. The resulting scale with remaining items and collapsed categories is provided in Figure 3, which demonstrates the adjustments made using Linacre’s guidelines for optimizing rating scale structure.

Final rating scale.
Rubric design
Following the finalization of the wind band rating scale, the results were analyzed and reviewed in order to develop a rubric that would better align with the empirical results of the rating scale. This rubric was developed to provide specific category descriptors that would allow all scale users to be able to clearly understand the resulting performance evaluations and provide information for improvement in the future. Three content experts reviewed the rubric for content and face validity. The rubric was formed using Vagias’ (2006) anchors from the categories of Affect (item 15), Appropriateness (item 9), Barrier (items 22, 37, and 39), Dichotomous (items 30 and 31), Frequency (items 7, 8, 9, 13, 14, 17, 18, 20, 21, 26, 32, 35, and 36), and Problem (items 3, 4, 5, 23, and 27). The final rubric is provided in Figure 4. The rubric should be viewed as a tool for feedback in order to guide future study by the ensemble. The rubric reveals basic categories of musicianship and the anchors that describe the ensemble’s ability within that category. The anchor statements were developed in close relation to the statements in the rating scale to increase the reliability of feedback when the rubric is used in coordination with the rating scale.

Final rubric.
Conclusion and further research
The first research question includes the evaluation of the psychometric quality of the scale developed for the assessment of wind band performance achievement. The specific psychometric qualities that were the focus of investigation were validity, reliability, and precision of the measurement instrument. Reliability is viewed in the light of how the information gained from the data is used to distinguish between the quality and/or the amount of the latent trait, which in this case is performance achievement of wind bands. Evidence of strong reliability was seen in the high statistic for reliability of separation of items, performances, and raters (RELitems =.94; RELperf =.98; RELraters =.98). Evidence of precision is observed in the standard errors and related rating scale category diagnostics. Both the standard errors for raters and items are small, which provides strong support for high levels of precision for the scale. Furthermore, the carefully evaluated rating scale category data and related optimization provides a more precise response to the items. The scale was able to clearly distinguish between performances of varying achievement within the ensemble as displayed along an equal-interval continuum seen in the variable map. The ability of the scale to achieve this goal with reliability and precision is evidence of strong validity and assumptions for the reproducibility in further studies.
The second question guiding this research study investigated item fit to the model and variability in difficulty. These ideas specifically hone in on the content and construct validity issues facing this measurement tool. After the items were analyzed, 14 of the original 39 item stems were removed for misfit. In other words, these 14 items caused a violation of the five requirements for Rasch measurement and do not appropriately fit the model to be included in a measurement tool demonstrating invariant measurement. It is probable to consider that these 14 items demonstrate multidimensionality that confounds their effective use within the scale. Due to the items’ nature of being either underfit/too predictable, or overfit/too unpredictable, they are items that would require either modification or deletion from the final scale. Because of the focus and scope of this specific study, these 14 items were deleted from the final scale. However, it is suggested that in future studies these items should be the focus of careful review such that they can be modified and transformed in an attempt to produce unidimensional items that exhibit appropriate model fit.
The third research question focused on categorical structures within each item, therefore, the remaining 25 items were reviewed for proper categorical behavior based on categorical usage, outfit, and sufficient logit separation of categories. Review of this data demonstrates strong evidence that all items in the scale did not share structural equality due to violations of monotonicity, disproportionate and skewed categorical usage, idiosyncrasies in model predictability, and poor logit separation. The resulting scale demonstrated items with categorical structures ranging from two to four categories within the Likert-type scale. This analysis provides the resulting scale with model fit that also improves precision of the scale structure thereby improving the overall construct validity of the measure.
It is important to remember the practical application of this information, or as Messick (1989) notes, the consequential validity. Consequential validity frames the implications this process has on its practical application for implementation and assessment. One of the most important initial steps needed to make this a usable assessment for large ensemble evaluations would be the development of performance standards or cut scores. These would be the defining points on the logit score of the performance achievement used to determine if an ensemble received a superior rating verses an excellent rating. In order for this to be done in an effective manner, it would require the cooperation of subject matter experts, psychometricians, and policy makers. This collection would have the task of devising the appropriate benchmarks for achievement that define the achievement levels needed for effective and practical use as an assessment. It would be unwise for a reader of this study to consider using the sum scores of this evaluation in place of the careful work of the collection of professionals listed above. It has been shown in the study that the items themselves are not equal in category, as the ordinal interpretation would imply. Rather, the items vary in difficulty across categories and items. Summative scoring based on the final scale produced in this study would negate the validity of the scale.
Further research in the area of item stem development would also be necessary for filling in some of the gaps in logit locations of the item stems. In viewing the variable map (Figure 2), the item column reveals several gaps in placement along the continuum. For example, a gap can be seen between the item row that contains item 2 and the item row that contains item 18. This indicates areas on the logit continuum where there is less ability to distinguish between performances. Further testing of new item stems that could be used to help fill in these gaps, as well as rewritten item stems that were removed from the final scale, may help complete a series of items that filled in the scale with fewer gaps. This would continue to strengthen the ability of our measure to distinguish between performances at all levels, thus providing stronger validity and precision of this measure.
Finally, it is important for teachers viewing this study to glean some practical understanding of performance evaluation from the given results. To do this, an investigation of which items were rated most difficult and easiest to endorse is needed. The easiest item to endorse in the final scale is “key signatures and key changes are performed correctly.” This may indicate that during these large ensemble evaluations proper note performance is considered to be one of the most basic skills evaluated. If a teacher is striving in their classroom to only be able to teach the pitches and rhythms then they are missing out on the advanced concepts of music. However, the hardest to endorse items, 24 and 37, describe ensemble intonation and balance. Teachers who maintain an effective classroom focus on these topics are likely to be those teachers who are leading performances that achieve at the highest levels. While it is true that one cannot forgo the teaching of notes and rhythms for a focus on intonation and balance, it is important to remember these ideas during repertoire selection. Repertoire selected should be appropriate so that the notes and rhythms are attainable in an amount of time that yields the teacher appropriate time to cover more difficult topics such as intonation and ensemble balance.
Footnotes
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
