Abstract
The content that test-takers attempt to convey is not always included in the construct definition of general English oral proficiency tests, although some English-for-academic-purposes (EAP) speaking tests and most writing tests tend to place great emphasis on the evaluation of the content or ideas in the performance. This study investigated the relative contribution of linguistic criteria and the elaboration of speech content to scores on a test of speaking proficiency. A speaking test was designed and administered to Japanese undergraduates to determine what criteria English teachers associate with general oral proficiency. Nine raters were recruited to rate 30 students’ monologues on three topics, using intuitive judgments of oral proficiency (referred to as Overall communicative effectiveness). Following this, they assigned scores to the monologues using five criteria: Grammatical accuracy, Fluency, Vocabulary range, Pronunciation, and Content elaboration/development. The raters were also asked to provide open-ended written comments on the factors contributing to their intuitive judgments. Statistical analyses of the scores – Rasch measurement, multiple regression, and multivariate generalizability (G) theory analysis – revealed that Content elaboration/development made a substantive contribution to the intuitive judgments and composite score. The present study enriches our understanding of general oral proficiency and the construct definition of proficiency tests.
Keywords
In performance testing, analytic rating scales are frequently used to assess the components of test-takers’ performance. It is common for the criteria included in the scales to vary according to the purpose of the test or what stakeholders wish to know. Interestingly, the content or ideas that the test-taker attempts to convey are almost always explicitly assessed in tests of writing performance, irrespective of the purpose of assessment. For example, the criterion of ideas is included in the well-known six-trait analytical writing-scoring guide developed by Spandel and Stiggins (1997). Furthermore, the ubiquitous analytic scale developed by Jacobs et al. (1981) not only assesses content but also gives it more weight than language use or vocabulary. In assessing speaking performance, however, speech content is not always included in analytic scales. Instead, the linguistic components of oral performance, including pronunciation, accuracy, and vocabulary, are most commonly assessed, implying that the oral performance of test-takers is simply taken as representative of their linguistic ability.
Recently, research has demonstrated that raters of several EAP speaking tests pay considerable attention to the content of test-takers’ speech (Brown, 2007; Brown, Iwashita, & McNamara, 2005; Eckes, 2009). Accordingly, the analytic rating scales used in the TOEFL iBT evaluate the ability to elaborate monologic speech through the category topic development (ETS, 2008). Brown (2007) claims that raters may believe that the maturity of ideas as well as their linguistic sophistication is relevant to success at university. Although there is research on raters’ perception of the relative importance of rating criteria in some other specific purpose contexts (e.g. McNamara, 1990), little is known about the extent to which speech content contributes to rater judgment of oral performance measured by general English proficiency tests. Thus, the present study administered an oral proficiency test to undergraduate university students in order to investigate the criteria underlying raters’ holistic and intuitive judgments.
Literature review
Since rating scales form a part of the construct definition assessed in a test (Luoma, 2004), what is included in them has a profound impact on test-takers’ scores and the validity of the performance test. Thus, scales need to be developed with caution so as not to lead to erroneous inferences about test-takers’ language proficiency and, as a result, unfair evaluations. In developing rating scales, test designers must decide what the test will assess, because it is virtually impossible to assess all the potential components of performance in a single test (Wigglesworth, 2008). In the process of defining the construct to be assessed, the test’s purpose should be the starting point, since it will be the basis of further decisions (Fulcher & Davidson, 2009). For example, if testers need information on only one discrete component of speaking performance (e.g. a particular phoneme) for their intended purpose (e.g. diagnosis), it is justifiable that the construct definition be restricted to the target component, excluding other components (Purpura, 2008). Once the test’s purpose is determined, testers need to provide a specific construct definition, usually on the basis of a course syllabus or needs analysis (Bachman & Palmer, 2010). These sources of construct definition, however, are not applicable unless there is a specific target language use (TLU) domain. For this reason, in the case of proficiency tests whose purpose is to assess overall proficiency levels, the construct definition is mainly based on theoretical models of language ability (Bachman & Palmer, 1996, 2010; Luoma, 2004).
Although theoretical models of language ability (e.g. Bachman & Palmer, 1996; Canale & Swain, 1980) are a major source of construct definitions for proficiency tests, these models have some caveats, one of which is that they do not necessarily take satisfactory account of the non-linguistic components of communicative ability and how they affect language performance. Luoma (2004) suggests that the models attempt to explicate the linguistic components of communicative ability in great detail but show a static view of communication that distracts our attention from other aspects. The narrow perspective of these models appears to influence performance testing. Hughes (2002) points out the discrepancy between the focus of speaking in general and that of oral tests, arguing that ‘speaking is not naturally language focused, rather it is people focused. Language testing is not naturally people focused but by its nature will tend to be language focused’ (p. 77). This perspective is echoed by McNamara (1996), who states that most general-purpose proficiency tests are language performance tests in the weak sense of being designed to elicit a language sample from test-takers and focus on the linguistic quality of their oral performance rather than to assess performance on real-world criteria or task fulfillment. To address this situation, the use of practical frameworks, together with the theoretical models, is advised when defining the construct of performance assessed in the test (Luoma, 2004). Otherwise, a factor significantly contributing to oral proficiency may not be attended to at all. Research has suggested that some non-linguistic components of communicative ability affect test-takers’ oral proficiency levels; these components include content meaningfulness or sophistication (Brown, 2007; Brown, Iwashita, & McNamara, 2005; Eckes, 2009) and non-verbal behaviors (Ducasse, 2009; Ducasse & Brown, 2009; May, 2009; Orr, 2002). Thus, it is necessary to further explore the construct of general oral proficiency, or what it means to be able to communicate effectively, without relying solely on existing communicative models.
The relative contributions of various criteria (i.e. the weights allocated to specific criteria in a scale) are also a crucial issue when testers attempt to develop rating scales (Wigglesworth, 2008) or provide a single total score based on componential scores obtained from analytic scoring (Bachman, 2004). Unfortunately, communicative language models do not have clear implications for weighting different criteria. In fact, a number of analytic scales for speaking tests appear to attach equal weight to all criteria; as a result, the various components of oral performance are treated equally. Weights can be determined either by testers’ beliefs about the relative importance of the criteria (nominal weights) or by analysis of empirical data applying G-theory (effective weights) (see Bachman, 2004). Even though there are studies on weighting based on test scores (e.g. de Jong & van Ginkel, 1992; Sawaki, 2007), comparisons in these studies were narrowly restricted to the linguistic components of oral performance. As noted above, there is little empirical research on the relative effects of linguistic criteria and non-linguistic criteria on test-takers’ oral performance.
In addition to theoretical models of communicative language ability, raters’ perception of test-takers’ oral performance can be another valuable source of components to be measured in the tests, and of their weights. Pollitt and Murray (1996) argue that ‘the starting point for scale development should surely therefore be a study of the perceptions of proficiency by raters in the act of judging proficiency’ (p. 76). They examined rater judgments of test-takers’ speaking performances and found that the features the raters heeded varied with the test-takers’ proficiency levels. Thus, they concluded that rating scales should assess only salient components at each level, suggesting that assessing pronunciation of high-proficiency test-takers may be pointless, since raters do not perceive it to be salient in oral performance. Other research has also examined raters’ perspectives on components of performance and suggested that the results should be incorporated into the test’s construct definition (Ducasse, 2009; Ducasse & Brown, 2009; Iwashita & Grove, 2003). In a similar vein, Bachman and Palmer (2010) argue that criteria used for tests should correspond closely to those typically used by stakeholders to assess performance. If the test’s purpose is to assess general English proficiency levels, testers must develop criteria corresponding to the criteria English speakers actually use (whether consciously or subconsciously) to judge L2 speakers’ oral proficiency in general. Therefore, it is necessary to investigate whether or not the components of performance frequently measured in a proficiency test are the same as those that real-life interlocutors or listeners are oriented toward when judging speakers’ communicative effectiveness.
To sum up, the validity of assessment criteria solely derived from theoretical models is called into question, given that the models do not necessarily delineate non-linguistic components of oral proficiency. Although some non-linguistic features have been found to be influential in oral performance, little research has focused on the importance of non-linguistic criteria vis-à-vis linguistic criteria. This provides a justification for scrutinizing the criteria – linguistic and non-linguistic – underlying raters’ judgments of oral proficiency as an additional source of construct definitions. This research is considered to be essential to rating scale development, since assessment criteria need to align with those used by speakers in the TLU domain.
The present study therefore addresses the following research question: What are the relative contributions of linguistic criteria and speech content to scores on a test of speaking proficiency? The linguistic criteria to be assessed in the present study are Grammatical accuracy, Fluency, Vocabulary range, and Pronunciation, because they are frequently assessed in speaking tests (Underhill, 1987) and are considered to be fundamental components of oral proficiency. Speech content was defined as Content elaboration/development, indicating the degree to which the test-taker is conveying relevant and well-elaborated/developed ideas on given topics. This definition was developed on the basis of the concept of topic development in the TOEFL iBT speaking rubrics (ETS, 2008). Furthermore, some research has shown that relevance and sophistication of ideas are features that raters of some EAP speaking tests heed as a content criterion (Brown, 2007; Brown et al., 2005).
Method
Participants
Test-takers in the present study were 156 undergraduate students (71 female and 85 male) studying at a university in Japan. They consisted of 69 freshmen, 86 sophomores, and one senior, taking intermediate and advanced general English language courses offered at the university; 135 and 21 students were enrolled in the intermediate and advanced courses respectively. Among them, 80 students were at EIKEN 1 Grade Pre-2 or higher, demonstrating that they were at least able to discuss everyday concerns in English. The aim of the courses was to develop the students’ overall English communication skills; students at both course levels were regularly required to discuss various issues and had ample opportunity to practice expressing their thoughts orally. In the advanced course, the attendees were also required to discuss global, political, and social issues. Since the regular classes were conducted in computer-assisted language learning (CALL) rooms, where a desktop computer was available to each individual, the participants were familiar with computer-mediated learning activities such as peer-to-peer oral interactions through the computer. They agreed voluntarily to participate in the study as test-takers.
Nine participants were recruited to rate oral performances, as shown in Table 1. Raters had English-teaching experience at the secondary or post-secondary levels (not at the test-takers’ university) and included eight in-service teachers and one PhD candidate specializing in applied linguistics. Among these participants, there were four Japanese raters and one Taiwanese rater, all with high levels of English proficiency, and four native English-speaking (NS) raters, from the USA, the UK, and Ireland. All the non-native English-speaking raters had a master’s degree in education or TESOL and had received instruction in English language education. The NS raters were all assistant language teachers and English instructors working in Japan, and their amount of teaching experience was the same as their number of years spent in Japan. Four raters had prior experience rating L2 learners’ oral performance in various situations: a public school teachers’ examination (Rater A), an in-class speaking test (Raters C and G), and a proficiency test (Rater D). These raters were recruited primarily on the basis of their willingness to participate in this research and rate L2 learners’ oral performance. Monetary compensation was offered for their cooperation.
Raters’ background information (n = 9)
Note: EIKEN Grade 1 and Grade Pre-1 are recognized as sufficient for admissions to international graduate and undergraduate programs. Grade 2 is designated as a benchmark for high-school graduation by the Japanese Ministry of Education (STEP, 2008).
Instruments
A speaking test was developed to elicit oral performances from the university students. Three prompts to elicit monologic performance – personal opinions and attitudes about general issues – were selected from a TOEFL textbook (Pierce & Kinsell, 2006) in light of their proficiency levels and the regular activities they were engaged in. (The task prompts are presented in Appendix 1.) This task format was chosen because it is a common format in study settings (Luoma, 2004) and because test-takers need to demonstrate their ability to use language effectively at some length (Underhill, 1987). Thus, this type of elicitation task was considered appropriate for collecting speaking data from the participants. The test was designed to be computer delivered; that is, test-takers were asked to read prompts presented on the computer screen and deliver speech using a headset. A computer-delivered test format was adopted in the study because of the participants’ familiarity with this format and because of its practicality for recording a large number of monologues.
Rating criteria and scales were also developed for the assessment (Appendix 2). The criteria provided to the raters were (a) Overall communicative effectiveness, (b) Grammatical accuracy, (c) Fluency, (d) Vocabulary range, (e) Pronunciation, and (f) Content elaboration/development. A detailed definition of Overall communicative effectiveness was not specified; instead, the raters were asked to rate this criterion using their intuition or impressions of the test-takers’ overall oral proficiency. In contrast, concise definitions of the other criteria were provided, in order to explain what features the raters had to assess. This study employed numerical rating scales, which divide the top and the bottom ends by number without descriptions of the levels. Descriptors were not developed because the primary purpose of the study was to examine the raters’ perception, and not to achieve high inter-rater reliability. The possible responses were on a five-point scale ranging from 1 (unsatisfactory) to 5 (completely satisfactory). This structure is commonly applied in order to consistently distinguish proficiency levels (Luoma, 2004).
A rater’s manual was designed in order for the raters to train themselves in preparation for scoring. Face-to-face rater training was not incorporated since some raters resided in distant locations and it was unfeasible to administer this type of training. The manual contained information on the purpose of assessment, the general background of the test-takers, rating procedures, and rating criteria. It was stipulated that the purpose of the assessment was to research how raters of a speaking test assign scores to test-takers’ oral performance. The information about the test-takers provided to the raters was that they were EFL learners at a private university (freshmen and sophomores) attending intermediate and advanced English courses offered at the university. Information on the courses (e.g. course objectives) was not given. Three sample performances were prepared to accustom the raters to the test format and rating criteria. They consisted of random samples from speaking data collected from the test but not analyzed in the study. The manual also contained scores and annotations that I added to the samples in order to help the raters internalize the analytic rating criteria. Scores for Overall communicative effectiveness were not provided, since the scores were to be given on the basis of the raters’ intuitive judgments of performance.
Procedures
The speaking test was administered during regular class time in the CALL rooms where the test-takers regularly took classes. Prior to data collection, test-takers were informed that the test would be administered purely for research purposes and would not affect their grades in their courses. After being presented with information on the test’s purpose and data-collection procedures, the test-takers practiced for the task and were asked to deliver monologues on the three prompts. They were provided with 30 seconds for preparation and one minute for speech delivery. Their performances were audio-recorded directly onto the computer. Following data collection, students who had failed to record three monologues were excluded from the analysis. Next, 15 intermediate and 15 advanced students were randomly selected (these students’collective profile is presented in Appendix 3). This method of selection was intended to include students with a wide range of proficiency levels. Ninety monologues (30 students × 3 monologues) were ultimately rated.
Each of the raters was given a USB memory stick storing the audio data and asked to submit scores within one month, so that they could assess the monologues at their own pace. Prior to the rating, they were asked to read the rater’s manual and listen to the benchmarks, checking the assigned ratings and annotations written in the manual. This form of rater training was expected to increase intra-rater reliability rather than inter-rater reliability, because this study did not require perfect agreement among the raters as long as their scoring was internally consistent. Each monologue was listened to twice by each rater. In the first round, the raters were asked to assess Overall communicative effectiveness. After assigning scores to all the performances, they listened to the monologues again, rating them on the five analytic criteria. This procedure was designed to effectively differentiate intuitive judgments from analytical judgments. Following the scoring, raters were asked to write open-ended comments on the feature(s) they heeded while they were rating Overall communicative effectiveness.
Data analysis
The test scores assigned to the 90 monologues by the nine raters were statistically analyzed using Rasch measurement, multiple regression, and multivariate G-theory (the scores on the three different prompts were aggregated). The raters’ written comments were also scrutinized to complement the statistical analyses.
The study first employed many-facet Rasch measurement to confirm intra-rater reliability and explore local independence of the criteria. Initially, rater fit statistics – outfit and infit – were examined to confirm intra-rater reliability. Rater fit indices indicate the extent to which the raters’ actual ratings match those expected from the Rasch model; outfit statistics are estimated based on the sum of squared standardized residuals of all ratings, including outliers, whereas infit statistics are an information-weighted sum and are less sensitive to extreme ratings (Bond & Fox, 2007). The present study applied Myford and Wolfe’s (2000) standard for accurate raters, which is that raters applying the scales inconsistently will have outfit and infit mean square indices greater than 1.30.Following this, the fit statistics of the criteria were examined to see whether there would be any highly interdependent criteria; that is, the major concern was with overfit as indicated by excessively low fit indices, showing lack of local independence. The study applied the acceptable range of outfit and infit mean square values suggested by Bond and Fox (2007), which is between 0.75 and 1.30. More specifically, a fit value higher than 1.30 would be considered underfit (violation of unidimensionality or low item discrimination), and a fit value lower than 0.75 would be overfit. The analysis was conducted using FACETS Version 3.65.0 (Linacre, 2009).
Next, multiple regression analysis was performed to examine the degree to which each criterion would predict Overall communicative effectiveness. This analysis was intended to demonstrate the relative importance of the criteria to the raters’ intuitive judgments of oral performance. For this analysis, scores for each criterion assigned by the nine raters to each of 30 test-takers’ three monologues were combined; thus, each criterion was composed of 810 scores. The five analytic criteria were chosen to be the independent variables (IVs), and Overall communicative effectiveness was the dependent variable (DV). After the assumptions of the analysis were evaluated, the standardized regression coefficients were scrutinized to determine the contribution of each IV to the DV. The multiple regression analysis was conducted using SPSS Version 17.0.
Third, multivariate G-theory was applied in order to examine the effective weights of the five criteria. As noted earlier, effective weights represent the relative statistical contribution of each criterion to a composite score (Bachman, 2004; Brennan, 2001a). The G-theory design incorporated in the study was a two-facet fully crossed design where all the test-takers were assigned the same three prompts and the performances of each prompt were rated by the same nine raters on the five fixed criteria. This analysis excluded Overall communicative effectiveness, because the purpose was to explore the relative contribution of the five criteria to the composite score. The composite was composed of the five criteria alone, and was different from Overall communicative effectiveness, which might embrace features not assessed through the analytic rating scales. The effective weights were determined by the nominal weights, the variance of each component score, and the covariances among the component scores (Bachman, 2004; Brennan, 2001a). The nominal weights were equal across the five criteria; that is, they each had a 20% weight. Computer software mGENOVA Version 2.1 (Brennan, 2001b) was used for the multivariate G-theory analysis.
Finally, the study examined the raters’ written comments, which described the features to which they paid the most conscious attention while rating Overall communicative effectiveness. The comments were sorted into the five analytic criteria assessed in the study and tallied to gain an understanding of the raters’ focus. Since their comments were provided in an open-ended format, some of the raters’ comments were on features that were not directly related to any of the rating criteria. Those features were categorized differently from the five criteria and examined individually.
Results
As discussed above, the intra-rater reliability and local independence of each criterion were analyzed using Rasch measurement. Rasch person separation reliability was .99, indicating that the test-takers were widely separated in terms of their levels of oral performance. Table 2 presents the nine raters’ infit and outfit mean square statistics and single rater-rest of the raters (SR/ROR) correlations. The result shows that all the raters’ fit statistics were less than 1.30, suggesting that the raters were able to comply with the given rating criteria and distinguish the test-takers according to the components being assessed. This finding is supported by the similar magnitude of SR/ROR correlations for all the raters (ranging from .31 to .44).
Rater fit statistics and SR/ROR correlations
The scores for each criterion were closely analyzed to explore interdependent patterns in the criteria. Table 3 displays the criteria measurement report, including infit and outfit mean square statistics and point-measure correlations. The measure of each criterion indicates that Pronunciation was scored most leniently (−0.41 logits) and Fluency and Vocabulary range were scored most severely (0.24 and 0.23 logits, respectively). Separation reliability was .96, indicating that the criterion difficulties were widely spread and different. The infit and outfit indices fell within the acceptable range adopted in the study (0.75 to 1.30). This result suggests that all criteria demonstrated neither totally unexpected nor totally interdependent scoring patterns; that is, both the holistically and the analytically assigned scores indicate unidimensionality and local independence. The positive point-measure correlation coefficients (in the Corr. PtBis column) support this result. Overall communicative effectiveness was expected to be an overfitting criterion, since it was considered to be a unitary criterion encompassing the other five criteria.However, its infit and outfit mean square indices (0.88 and 0.87, respectively) indicate that this was not strongly so. The Rasch analysis instead demonstrated that Overall communicative effectiveness contributed to a unidimensional construct as an independent criterion similar to the other criteria. It is thus unlikely that this holistically assessed criterion was a simple summary of the scores for the analytic criteria. For this reason, the Rasch analysis did not find any criteria demonstrating similar or identical scoring patterns with or showing close relationships with Overall communicative effectiveness.
Criteria measurement report
Note: Model SE = standard error; MNSQ = mean square; Corr. PtBis = point-measure correlation.
Table 4 displays the correlations among the criteria and the standardized regression coefficients, demonstrating the extent to which the individual IVs predict the DV. All the correlation coefficients between the IVs and the DV were statistically significant (p < .001) and considered large according to Cohen’s (1992) rules of thumb for effect sizes (r ≥ .50). Nevertheless, their coefficients ranged from .55 to .77, indicating that Pronunciation appeared to have the weakest relationship with Overall communicative effectiveness among the criteria. In contrast, Content elaboration/development and Fluency showed the highest correlations with the DV. With regard to prediction, R for regression shows the statistically significant difference from zero, F(5, 804) = 331.56, p < .001, which indicates that the five analytic criteria significantly predicted Overall communicative effectiveness. In addition, the value of R2 (.67) shows that 67% of the DV was predicted by the IVs. In other words, approximately one-third of Overall communicative effectiveness was explained by features that were not assessed by the rating scales employed in the study. Each of the IVs also separately significantly predicted the DV (p < .05). The standardized regression coefficients show that Content elaboration/development was the strongest IV (β = .42), followed by Fluency (β = .25), indicating that rater judgments of Content elaboration/development and Fluency appeared to most strongly affect their intuitive judgments of Overall communicative effectiveness.Conversely, the other three IVs turned out to be less important for determining Overall communicative effectiveness; Pronunciation was the weakest (β = .07). The degrees of contribution of Grammatical accuracy and Vocabulary range to Overall communicative effectiveness were similar and seemed to be small (β = .10 and .11, respectively).
Multiple regression of the effect of the five criteria on overall communicative effectiveness
Note: Overall = Overall communicative effectiveness (DV); Vocab = Vocabulary range; Pronun = Pronunciation.
p< .001. ** p< .01. * p< .05.
The test scores were further analyzed using multivariate G-theory. The composite phi coefficient of the test (nine raters and three prompts) was .91. This high composite coefficient indicates that it is highly reasonable to form a composite of the scores and that multiple components represent an underlying dimension (Webb, Shavelson, & Maddahian, 1983). Therefore, the result shows that the five criteria were reasonably highly correlated, suggesting that the composite of their scores is highly interpretable. Table 5 displays the results of the composite score analysis, including the nominal weights and effective weights. It shows that Fluency and Content elaboration/development accounted for 24.47% and 24.08% of the composite score variance respectively, while the effective weights of the other three criteria were smaller than the nominal weights, ranging from 14.57% (Grammatical accuracy) to 19.16% (Vocabulary range). An effective weight becomes small when the variance of the criterion and its correlations with other criteria are small. Examining the variances and correlations of the five criteria, it became clear that those of Fluency and Content elaboration/development were substantially greater than those of the other three criteria. Thus, Grammatical accuracy, Vocabulary range, and Pronunciation, whose universe score variances and correlations were smaller than those of the other two criteria, were considered to contribute less to the composite.
Nominal weights and effective weights of each criterion (%)
Note: The results are based on the D study for nine raters and three prompts. Univ score var = universe score variance.
Lastly, the raters’ retrospective comments were categorized according to the five analytic criteria and tallied (see Table 6). The result indicates that the raters paid the most attention to Fluency out of the analytic criteria when rating Overall communicative effectiveness; six raters’ comments referred to this criterion. Following Fluency, the next most-heeded criteria were Pronunciation and Content elaboration/development, which were focused on by two raters each. Unlike the above-mentioned three criteria, Grammatical accuracy and Vocabulary range of the test-takers’ monologues did not appear to be consciously heeded by the raters. It should be noted that Raters B and D clearly stated that it is not particularly important for the speakers to avoid grammatical mistakes or use complex syntactic structures. Rater B also commented that Vocabulary range is not crucial as long as the test-takers demonstrate command of sufficient and appropriate lexical items to express their ideas in given contexts. Five raters (A, C, D, E, and I) alluded to features that were irrelevant to any of the analytic criteria incorporated in the study. Raters D and E mentioned similar components: the amount of information that the test-takers provided and the quality of their voices. Raters A and C rated Overall communicative effectiveness focusing on a broader perspective beyond linguistic abilities, such as what the test-takers would communicate using English. Raters D and I respectively mentioned variation of expressions and schematic organizational skills.
Criteria heeded by the raters
In summary, the many-facet Rasch analysis revealed an acceptable level of intra-rater reliability and unidimensionality of criteria. The regression coefficients of the criteria indicated that Content elaboration/development was the strongest predictor of Overall communicative effectiveness. The multivariate G-theory analysis and the raters’ written comments also demonstrated the important role of Fluency, both in the composite score and in their conscious attention. In contrast, as the analyses of the test scores and raters’ written reports showed, Grammatical accuracy, Vocabulary range, and Pronunciation were less closely connected to raters’ intuitive judgments of test-takers’ oral performance than were the other two criteria.
Discussion
The raters’ infit and outfit mean square indices demonstrate that intra-rater reliability was acceptable, suggesting that raters applied the analytic rating scales in a consistent manner to assess test-takers’ oral proficiency. This result indicates that the raters were able to judge both linguistic components of performance and Content elaboration/development, even though specific descriptors were not provided in the study. The raters appeared to have general ideas about the nature of ability to elaborate speech, although this ability is deemed fuzzier than command of discrete items of grammar (Brown, 2007). This result suggests that it might be feasible to assess this non-linguistic component in an oral performance test.
The results of the statistical analyses suggest that Content elaboration/development was not completely independent of the linguistic components measured in the present study. Many-facet Rasch measurement statistically demonstrated the unidimensionality of the test scores, indicating that the multiple components assessed in the test demonstrated a single underlying scoring pattern; that is, a person who achieved high English oral proficiency obtained a high score for Content elaboration/development, and vice versa. Although unidimensionality confirmed by Rasch measurement does not automatically mean that the test is measuring a psychologically unidimensional construct (McNamara, 1996), this statistical result suggests that Content elaboration/development, as well as linguistic components, can be legitimately included in the assessment of oral proficiency. The high composite phi coefficient obtained from the multivariate G-theory analysis corroborates this claim. Furthermore, it is important to note that Content elaboration/development made a major contribution to Overall communicative effectiveness and the composite score of the criteria. As noted earlier, speech content (including sophistication, meaningfulness, and relevance of ideas) is likely to be perceived as one of the crucial components of oral performance on some EAP tests because intellectual maturity pertains to success in higher education (Brown, 2007; Brown et al., 2005; Eckes, 2009). A similar result was obtained from the present study, which investigated rater judgments on a general English proficiency test. Even though the test’s purpose was to measure general oral proficiency levels, speech content was believed to be not only a means of demonstrating linguistic competence but also an important criterion underlying evaluation of overall performance quality.
Fluency was perceived to be the most salient among the linguistic criteria. Its effective weight was the largest, and it was the second-strongest predictor of intuitive judgments. Moreover, the largest number of raters mentioned this criterion in their written comments as a feature influencing their intuitive judgments. These results are consistent with the findings of prior studies that have found fluency to be one of the decisive factors in raters’ holistic judgments (Barnwell, 1989; Iwashita et al., 2008; Iwashita & Grove, 2003). Conversely, it was revealed that the other linguistic criteria – Grammatical accuracy, Vocabulary range, and Pronunciation – contributed less to raters’ intuitive judgments, although they were largely correlated with the holistic scores. A plausible reason for the large contribution of Fluency could be that the test-takers in this study had a relatively high level of English proficiency and did not have problems with grammar and pronunciation serious enough to impede message conveyance. This interpretation is partially supported by the relatively low logit values for Grammatical accuracy and Pronunciation, indicating that raters’ evaluation of these criteria was relatively high. In particular, the test-takers might have already achieved intelligible pronunciation or the ‘first level hurdle’ for test-takers (Iwashita et al., 2008, p. 44), which led the raters’ attention to other features. This is consistent with Pollitt and Murray’s (1996) argument that assessing the pronunciation of high level test-takers is pointless because problems with pronunciation become less serious in speakers with higher proficiency. Pollitt and Murray demonstrated that raters’ primary focus in evaluating the performance of high-proficiency test-takers was on stylistic devices including content, elaboration, and creativity, rather than on linguistic features such as vocabulary, grammar, or pronunciation. Similarly, de Jong and van Ginkel (1992) show that the relative weight of fluency becomes higher and that of pronunciation becomes lower as ratings of global oral proficiency rise. The present study, in keeping with the research cited above, indicates that the raters might have placed more emphasis on fluent speech than on linguistic resources and pronunciation when rating high-proficiency test-takers’ overall speaking performance.
Language teachers’ professional perspectives appear to be a possible explanation for the profound influence of Content elaboration/development and Fluency on intuitive judgments of oral proficiency. It can be considered that currently prevailing pedagogical approaches might have influenced the raters’ (i.e. English teachers’) judgments, since their perspectives aligned with the primary concern of communicative language teaching: mastering fluency, rather than formally accurate use of language. Moreover, attaining message-conveyance ability is the principal objective of content-based and task-based language teaching. It is conceivable that teacher perspectives on good oral performance are gained from their professional training, where these teaching approaches are taught and trained (Berry, 2007; Chalhoub-Deville, 1996). Furthermore, as Iwashita and Grove (2003) argue, the practice of communicative language teaching and assessment may result in a belief that grammatical and lexical accuracy are not the most crucial aspects of oral performance. The raters participating in this research are likely to have received instruction in appropriate pedagogical approaches as part of their teacher training and to have applied them in class in teaching and/or assessing their students. Consequently, rater judgments of student performance in the present study might have been influenced by communicative models of language teaching more than by grammar-centered approaches.
The Rasch analysis revealed that Overall communicative effectiveness was not an overfitting criterion. This result is counter-intuitive, because this criterion was expected to be predominantly composed of the criteria used in many speaking tests. The five analytic criteria used in the current study did not fully account for raters’ holistic judgments of overall oral proficiency, given that only 67% of Overall communicative effectiveness was predicted by them. The raters’ written comments also demonstrated that their holistic judgments were influenced not only by the five criteria, but also by components that were not included in the analytic scales. It is not uncommon that even trained and accredited raters heed non-criterion features while rating (Brown, 2007). For instance, Orr (2002) identified 12 non-criterion categories heeded by official raters of the Cambridge First Certificate in English Speaking test, although these raters were only required to assess oral performance using analytic rating scales composed of four criteria. Likewise, May (2006) found that trained raters employed non-criterion features; more than 30% of rater comments studied by May were irrelevant to the analytic criteria given to them for assessment. Similar to the findings of the research above, Overall communicative effectiveness in the present study appeared to consist of the criteria included in the analytic scales as well as features irrelevant to the analytic criteria, the latter of which accounted for approximately 33% of rater judgments of oral proficiency.
Conclusion
The primary objective of this study was to explore the relative contribution of linguistic criteria and speech content to raters’ holistic judgments of general English oral proficiency. A speaking test was administered to Japanese university students in order to explore nine raters’ perception of the relative importance of Grammatical accuracy, Fluency, Vocabulary range, Pronunciation, and Content elaboration/development. A key finding from this research is that the scores for Content elaboration/development most strongly predicted the raters’ intuitive judgments of oral proficiency and made a substantive contribution to the composite score. Fluency was found to be another influential criterion, given its large effective weight and the number of rater comments mentioning it. Raters also appeared to heed non-criterion features, which were deemed to be part of their intuitive judgments of Overall communicative effectiveness.
The results of the present study enrich our understanding of general oral proficiency in second languages and of the construct definition of proficiency tests. The present study has identified an additional dimension that is highly relevant to language proficiency but is not fully delineated by existing communicative language ability models. An ability to elaborate speech content is found to form a major part of the non-linguistic components affecting rater judgments of the quality of second language oral performance. Additionally, this study has explored the relative weights of criteria that are not explicated by the theoretical models. It can be argued that rater judgment of fluency outweighs that of grammatical accuracy, vocabulary range, or pronunciation. The findings of this study are readily applicable to the process of developing construct definitions and scales for oral proficiency tests. As stated earlier, research on raters’ perceptions of general oral proficiency is critical for scale development because it is necessary to connect test criteria with those actually used by English speakers to assess performance on TLU tasks (Bachman & Palmer, 2010). This study suggests that the quality of the ideas that test-takers attempt to convey should be treated as a criterion in oral assessments involving monologic performance, as is the case in most writing performance tests. Narrowly restricting our focus to linguistic features may lead to erroneous inferences about L2 learners’ ability to communicate effectively.
One possible limitation of this research is that the implications stated above may only be applicable to (a) high-proficiency test-takers, (b) monologic elicitation tasks, or (c) English-teacher raters. Raters are likely to heed speech content and fluency to a greater extent when test-takers do not have serious problems with linguistic resources. If raters are distracted by grammatical and phonological errors that hinder their comprehension, these problematic features are likely to contribute to negative rater judgments (Brown, 2007; Pollitt & Murray, 1996). Raters’ perception of low- or intermediate-proficiency students’ oral performances needs to be investigated in future research. Further, it is to be remembered that this research employed only one type of monologic elicitation task, in which the test-takers were required to express their thoughts on given prompts. Rater judgment may differ when dialogic tasks, such as interviews or group discussions, are used to elicit performance. Research has found that pivotal roles are played by supportive non-verbal behaviors on the part of test-takers, listening comprehension, and conversational management in face-to-face interaction (e.g. Ducasse, 2009; Ducasse & Brown, 2009; May, 2009; Orr, 2002). In addition, it is relevant that in the current study, the raters had English-teaching experience. As discussed earlier, their judgments might have been affected by their practice of communicative language teaching or assessment. Raters with different professional backgrounds might heed different components of oral performance from those heeded by English teachers (Chalhoub-Deville, 1996). Without confirming the nature of non-teachers’ perceptions, Pollitt and Murray (1996) argue, ‘we will not know that ‘proficiency’ as judged by teachers is the same thing as the world knows by that name, and the validity of the test will be compromised accordingly’ (p. 89). In summary, further research is recommended to investigate non-teacher or linguistically naïve perceptions of oral performances on various TLU tasks as completed by low- or intermediate-proficiency test-takers.
A final point remains to be made. Multiple statistical analyses of test scores have manifested the relative contribution of four commonly used linguistic criteria and the elaboration of speech content, respectively, to raters’ intuitive judgments. Nevertheless, since non-criterion features may greatly affect rater judgments, further research is also needed to explore the components of Overall communicative effectiveness that cannot be accounted for by the five criteria incorporated into the study.
Footnotes
Appendix
The test-taker profile (n = 30)
| Gender | Male: 16 Female: 14 |
| Academic Year | First Year: 19 Second Year: 11 |
| English Course Level | Advanced: 15 Intermediate: 15 |
| EIKEN Grade | Grade Pre-1: 1 Grade 2: 11 Grade Pre-2: 2 Grade 3: 6 Grade 4: 1 NA: 9 |
Acknowledgements
This research article is derived from my master’s thesis submitted to Sophia University. I would like to thank the participants in the study as test-takers and raters. I would like to express my gratitude to my thesis committee members, Yoshinori Watanabe, Kensaku Yoshida, Junichi Kasajima, and Mitsuyo Sakamoto for their constant guidance. My thanks also go to Tim McNamara and the three anonymous Language Testing reviewers for their constructive comments on earlier versions of the paper.
