
Editorial
Select search scope: search across all journals or within the current journal

The overarching aim of the study is to explore the extent to which test takers’ performances on monologic speaking tasks provide information about their interactional competence. This is an important concern from a test use perspective, as stakeholders tend to consider test scores as providing comprehensive information about all aspects of L2 competence. One hundred and fifty test takers completed a TOEFL iBT speaking section consisting of six monologic tasks, measuring speaking proficiency, followed by a test of interactional competence with three monologues and three dialogues, measuring pragmalinguistic skills, the ability to recipient design extended discourse, and interactional management skills. Quantitative analyses showed a medium to high correlation between TOEFL iBT speaking scores and interactional scores of
Over the past decade, testing and assessing spoken-language interpreting has garnered an increasing amount of attention from stakeholders in interpreter education, professional certification, and interpreting research. This is because in these fields assessment results provide a critical evidential basis for high-stakes decisions, such as the selection of prospective students, the certification of interpreters, and the confirmation/refutation of research hypotheses. However, few reviews exist providing a comprehensive mapping of relevant practice and research. The present article therefore aims to offer a state-of-the-art review, summarizing the existing literature and discovering potential lacunae. In particular, the article first provides an overview of interpreting ability/competence and relevant research, followed by main testing and assessment practice (e.g., assessment tasks, assessment criteria, scoring methods, specificities of scoring operationalization), with a focus on operational diversity and psychometric properties. Second, the review describes a limited yet steadily growing body of empirical research that examines rater-mediated interpreting assessment, and casts light on automatic assessment as an emerging research topic. Third, the review discusses epistemological, psychometric, and practical challenges facing interpreting testers. Finally, it identifies future directions that could address the challenges arising from fast-changing pedagogical, educational, and professional landscapes.
The aim of this study was to investigate how test methods affect listening test takers’ performance and cognitive load. Test methods were defined and operationalized as while-listening performance (WLP) and post-listening performance (PLP) formats. To achieve the goal of the study, we examined test takers’ (
In this study, we present the development of individualized feedback for a large-scale listening assessment by combining standard setting and cognitive diagnostic assessment (CDA) approaches. We used the performance data from 3,358 students’ item-level responses to a field test of a national EFL test primarily intended for tertiary-level EFL learners. The results showed that proficiency classifications and subskill mastery classifications were generally of acceptable reliability, and the two kinds of classifications were in alignment with each other at individual and group levels. The outcome of the study is a set of descriptors that describe each test taker’s ability to understand certain level of oral texts and his or her cognitive performance. The current study, by illustrating the feasibility of combining standard setting and CDA approaches to produce individualized feedback, contributes to the enhancement of score reporting and addresses the long-standing criticism that large-scale language assessments fail to provide individualized feedback to link assessment with instruction.
This paper investigates what matters to medical domain experts when setting standards on a language for specific purposes (LSP) English proficiency test: the Occupational English Test’s (OET) writing sub-test. The study explores what standard-setting participants value when making performance judgements about test candidates’ writing responses, and the extent to which their decisions are language-based and align with the OET writing sub-test criteria. Qualitative data is a relatively under-utilized component of standard setting and this type of commentary was garnered to gain a better understanding of the basis for performance decisions. Eighteen doctors were recruited for standard-setting workshops. To gain further insight, verbal reports in the form of a think-aloud protocol (TAP) were employed with five of the 18 participants. The doctors’ comments were thematically coded and the analysis showed that participants’ standard-setting judgements often aligned with the OET writing sub-test criteria. An overarching theme, ‘Audience Recognition,’ was also identified as valuable to participants. A minority of decisions were swayed by features outside the OET’s communicative construct (e.g., clinical competency). Yet, overall, findings indicated that domain experts were undeniably focused on textual features associated with what the test is designed to assess and their views were vitally important in the standard-setting process.
This study evaluated the validity of the Michigan English Test (MET) Listening Section by investigating its underlying factor structure and the replicability of its factor structure across multiple test forms. Data from 3255 test takers across four forms of the MET Listening Section were used. To investigate the factor structure, each form was fitted with four Bayesian confirmatory factor analysis (CFA) models: (1) a three correlated-factor model, (2) a bi-factor model, (3) a higher-order factor model, and (4) a single general-factor model. In addition, a four-pronged heuristic comprising construct delineation, construct operationalization, factor structure analysis, and congruence coefficient was developed to examine the replicability of factor structures across the test forms. Results from the CFA models showed that the test forms were unidimensional and the four-pronged heuristic indicated that the test construct was consistently operationalized across forms. Furthermore, the congruence coefficient indicated that the factor structure representing listening was highly similar and replicable across test forms. In sum, the construct of the MET Listening Section did not comprise divisible subskills. Yet, the unidimensional factor structure of the test was replicable across the test forms.
