Abstract
Over the past decade, testing and assessing spoken-language interpreting has garnered an increasing amount of attention from stakeholders in interpreter education, professional certification, and interpreting research. This is because in these fields assessment results provide a critical evidential basis for high-stakes decisions, such as the selection of prospective students, the certification of interpreters, and the confirmation/refutation of research hypotheses. However, few reviews exist providing a comprehensive mapping of relevant practice and research. The present article therefore aims to offer a state-of-the-art review, summarizing the existing literature and discovering potential lacunae. In particular, the article first provides an overview of interpreting ability/competence and relevant research, followed by main testing and assessment practice (e.g., assessment tasks, assessment criteria, scoring methods, specificities of scoring operationalization), with a focus on operational diversity and psychometric properties. Second, the review describes a limited yet steadily growing body of empirical research that examines rater-mediated interpreting assessment, and casts light on automatic assessment as an emerging research topic. Third, the review discusses epistemological, psychometric, and practical challenges facing interpreting testers. Finally, it identifies future directions that could address the challenges arising from fast-changing pedagogical, educational, and professional landscapes.
In the field of language testing, the assessment of oral communication (e.g., L2 speaking) represents a vibrant line of research, spawning much scholarly discussion and debate over the past three decades (for a historical review, see Fulcher, 2015). However, one area of oral communication—spoken-language interpreting—seems to have drawn far less attention from language testers than it rightfully deserves, given that interpreting, in and of itself, is language-mediated communication. A review of the two prominent language testing journals, Language Testing (1984–2020) and Language Assessment Quarterly (2004–2020), produced only four publications that center on spoken-language interpreting (i.e., Han, 2016, 2019; Stansfield & Hewitt, 2005; Zhao & Gu, 2016) or what is known as interpreting testing and assessment (hereafter ITA). The lack of testers’ attention to (oral) interpreting stands in sharp contrast to its counterpart of written translation, which has long been practiced as a pedagogical exercise in L2 teaching (e.g., grammar translation method), studied by early language testing researchers (e.g., Lado, 1961), and which has recently reinvigorated interest among L2 educators and applied linguists (e.g., Cook, 2010). Although the historical interplay between translation and language learning may explain the visibility of translation among L2 teachers and/or testers, it would seem that interpreting has seldom been explicitly associated with L2 learning (e.g., as an exercise to promote L2 acquisition), and has hardly been a concern of L2 testers. 1
The enterprise of ITA (i.e., interpreting testing and assessment) and its developmental trajectory are shaped by forces originating from the interpreting profession. The fundamental force driving ITA is the societal and political demand for high-quality interpreting services, in response to the global mobility of goods, services, technologies, ideas, and people. For example, recent waves of immigration (e.g., skilled labor, refugees, asylum seekers), which have given rise to social challenges of enabling equitable access to legal, medical and other public services in many countries (e.g., Australia, Canada, the UK), have precipitated the shortage of qualified community interpreters (see, e.g., González et al., 2012, p. 7). The trend of globalization, which has boosted business, cultural, and people-to-people exchanges, has also scaled up the need for interpreter-mediated communication. Furthermore, the provision of interpreting services has been institutionalized in bi/multilateral diplomatic situations and international organizations, notably the United Nations and the European Union.
Consequently, the surging demand for better interpreting services puts growing pressure on interpreter training and professional certification. In both these areas, ITA can play a gatekeeping role, preventing candidates without sufficient competence from entering the professional market (Setton & Dawrant, 2016). Another area that relies heavily on ITA is empirical interpreting research, in which interpreting assessment outcomes are operationalized as (in)dependent variables for research purposes (Han, 2018b). In all three areas (i.e., training, certification, and research), measurements from ITA constitute a critical evidential basis, on which inferences and actions are grounded for such decision-making as selection, certification, and hypothesis testing.
As is to be elaborated below, the fast developments over the past decade(s) regarding the establishment of postgraduate interpreting programs, of interpreter certification testing schemes, and of translation and interpreting studies as an academic discipline, have become the three-pronged driving force that fuels the practice of and research on ITA. Such developments are also accompanied by an increasing number of ITA-themed publications. For instance, there have been special issues published in mainstream translation and interpreting journals (Koby & Lacruz, 2017; Melby, 2013), monographs (Sawyer, 2004; Setton & Dawrant, 2016), edited books (Angelelli & Jacobson, 2009; Chen & Han, 2021; Huertas-Barros et al., 2018; Tsagari & van Deemter, 2013), doctoral theses (Clifford, 2004; Collados Aís, 1998; Han, 2015a; Wu, 2010), and research reports (e.g., ALTA Language Services, 2007; Hale et al., 2012; Roat, 2006).
Despite the increasing availability of relevant literature, and in contrast to our in-depth knowledge of speaking assessment (Fan & Xu, 2020; Fulcher, 2015; Luoma, 2004) and translation quality assessment (Han, 2020; House, 2015), it would appear that little effort has been initiated to provide a comprehensive, state-of-the-art review of ITA practice and research. In light of this, in the present article we aim to take stock of practice and research associated with testing and assessing spoken-language interpreting over the past decade or so, covering three major areas where ITA plays a critical role: interpreter training and education, professional certification, and interpreting research. In particular, we try to synthesize the state of affairs, identify emerging trends, and discover potential lacunae and challenges. By doing so, we want to provide an updated understanding of ITA as an emergent endeavor and a worthwhile research area. Ultimately, we hope to increase the visibility of ITA among applied linguists, particularly language testers, and to seek cross-disciplinary collaboration with colleagues from cognate disciplines.
The review comprises eight sections: in section one, we contextualize the review by taking account of societal factors driving ITA and its recent developments; in section two, we summarize three main approaches to defining the construct of interpreting ability/competence, including the cognitive processing, multi-componential, and interactional approaches, and relevant research on interpreting ability/competence; in section three, we provide a descriptive overview of ITA practiced in educational, professional and research contexts; in section four, we segue into a synthetic analysis of important aspects of ITA practices (e.g., assessment criteria, scoring methods); in section five, we focus on main topics, findings and trends in evidence-based ITA research; in section six, we discuss major challenges facing interpreting testers, by highlighting lacunae derived from the above review; in section seven, we identify future directions that could help to address these challenges; and in the last section, we conclude the review by recapitulating the main features of ITA, discussing the potential limitations of the review, and calling for sustained attention to ITA from all relevant stakeholders.
Interpreting ability/competence: Construct definition
In the interpreting literature, there has been much scholarly discussion on interpreting ability/competence. One popular method of defining interpreting ability can be called the cognitive processing approach (e.g., Gerver, 1975; Gile, 2009; Moser, 1978; Setton, 1999), in which researchers try to model essential linguistic-cognitive operations and processes involved in interpreting. One of the early models was proposed by Gerver (1975), who postulated mental structures and processes in simultaneous interpreting. The most comprehensive and detailed model was produced by Setton (1999), who described every possible step involved in the comprehension of source-language input, cross-linguistic conversion, and articulation of target-language output. Furthermore, the Effort Models by Gile (2009) are probably the most widely known models of interpreting processes, in which four major efforts of listening and analysis, memorization, production, and coordination were hypothesized.
Another useful method of conceptualizing interpreting competence is the multi-componential approach (e.g., Pöchhacker, 2000; Setton & Dawrant, 2016; Wang, 2007), in which researchers enumerate core traits and characteristics that must be possessed by interpreters. There has been a general consensus on two fundamental and indispensable components of interpreting: bilingual language competence and linguistic transfer competence. In addition, two frequently mentioned components are subject matter or topical knowledge and professionalism, particularly ethical competence. Several other components have also been discussed, including interaction management skills and physio-psychological qualities (e.g., stamina, motivation).
A recent method of defining the construct for the purpose of testing and assessment is known as the interactional approach (Han, 2015a; Wang et al., 2020), in which interpreting performance is viewed as a sign of underlying traits and is influenced by the context in which it occurs, and is thus considered a sample of performance in similar contexts (see also Bachman, 2007; Chapelle, 1999). In particular, the interactionalist approach accentuates the role of (meta-)cognitive processes/abilities that regulate linguistic-cognitive processes in response to external task-specific characteristics. Consequently, this approach links the cognitive process modeling and the multi-componential modeling of interpreting ability/competence; establishes a relationship between interpreters’ internal traits and external task/context-specific characteristics; and relates performance consistency to interpreters’ ability and task features.
Apart from the theoretical definitions described above, empirical research has also been conducted covering such topics as development and progression of interpreting ability/competence, acquisition and patterns of interpreting strategies, and cognitive processes involved in interpreting. First, some researchers have examined interpreting ability/competence from the expert-performance perspective (Ericsson, 2000; Moser-Mercer, 2008), investigated its development (Albl-Mikasa, 2013; Cai et al., 2015), and designed questionnaire instruments to measure interpreting competence (Schaeffer et al., 2020). One interesting study by Albl-Mikasa (2013) analyzed semi-structured in-depth interviews with professional interpreters to identify critical competences, their acquisition over time, and enhancement strategies. Second, another group of researchers have focused on interpreters’ use of strategies, reporting that strategy use may ease cognitive burden and improve delivery (e.g., Bartłomiejczyk, 2006; Gile, 2009; Li, 2013). Third, an increasing amount of scholarly attention has been directed to the cognitive processing involved in interpreting. Two specific topics have been thoroughly investigated: working/short-term memory and cognitive load. Regarding the role of memory in interpreting, a meta-analytic study by Wen and Dong (2019) showed an interpreter advantage in both working memory and short-term memory spans, compared to bilingual controls. In addition, Mellinger and Hanson’s (2019) meta-analysis indicated that working memory capacity was positively correlated with measures of interpreting quality. In terms of cognitive load in interpreting, recent empirical studies have shed light on how interpreters process source-language input and cope with problem triggers. For instance, based on the corpus analysis of filled pauses (e.g., uh, um) as a window into interpreters’ cognitive load, Plevoets and Defrancq (2018) found that pause frequency increased with the lexical density of the source-language texts, but was negatively associated with formulaicity of both the source-language and the target-language texts. In addition, such techniques as digital pen recording, eye-tracking and event-related potential have been applied to understand interpreters’ cognitive load (e.g., Chen, 2020; Koshkin et al., 2018; Seeber 2011; Tiselius & Sneed, 2020). For instance, a recent eye-tracking investigation by Tiselius and Sneed (2020) into dialogue interpreting, an under-examined mode of interpreting, revealed no significant difference regarding gaze patterns between experienced and inexperienced interpreters, but indicated that interpreting into L2 in a dialogue may incur more cognitive effort than interpreting into L1. Similarly, in a pen-recording and eye-tracking study by Chen (2020) to examine cognitive processes during consecutive interpreting (with note-taking), it was found that L2-to-L1 interpreting appeared to be less cognitively demanding than the other direction, although a higher level of cognitive load was involved in note-taking in L2-to-L1 direction than the other way around.
Interpreting testing and assessment: An overview
Given the above approaches to construct definition, we now proceed to review how ITA has been conducted in the field. In general, the history of ITA is severely under-documented; at best, we are only able to piece together scattered information in the literature concerning the history of the conference interpreting profession, mostly contributed by Jesús Baigorri-Jalón (2005, 2014), to partially reconstruct the development of ITA. Arguably, the practice of ITA in the contemporary era could be traced back to the Paris Peace Conference of 1919 when the interpreting profession was born. According to Baigorri-Jalón (2005, p. 990), “selection tests for staff interpreters and translators in the League of Nations began in December 1919”. One of the key milestones in the history of ITA was the founding of the International Association of Conference Interpreters (AIIC) in 1953, following the Nuremberg trials after the World War II, to ensure the highest level of interpreting quality and ethics at international conferences. Another important milestone that relates directly to ITA was the establishment of the National Accreditation Authority for Translators and Interpreters (NAATI) in 1977 in Australia. This was the first country in the world to have a government-instituted accreditation authority that has since played a critical role in testing and certifying interpreters in the country. Meanwhile, the Court Interpreters Act of 1978 in the USA led to the creation and implementation of the Federal Court Interpreter Certification Examination (FCICE) in 1980, which was the first certification test of its kind and has become an exemplar for many other certification tests (see Setton & Dawrant, 2016). 2 It was only in the 1990s that the practice of ITA began to be examined and documented (e.g., Arjona-Tseng, 1993); only in the recent decade or so, has systematic and empirical research been initiated by a small group of researchers (e.g., Hale et al. 2012; Liu, 2013; Wu, 2010). In this section, we aim to offer a general review of ITA as practiced in three areas: interpreter training and education, professional certification, and interpreting research. We first describe the practice of ITA in the educational context, focusing on formative and summative assessment. We then shed light on professional certification in which ITA features as a major component. Lastly, we provide an overview of the use of ITA in empirical interpreting studies to generate measurements for research purposes.
Interpreter training and education
The past decade or so has witnessed a tremendous growth in the number of postgraduate-level interpreting programs in different parts of the world, especially in emerging markets such as mainland China. This educational development has subsequently led to a huge demand for educational assessment that serves formative and summative purposes.
On the one hand, formative assessment, often conceptualized by interpreter educators as an on-going and progressive process, is generally of low stakes and is aimed at revealing students’ strengths and weaknesses, generating formative feedback, and promoting self-directed learning and metacognitive awareness (Arumí Ribas, 2010). In particular, four main forms of formative assessment have been documented: (1) self-assessment (Han & Riazi, 2018; Postigo Pinazo, 2008), (2) peer assessment (Lee, 2017; Su, 2019b), (3) a combination of self and peer assessment (Fowler, 2007; Hartley et al., 2003), and (4) student portfolios (Arumí Ribas, 2010; Sawyer, 2004).
On the other hand, summative assessment, conducted at the end of a given training period to ascertain the level of student achievement, could be of high stakes for students and/or instructors (Liu et al., 2008; Setton & Dawrant, 2016). This is because relevant scores can be used to select better-performing students for advanced-level training (Sawyer, 2004), to inform the decision of degree conferral at the end of an interpreting program (Lee, 2008; Liu et al., 2008), and to evaluate the efficacy of course/curriculum design (Hale & Ozolins, 2014). There are three notable sources that provide valuable information on summative assessment: (1) Liu et al. (2008), who surveyed assessment practices in 11 postgraduate-level interpreting programs from Taiwan, mainland China, Britain, and the USA; (2) Sawyer (2004), who provided an in-depth case analysis of summative assessment in the Middlebury Institute of International Studies at Monterey in the USA; and (3) Setton and Dawrant (2016), who drew on their extensive teaching, research, and evaluation experience to describe and problematize high-stakes summative assessment conducted in conference interpreting programs.
Professional certification
Another area where testing and assessment plays an important role is professional certification, as interpreter certification performance testing (ICPT) is commonly recognized as a major avenue to certification. ICPT has also experienced considerable development over the past decade, as is demonstrated by the establishment of new certification testing programs, the expansion of current testing programs (e.g., an increasing number of test candidates), and the enhancement and validation of certification testing procedures (see Han & Slatyer, 2016). Overall, ICPT is of high-stakes, not only because it can impact interpreters’ livelihood by regulating their access to professional practice, but also because it guards against potential risks to users, especially in medical and legal settings.
Although ICPT may target diverse domains of practice in a range of socio-cultural settings, the nature of ICPT is largely dictated by practical needs in a given country. For example, in countries with large-scale immigration such as Australia, Canada, New Zealand, the UK, and the USA, ICPT tends to concentrate on the certification of community interpreters who mainly work in public service facilities (e.g., hospitals, courts, police stations, schools, post offices). In contrast, in other countries such as China, the main purpose of ICPT is to certify interpreters who can work in the private sector (e.g., international conferences) and government-related workplaces (e.g., government departments). For detailed reviews, the special issue on interpreter certification is a good source of relevant information (see Melby, 2013).
For ICPT, scores derived from a given test are often used as an indicator of what test candidates know and can do in a certain domain of practice. That is to say, test scores are used to describe test candidates’ mastery of relevant knowledge, skills, and abilities, and their interpreting performance (e.g., accuracy, delivery, language quality) in practice domains of interest. Such score-based inferences also form a critical basis for certification decisions. In particular, to be certified as an interpreter, test candidates must outscore a predetermined cut-off point. However, for many certification programs, obtaining a high score on ICPT does not necessarily lead to certification. Test candidates also need to pass written tests on language competence and/or professional ethics (see, e.g., Australia’s National Accreditation Authority for Translators and Interpreters/NAATI testing, the China Accreditation Tests for Translators and Interpreters/CATTI, the USA’s Federal Court Interpreter Certification Examination/FCICE).
Interpreting research
In evidence-based interpreting research, testing and assessment represents an important means to collect quantitative data, although this role receives scant attention (Han, 2018b). Specifically, there are at least four ways in which the measurements from ITA can influence research. First, ITA-based measurements can function as a grouping variable to categorize test candidates into different ability/performance groups, so that between-groups comparisons could be made on the basis of other variables of researchers’ interests. Second, scores obtained from ITA can form a dependent variable which is in turn to be predicted by other variables based on statistical modeling such as multiple regression analysis (e.g., Lee, 2015). Third, ITA results can also be operationalized as a dependent variable to be correlated with other variables so as to understand inter-variable relationships (e.g., Yu & van Heuven, 2017). Fourth, ITA outcomes, functioning as a dependent variable, can be statistically compared across two or more cohorts of interpreters (e.g., Hale & Ozolins, 2014).
As can be seen above, the three major areas where ITA has been practiced have their distinct assessment foci and purposes and thus involve different stakes. These discrepancies are also reflected by varying approaches and designs that characterize testing and assessment practices reviewed in the next section.
Testing and assessment practices: A synthetic analysis
This section attempts to synthesize testing and assessment practices and also provide some critical comments where appropriate and relevant. In particular, it summarizes operational assessment procedures by shedding light on a diverse range of topics regarding ITA practices including three major aspects of ITA: (1) specificities of interpreting assessment (i.e., modes of interpreting, language pairs and directionality), (2) assessment design (i.e., assessment tasks, assessment criteria), and (3) scoring and rater training (i.e., scoring methods, specificities of scoring operationalization, rater recruitment and selection, rater training and calibration).
Specificities of interpreting assessment
Interpreting modes
In ITA, three major modes of interpreting are frequently tested and assessed, including simultaneous interpreting (SI), consecutive interpreting (CI), and sight interpreting (SiI). When an interpreter performs SI, he or she listens to a source-language speech and interprets simultaneously in the target language (this usually happens in a sound-proof booth and with the help of modern technology). When it comes to CI, an interpreter usually listens to a speaker for a certain amount of time while taking notes and then interprets what the speaker has said when he or she stops (which is also known as long consecutive or classic consecutive). Dialogue interpreting (DI) can be considered a special type of CI in which an interpreter usually interprets dialogue-like interactions rather than speeches. Lastly, SiI involves an interpreter’s reading of a text from a source language into a target language simultaneously.
Depending on different purposes, some modes of interpreting may figure more prominently in testing and assessment than others. For instance, in most interpreting programs—given that students learn DI at early stages of training before proceeding to CI and SiI at intermediate stages and finally to SI at advanced stages—educational testing and assessment tends to reflect this sequence of learning by incorporating different modes of interpreting at different learning stages (see Setton & Dawrant, 2016, “incremental realism”). In professional certification, modes of interpreting included in ICPT need to reflect professional practice in a target domain. For example, although a certification test designed for public services settings usually incorporates CI, SI and SiI (e.g., the U.K. Diploma in Public Service Interpreting test), other tests that target business settings may only sample CI (e.g., the Shanghai Business Interpretation Accreditation Test).
Language pairs and directionality
The issue of language pairs mostly concerns professional certification testing. Although some certification tests focus solely on a single language pair, others may involve multiple language pairs.
Another related issue is interpreting directionality, or the direction in which a test candidate needs to interpret. In training, certification, and professional practice, an interpreter is only allowed to interpret between her or his native language (also known as an A language) and another working language(s) (as a B language); extremely rare is the situation where an interpreter works between two B languages. Consequently, for a given language pair, an interpreter could interpret into her or his A language (i.e., B-to-A) or into B language (i.e., A-to-B). Historically, interpreting into B is frowned on by some early trainers and educators, on the grounds that such practice may potentially undermine interpreting quality (Seleskovitch & Lederer, 1989). However, owing to practical constraints (e.g., unavailability of interpreters for certain language pairs), interpreting is fully bidirectional in many markets. This reality is reflected in interpreter educational assessment (see Liu et al., 2008; Setton & Dawrant, 2016) and certification testing (Han, 2016), as interpreting into and from A language(s) is customarily tested.
Assessment design
Assessment tasks
Regarding the development of assessment tasks, interpreting testers need to consider a number of important issues concerning the characteristics of tasks, including, for example, the difficulty, variety, number, and length of tasks. To control for task difficulty, testers usually rely on a loose set of variables (also known as “difficulty factors” or “problem triggers”), such as the delivery speed, accent, terminology, information density, and availability of background materials (Setton & Dawrant, 2016; Wu, 2019). In interpreter training, assessment tasks are often developed in line with incremental realism (Setton & Dawrant, 2016), meaning that the difficulty of source-language speeches increases with the time of training and that certain problem triggers are deliberately not operationalized in source-language materials until a later learning phase. Regarding ICPT, source-language speeches usually contain a full range of difficulty factors that could be encountered in a given real-life practice domain. Regarding assessment for research purposes, an array of delivery features of source-language speeches (e.g., fast speech rate, strong accent, convoluted syntactical structures) have been investigated to understand their impact on interpreters’ performance (see, e.g., Han & Riazi, 2017; Meuleman & Van Besien, 2009). Currently, it seems that what interpreting testers really need is a descriptive framework of task characteristics that can be used to develop task specifications and ensure task comparability (see Campbell & Hale, 2003; Setton & Dawrant, 2016).
In terms of task variety, number and length, Liu et al.’s (2008) survey of 11 interpreting programs showed that summative assessment carried out in training institutions included one task for each direction, and that the duration of source speeches ranged from 45 seconds to 12 minutes. However, it would seem that none of the training institutions provided any justification for their decisions in their assessment design, which was deemed inadequate by Setton and Dawrant (2016), who argued in favor of a more stringent standard for summative assessment. They envisaged that the Professional Examination in Conference Interpreting (PECI) would incorporate three major tasks of about 40 minutes in total, based on a series of mini-tasks involving a variety of interpreting modes. When it comes to ICPT, it is commonly agreed that assessment tasks must represent and capture the essential features of professional interpreting practice in a target domain (Campbell & Hale, 2003; Hale et al., 2012). For example, California’s Court Interpreter Certification and Registration Testing (CCICRT) conducts systematic analyses of real-life practice involving the collection of both quantitative and qualitative data from multiple stakeholders (e.g., practicing interpreters, subject matter experts) to inform test/task design and development (ALTA Language Services, 2007). As a result, the CCICRT consists of two SiI tasks (225 words each), one CI task, and one SI task (seven minutes or 800–850 words).
Assessment criteria
There has been a variety of quality criteria used to assess interpreting, many of which have their roots in Bühler’s (1986) survey study (e.g., fluency of delivery, correct terminology, correct grammar, sense consistency with the original). Over the years, researchers seem to agree on three major quality criteria: content, delivery, and language quality (Lee, 2008; Liu, 2013; Setton & Dawrant, 2016). However, specific operationalization of these criteria may differ from one assessment to another (see Lee, 2008; Liu, 2013). One noteworthy criterion in ITA is the informational correspondence/equivalence between what a speaker delivers and what an interpreter renders, also labeled as “information completeness” or “fidelity,” which constitutes a crucial component of the criterion of content. In the rest of the review, we use “fidelity” to refer to the concept of informational equivalence. Certainly, although general criteria such as delivery and language quality could appear in any type of speaking performance assessment (see speaking assessment criteria for IELTS and TOEFL), it is fidelity that represents a unique and dominant quality concern in ITA (Han et al., 2021). As such, fidelity is usually assigned more weight than delivery and language quality, especially for ICPT in specific contexts (e.g., medical or legal settings).
Scoring and rater training
Scoring methods
One common method of evaluating interpreting can be called itemized/atomistic analysis, in which raters concentrate on and analyze interested phenomena in an interpretation such as errors and (para)linguistic features (e.g., ill-grammaticalities, disfluencies), and tally frequencies of their occurrence as an indicator of interpreting quality. The itemized/atomistic method is epitomized by proposition-based analysis, error analysis, and (dis)fluency analysis as different means to assess interpreting. Although this method has the potential to provide an in-depth understanding of interpretation, it is said to be reductionist, labor-intensive, time-consuming and irreproducible. Consequently, such analyses are often conducted for pedagogical and research purposes where a relatively small number of renditions make them feasible.
Another common method relies on paper-based or electronic questionnaires presented in the form of a checklist or an assessment grid in which an inventory of quality criteria is structured hierarchically and aligned to a Likert-type scale (e.g., highest quality = 5, lowest quality = 1). Although ratings assigned to each criterion can be calculated in one way or another as a global index of interpreting quality (Lee, 2015), qualitative comments on each criterion could function as formative feedback (Hartley et al., 2003). This method may result in increased cognitive load for raters, as they need to differentiate among a series of conceptually related, yet individually presented, assessment criteria (e.g., as many as 30 micro-criteria in Hartley et al., 2003).
An additional method incorporating both itemized/atomistic analysis and rating scale-based assessment can be called multi-methods scoring. This scoring approach is practiced by a number of testing agencies in the USA (see CCHI, 2012; National Center for States Courts, 2013; PSI Services LLC, 2013). On the one hand, a list of scoring units (usually source-language words or phrases preselected by test developers) is provided to raters. They then determine whether test candidates have (in)correctly rendered such units. On the other hand, raters also use a (holistic/analytic) Likert-type rating scale to evaluate the overall effectiveness of interpreting. One potential problem, however, is the lack of a theoretical and empirical basis for integrating the two strands of scores for decision-making.
A trending method in recent years has been rubric-referenced, rating scale-based assessment, also known as rubric scoring. Essentially, a graduated series of descriptors, developed to capture features of interpreting performance at different levels, is used by raters to assign a rating(s) to an interpretation. The use of rubric scoring has gained popularity in interpreter educational assessment (Lee, 2008; Setton & Dawrant, 2016), certification testing (Liu, 2013), and research-oriented assessment (Tiselius, 2009). However, to harness the potential of rubric scoring, test developers need to write accurate descriptors, conduct rigorous rater training, and mitigate undesirable rater effects as much as possible (see Han, 2018b), all of which require substantial investment of human and logistical resources.
A less frequently used yet potentially efficient method to evaluate interpreting is comparative judgment. Rooted in psychophysical analysis and applied in educational settings, comparative judgment has been trialed by a few interpreting researchers (see Han, 2021; Wu, 2010) with some encouraging results (e.g., high scale separation reliability). Comparative judgment requires raters to compare two similar objects (e.g., two renditions) in order to choose the one with better quality. Results from repeated comparisons are then modeled statistically to produce a scaled rank order of all renditions compared from the worst to the best. Given its limited publicity in ITA, comparative judgment has rarely been practiced in applied settings.
Finally, we draw attention to an emergent approach to interpreting evaluation: automatic assessment (see Han et al., 2020; Ouyang et al., 2021; Stewart et al., 2018; Yu & van Heuven, 2017; Zhang, 2016). Its key objective is to make evaluation automatic, reliable, and affordable. As will be seen in section five, where we provide a detailed analysis, automatic assessment in ITA is in its infancy, with many of its much-touted benefits being slow to materialize. It is dwarfed by the impressive progress in automated assessment of L2 speaking performance, as several automated scoring systems (e.g., ETS’s SpeechRater, Pearson’s Versant) have already been commercialized.
Specificities of operational scoring
In this section, we highlight four areas of operational scoring that are potentially specific to ITA. The first concerns the visual versus auditory mode in which fidelity of interpreting is assessed. In fidelity assessment, raters need to constantly compare target-language renditions against original information. There are two ways such comparisons can be made: (1) listening to audio-recordings of renditions (i.e., auditory mode) while checking the source text, and (2) reading transcripts of renditions (i.e., visual mode) against the source text. The former appears to be mostly practiced in certification testing (e.g., ALTA Language Services, 2007; CCHI, 2012; National Center for States Courts, 2013), whereas the latter is preferred in interpreting research (e.g., Liu & Chiu, 2009; Tiselius, 2009). Although the latter mode may help raters identify different errors in interpreting (because the transcript is static and can be examined multiple times), the former mode seems to be more in line with spoken-language interpreting being used and evaluated by users and listeners in reality.
The second specificity has to do with the use of source versus target text as the reference material in fidelity assessment (see Han et al., 2021). The source information conveyed by a given speaker can be accessed by raters based on the source-language text (see, e.g., Lee, 2008; Setton & Motta, 2007) or a scripted exemplar (target-language) rendition carefully prepared by expert interpreters (see, e.g., Tiselius, 2009; Yeh & Liu, 2006). The rationale underlying the latter approach is that the exemplar or ideal rendition contains all information intended by the source-language text (Carroll, 1966). Using this approach, raters are in effect comparing the informational content in the same language. Although the target text approach is a common practice in machine translation research (Carroll, 1966), it could trigger controversy in the professional interpreting community because this approach essentially contradicts interpreting as a creative meaning-making activity.
The third specificity is scoring in situ versus scoring post-hoc. In Liu et al.’s (2008) review of summative assessment in 11 interpreter training institutions, students’ performances are mostly assessed in real time and on the spot (see also Setton & Dawrant, 2016). Although scoring in situ enables assessors to observe the whole process of interpreting, it requires a great deal of concentration from assessors (because interpreting is delivered real time) and may also intimidate student interpreters (because of the presence of assessors). In contrast, post-hoc scoring is more common in professional certification and interpreting research (see, e.g., CCHI, 2012; Lee, 2008; Liu, 2013), and it allows assessors to listen to interpreting recordings at their own pace.
The fourth specificity refers to group deliberation versus independent scoring. When raters need to listen to recorded interpretations, they could work as a scoring team, listening to recordings together and making pass/fail decisions based on group deliberation, especially in educational summative assessment; alternatively, they could work at their own pace, listening to recordings individually and providing independent scores. In Liu et al.’s (2008) review, at the Middlebury Institute of International Studies at Monterey, three raters usually formed a group to listen to exam recordings. This practice contrasts with that at the University of Westminster where raters preferred independent listening so that they could play recordings on their own. It is debatable whether one way is superior to the other.
Rater recruitment and selection
In ITA, human raters constitute one of the most important factors to ensure scoring quality. An ideal rater would thus be the one who has the following characteristics (see also Han, 2018b; Setton & Dawrant, 2016): (1) excellent command of the language pairs appropriate to a given assessment, (2) history of study in an interpreting program (preferably at postgraduate level), (3) a professional interpreting background, (4) experience in teaching interpreting, (5) extensive involvement in interpreting assessment, and (6) assessment literacy (i.e., possession of knowledge about the basic principles of sound assessment practice). However, in most cases rater recruitment is often based on a compromise among the above desired characteristics.
For summative interpreting assessment (e.g., PECI), in Liu et al.’s (2008) review, most of the interpreting programs surveyed had internal examiners (e.g., staff interpreting teachers) and/or external examiners (e.g., renowned professional interpreters) to sit on the peer review panel. In addition, seven out of 11 programs customarily recruited at least three raters, with some even involving as many as five to six raters. In certification testing, recruited raters usually satisfy a number of the requirements listed above (see ALTA Language Services, 2007; Wu et al., 2013), and two to three raters are often used to assess each interpretation (Campbell & Hale, 2003; Roat, 2006). In interpreting research, both expert and novice raters are used. The novice raters may be student interpreters (Han, 2015b; Lee, 2008), translators and translation teachers (Wu, 2010), or even bilinguals with no experience of interpreting training (Tiselius, 2009).
Rater training and calibration
In parallel to the emphasis on ideal rater characteristics, many educators, certifiers, and researchers consider rater training and calibration as a critically important step to ensure scoring quality (see Hale et al., 2012; Liu et al., 2008). Nevertheless, Liu et al. (2008) found that most of the 11 interpreting institutions they surveyed did not organize rater training of any sort, with the reason being that teachers/raters were said to be already familiar with in-house assessment criteria and relevant procedures. For certification testing, regular training and calibration is conducted for many ICPT programs (up to two or three days), as can be seen in ALTA Language Services (2007), CCHI (2012) and PSI Services LLC (2013). In interpreting research where rater-mediated assessment is used as a data-generation instrument, rater training is also practiced, although its duration may vary from ten minutes (Tiselius, 2009) to about five hours (Liu, 2013). Comparatively, rater training in L2 speaking assessment usually takes longer and involves lengthier pilot scoring sessions (see Xi & Mollaun, 2011). Lower levels of training effort in ITA, however, should not be interpreted as indicating that there is less complexity involved in interpreting assessment. Indeed, it is argued that assessing the quality of interpreting is a multi-tasking and cognitively demanding activity, requiring both expertise and experience (Han, 2015b; Wu, 2010).
Apart from the concerns over the practice and length of rater training, several other issues also merit attention. One issue is related to the procedures and components involved in rater training. Han (2018b) identified four common training components in the interpreting literature, including (1) an introduction and explanation of relevant content, (2) a practice session in which raters apply scoring guides to assess interpreting, (3) a norming session where raters evaluate benchmarked interpreting samples, and (4) discussion sessions in which raters exchange their views. Another noteworthy issue is training mode: whether it is organized face-to-face or remotely; and whether it is provided collectively to a group of raters or individually to a single rater. Rater training could be conducted via email/videophone or to a single rater (see, e.g., Lee, 2008; Wang et al., 2015), although most training sessions tend to be provided to a rater group on site (e.g., Roat, 2006; Tiselius, 2009). Furthermore, a related issue is the theoretical rationale underlying rater training, which could be based largely on a behaviorist approach (e.g., striving for correct error detection), a cognitive approach (e.g., achieving similar cognitive processes in scoring), or both. Finally, the effectiveness and assumed benefits of rater training (e.g., mitigating rater effects) is an important issue that still remains under-explored in ITA (see Liu, 2013, for an exception).
Evidence-based research on ITA: Main topics, findings, and trends
Validity evidence concerning score-based inferences and uses
We intend to frame the validity of ITA using an argument-based approach (Chapelle et al., 2008, 2010). Han and Slatyer (2016) demonstrated how such an approach could be utilized to inform validation research in ITA. For test scores to be understood and used appropriately, interpreting researchers need to collect empirical evidence to support a number of interlocking inferences, including domain analysis and modelling (i.e., content validity), evaluation (i.e., scoring validity), generalization (i.e., generalizability), explanation (i.e., theory-based explanation or construct validity), extrapolation (i.e., authenticity), and utilization (i.e., consequential validity or washback effects) (for details, see Han & Slatyer, 2016, pp. 242–249).
To map relevant evidence to each inference, we find that there is relatively more evidence in support of the evaluation inference, particularly in high-stakes ICPT. For example, according to Han’s (2015b, 2016) review of ICPT practices worldwide, most of the certification tests surveyed emphasized rater reliability, and relied on Pearson’s correlation coefficient and/or Cronbach’s alpha to generate reliability evidence. Similarly, in interpreting research, researchers usually provide estimates of rater reliability to justify their further statistical analysis (Setton & Motta, 2007). For interpreter educational assessment, Liu et al. (2008) observed that reliability evidence seems to be lacking as most of the surveyed interpreting institutions did not provide any inter-rater reliability evidence for their summative assessment.
In addition, a number of studies have generated evidence for the inference of domain analysis and modelling, including ALTA Language Services (2007) and CCHI (2012), in which practice analysis surveys are used to profile interpreting practice to inform test design. When it comes to the generalization inference, Han (2016, 2019) presented preliminary generalizability evidence for ICPT and summative assessment of English–Chinese interpreting.
However, in terms of the other inferences, it would appear that rigorous empirical evidence has not been documented. In particular, the lack of solid validity evidence concerning the extrapolation and utilization inferences looms large in interpreter certification testing (Hale et al., 2012), and also in professional qualification examinations conducted in the educational context (Setton & Dawrant, 2016).
Scoring procedures
One of the most vibrant lines of research concerning scoring procedures is the development and/or validation of rubric-referenced rating scales for assessing interpreting (Lee, 2008; Lee, 2015; Liu, 2013; Tiselius, 2009; Wang et al., 2020; Yeh & Liu, 2006). Some of the rating scales are used in educational summative assessment (e.g., Clifford, 2005; Lee, 2008), others are intended for certification purposes (Liu, 2013), and still others are designed with the aim of establishing a national framework of reference for interpreter education (analogous to the Common European Framework of Reference for Languages) (Wang et al., 2020). Several rating scales have been initially validated (Han 2017; Lee, 2008; Liu, 2013) based on quantitative procedures such as correlational analysis and many-facet Rasch measurement to ascertain inter-rater reliability, concurrent validity, and scale utility; or on qualitative methods (e.g., questionnaires, interviews) to elicit rater perceptions of scale effectiveness. Although the research results generally support the utility of the rating scales, several potential problems emerge including non-specificity of scalar descriptors (Lee, 2008), inappropriate number of scale bands (Han, 2017; Lee, 2008), lack of sufficient benchmarked samples to each scale band (Liu, 2013), and difficulty in distinguishing between assessment criteria (Lee, 2008; Liu, 2013).
Another line of research concerns how different operationalization of scoring specifics would affect assessment outcomes. For example, Gile (1999) studied the effects of auditory versus visual presentation mode on fidelity assessments for English-to-French SI. Gile found that target-language renditions were assessed more leniently when presented in the auditory mode than the visual mode. In addition, a number of researchers are interested in how the use of source versus target text would impact on English–Chinese CI fidelity assessments. Yeh and Liu (2006) found that inter-rater reliability was higher in the group using the exemplar interpretation (r = .812, p < .001) than the group using the original speech transcript (r = .767, p < .001), although such a difference may not be practically significant.
A less explored strand of research relates to scoring methods, particularly the relative utility of one scoring method vis-à-vis another. Liu (2013) examined how fidelity scores derived from the itemized/atomistic analysis (i.e., proposition-based rating) would compare with those from rubric scoring. The correlational analysis showed that the two sets of scores achieved very high correlation (r = .945, p < .001). In addition, Han (2021) compared the assessment results obtained from comparative judgment and analytic rubric scoring of English–Chinese CI and found that overall rubric scoring outperformed comparative judgment in terms of concurrent validity and raters’/judges’ confidence in using these methods.
Rating behavior and rater-generated measurements
Given that raters play a critical role in ITA, a large proportion of the research is centered on raters, shedding light on rater effects, rating behavior and rater-generated measurements. One group of studies concentrated on rater effects such as rater severity (Han, 2015b), rater accuracy (Han, 2018a; Han & Zhao, 2020) and halo effects (Wu et al., 2013). For example, Han (2015b) introduced many-facet Rasch measurement (MFRM) to analyze rater severity/leniency in rater-mediated assessment of English-to-Chinese SI; Han (2018a) also trialed an MRFM-based Rater Accuracy Model to examine rater accuracy in peer assessment of English–Chinese CI. Similarly, Wu et al. (2013) relied on MFRM to examine potential halo effect, but did not detect such an effect in an assessment of English–Chinese CI.
Another group of empirical studies revolves around rater characteristics and their potential effects on rating behavior and/or rater-generated measurements. The rater characteristics under investigation include whether raters have received formal interpreting training, practiced professionally in a given market, and taught interpreting in educational institutions. Other variables of interest include raters’ language background/combination and the number of raters involved in assessment. For instance, raters could be classified into different cohorts, depending on the extent of formal interpreting training (Tiselius, 2009; Wu, 2010; Yeh & Liu, 2006), professional interpreting practice (Lee, 2008), interpreting teaching (Wang et al., 2015), and their language background (Su, 2019a). Regarding the effects of the rater characteristics on rater reliability, Wang et al. (2015) reported higher inter-rater reliability between two interpreter educators than between each interpreter educator and an interpreting practitioner. Tiselius (2009) also found that when assessing English-to-Swedish SI, the PR group (i.e., professional interpreters as raters) achieved higher inter-rater reliability than the SR group (i.e., student interpreters as raters). This finding, however, runs counter to Lee’s (2008) finding that higher inter-rater reliability was obtained for the SR raters than for the PR raters in the assessment of English-to-Korean CI. In addition, exploring how native Chinese speakers versus native English speakers as raters would influence their behavior in the assessment of Chinese-to-English SiI, Su (2019a) reported that the two rater groups were influenced by their socio-cultural background, and one rater group tended to display certain assessment behavior (i.e., notating, underlining, prior marking, post hoc marking) more frequently than the other group. Moreover, a number of studies on English–Chinese CI suggested that raters’ language background/combination may interact with interpreting directionality to affect rater reliability and score generalizability (see Han, 2017; Han & Riazi, 2018). Invariably, these authors reported that when raters with Mandarin Chinese as their L1 and English as their L2 were recruited to assess English–Chinese, bidirectional CI, they generally displayed greater self-consistency in assessing English-to-Chinese interpreting than the other way around; inter-rater reliability and score generalizability estimates were consistently higher for the English-to-Chinese direction than for the opposite direction. This effect could be termed the “mother tongue effect”.
Automatic assessment of spoken-language interpreting
With automated scoring engines being developed and deployed for L2 speaking, preliminary research efforts have also been made to explore the mechanism(s) underlying automatic assessment of spoken-language interpreting. Overall, there are two strands of research on this emergent topic (for a detailed review, see Han & Lu, 2021), with some researchers utilizing linear regression modelling to predict interpreting quality based on a range of (para)linguistic indices (Han et al., 2020; Ouyang et al., 2021; Yu & van Heuven, 2017), and others relying on statistical procedures originating from machine translation research to estimate quality (Stewart et al., 2018; Zhang, 2016).
On the one hand, Han et al. (2020) and Yu and van Heuven (2017) investigated the predictability of human-judged fluency ratings on the basis of objectively measured acoustic variables such as the number of (un)filled pauses, mean length of run, speech rate, and phonation/time ratio. Based on relatively small samples, a number of variables were initially identified as significant predictors of fluency ratings, including (effective) speech rate, mean length of salient pauses, and mean length of run. In a similar vein, Ouyang et al. (2021) sought to predict human ratings of interpreting proficiency using linguistic indices obtained from Coh-Metrix analysis (i.e., word count, lexical diversity, hypernymy of verbs, and frequency of first person singular). Results showed that the four predictors accounted for 60% of the variance in the human ratings.
On the other hand, to estimate SI quality for three language pairs (English–French, English–Italian, English–Japanese), Stewart et al. (2018) incorporated a number of interpreting-specific features (e.g., ratio of pause/hesitations/incomplete words, ratio of non-specific words) to the default quality estimation (QE) methodology in machine translation research. They found that the new QE system obtained respectable correlation scores between the predicted and true METEOR (Metric for Evaluation of Translation with Explicit Ordering) evaluation metric, thus supporting the feasibility of the new QE system. Zhang (2016) also proposed semantic-scoring metrics (i.e., MinE, MaxE) to estimate English-to-Chinese SI, and found that these metrics had higher correlations with human judgment than the traditional BLEU (Bilingual Evaluation Understudy) metric.
Challenges and directions
This section aims to reflect critically on the practice of and research on ITA, discussing three types of existing challenges concerning epistemology, psychometrics, and assessment practicality. In light of these challenges, it also provides suggestions for future practice and research.
Existing challenges
Epistemological challenges
The first type of challenge is epistemological and particularly concerns the knowledge gaps between ITA and language testing. The enterprise of ITA is essentially interdisciplinary or even multidisciplinary requiring intellectual contribution from myriad research fields. Ideally, practitioners and researchers of ITA need to have a mixed and well-balanced body of knowledge and competencies to ensure the reliability, validity, practicality, and (legal) defensibility of ITA. However, ITA has long been dominated by interpreting practitioners, educators, and researchers, also known as practitioners-cum-researchers or practisearchers (Gile, 1994), who are not necessarily equipped with the relevant know-how. In particular, although practisearchers are generally well versed in translation and interpreting studies, they usually lack sufficient knowledge and expertise of testing and measurement theory, language assessment literacy, and statistical analysis. Such epistemological gaps have been consistently alerted by a number of translation and interpreting researchers (Angelelli, 2009; Campbell & Hale, 2003; Hale et al., 2012; Han, 2018b, 2020; Sawyer, 2004; Wu, 2010). Although language testing theory has been used to inform ITA design (Hale et al., 2012) and sophisticated analytic techniques such as generalizability theory and MFRM trialed to examine rater-generated measurements, such efforts have been initiated only recently by very few researchers.
Psychometric challenges
Psychometric issues concerning the validity of score-based inferences and actions also pose challenges to the enterprise of ITA. The root causes of these challenges are as follows: (1) the lack of a scholarly consensus on a research agenda, and (2) the lack of a rigorous execution of such an agenda. Eventually, psychometric challenges boil down to the paucity of research-based, substantive evidence informing and/or supporting test design and development, operational scoring, and the attainment of positive washback.
First, there has been little hands-on, tailor-made, and research-informed guidance on how to design and develop an interpreting performance test in different assessment contexts. There is a genuine need for clear guidelines that help educators and testers develop interpreting tests and to accommodate a wide range of specificities associated with spoken-language interpreting. For instance, we need a customized descriptive framework that accounts for important parameters of an interpreting task (Campbell & Hale, 2003). Further insights are also needed to illuminate the processes of material selection, task authoring, task banking, determining task difficulty and test comparability, test equating, and standard setting.
Second, we have very limited understandings of how specificities of ITA would affect rating processes and rater-generated measurements. Such specificities include, for example, the use of analytic versus holistic rating scales, raters with different profiles of background characteristics, and rater training programs based on different formats and procedures. Our insufficient knowledge of their potential effects constitutes a threat to the psychometric soundness of our measurements.
Third, if we draw on the argument-based approach to test validity outlined in the section above, Validity evidence concerning score-based inferences and uses, we find that little empirical data has been collected to elucidate the consequential validity of score-based actions in ITA. Our scant knowledge of the consequences that ITA could have for interpreting students, trainers, and educators is a limiting factor in our understanding of the usefulness of ITA.
Practical challenges
There are a host of practical challenges to the practice of ITA, owing to fast-changing professional and educational landscapes. Prominent among them are technological and logistical challenges caused by the incorporation of real-life complexities in a testing environment, the recruitment of qualified raters to assess interpreting into languages with limited diffusion, and the scaling-up of testing capacity to accommodate a growing number of test candidates. In recent years, we have witnessed a burgeoning demand for technology-assisted interpreting for public services, especially during the global pandemic of COVID-19 (e.g., video remote interpreting) and interpreting for a diverse range of highly specialized subject matters. These developments pose a challenge to ITA, logistically and technologically, as test developers may want to recreate real-life complexities in a test. Available human resources may also be an issue when a certifying organization needs to certify an interpreter who interprets into a language of limited diffusion. Usually, qualified raters with a decent command of a rare language are difficult to find. Furthermore, the increasing number of test candidates for certification programs has to be accommodated. To test organizers, this means either the deployment of more logistical resources to cope with the demand, or resorting to efficient screening techniques and computerized testing and assessment.
Future directions
Given the above challenges, we venture to provide some pointers on future directions. First, to redress the epistemological gaps, interpreting testers need to partake in continuous professional development. They could also initiate cross-disciplinary collaboration with colleagues from mature disciplines, especially L2 language testers (see also Angelelli, 2009; Hale et al., 2012; Wu, 2010). There have been some joint efforts between interpreting researchers and language testers to enhance certification testing programs in Australia and the USA (Elder et al., 2016; National Center for States Courts, 2013; Stansfield & Hewitt, 2005). Some test developers also seek to involve subject matter experts in test design and development by conducting empirical practice analysis (NAATI, 2016) or eliciting expert opinions to inform test content (ALTA Language Services, 2007).
Second, to mitigate the psychometric challenges, interpreting educators and researchers should map out a research agenda and follow it up by conducting evidence-based substantive inquiries. In particular, four lines of research merit special attention which would benefit the practice of ITA considerably long term but have been severely under-researched or not investigated at all. The first and most important line of research would be to gain an in-depth understanding of interpreting ability/competence and examine its developmental trajectory via systematic collection of linguistic, cognitive, behavioral, and other types of evidence. Another line of research urgently needed is to investigate how operationalization of scoring specifics could affect rating processes and assessment outcomes. The next line of future investigation is to examine the interplay between rater characteristics, raters’ cognitive processes, and measurement results. The last line of research is to investigate the consequences or the washback effects of ITA on relevant stakeholders (e.g., students, test takers, educators and teachers, certifiers, consumers).
Third, to improve the efficiency and manageability of large-scale ITA, testing agencies and certifying bodies may need to consider using technology-assisted testing and assessment systems. For example, test delivery could be computerized; rater training could be made online; and operational scoring could be managed by a centralized system, although test security, cost-effectiveness, and feasibility may also be weighed in decision-making. In addition, given the fast-paced development of automatic assessment, testing agencies could use automated scoring engines for initial screening, meaning that only those test candidates who pass the screening test are able to be qualified for rater-mediated performance assessment. The coupling of human and machine power may save a considerable amount of financial, logistical, and human resources for certifying bodies (see also Han & Lu, 2021).
Conclusion
Interpreting testing and assessment has gained traction over the past decade thanks to the growing demand for high-quality interpreting services in different parts of the world. Despite the increasing amount of testing practice and research, few efforts have been made to generate an in-depth understanding of the state of affairs. This state-of-the-art review thus set out to summarize the existing literature, discover potential lacunae, and identify future trends. It foregrounded three major areas where the needs of testing and assessment have been rising over the years: interpreter training and education, professional certification, and interpreting research. The review also synthesized relevant practices and scoring procedures involved in the three areas, highlighting the operational diversity and methodological multiplicity of ITA. The review cast light on empirical studies on testing and assessing spoken-language interpreting, describing some of the preliminary investigations led by a small, but growing, group of interpreting researchers and testers to understand the intricacies and complexities of assessing language interpreting. Furthermore, the review identified existing problems, including the epistemological, psychometric, and practical challenges that hold back interpreting testing and assessment. Finally, the review ventured to provide some suggestions for moving forward practice and research. We have tried our best to provide a systematic review of the current state of the affairs and to incorporate diverse perspectives. Nonetheless, readers are duly cautioned that the publications cited in the article may be potentially biased in favor of published and widely available literature. Our personal preferences may also constitute a potential limiting factor. Overall, we believe that the review should afford valuable insights into the general features and trends of interpreting testing and assessment as a worthwhile endeavor and a promising research field.
