Abstract
Despite consistent calls for authentic stimuli in listening tests for better construct representation, unscripted texts have been rarely adopted in high-stakes listening tests due to perceived inefficiency. This study details how a local academic listening test was developed using authentic unscripted audio-visual texts from the local target language use (TLU) domain without compromising the reliability of the test results and validity of the score interpretations. The purpose of the listening test was to identify international students who need additional language support at a U.S. university. We show that efficiency persists when using authentic unscripted texts that are representative of the local context both at the test development phase and at the classification phase where placement decisions are made in a dependable manner. Expert judgments highlighted the improved correspondence of the listening test using locally sourced audio-visual texts to the local TLU domain, providing additional support for using the listening test for local placement purposes. Additionally, dimensionality assessments demonstrated that test design decisions inevitably entailed with using authentic unscripted texts did not threaten the internal structure of the test. We argue that local resources are indispensable in developing authentic test stimuli and in supporting the validity of local test interpretation and use.
Keywords
Using authentic, unscripted texts sourced directly from the target language use (TLU) domain instead of scripted texts helps increase confidence in the inferences we make from test scores (Ockey & Wagner, 2018). However, scripted texts have been predominantly used in large-scale standardized listening tests, the inauthentic nature of which may undermine the interpretations of the test scores (Wagner & Wagner, 2016). In addition, there has been little published research that details the development and validation of a listening test using authentic, unscripted texts. There is not only a general lack of established large-scale listening tests to model after but also limited references available for local test developers. These may have added to the difficulty of forgoing the convention and comfort of scripted texts in favor of using authentic unscripted texts in local tests, resulting in construct underrepresentation. Local tests, contrary to large-scale standardized assessments, have a certain degree of flexibility in the development process to circumvent some foreseeable challenges (Dimova et al., 2020), such as less immediate need to create a large number of test forms. Local testing settings are thus suited to exploring and instantiating the use of authentic unscripted stimuli that are readily available from local resources, a practice rarely feasible in large-scale standardized tests.
In response to the above observation in testing practices and the recent call for the use of authentic input texts in listening tests for stronger validity arguments (Ockey & Wagner, 2018), this paper details the development and validation of a local listening test using authentic, unscripted audio-visual texts sourced from the local context. The purposes of this study are (a) to document the process of developing a local listening test using authentic unscripted audio-visual texts and (b) to examine the extent to which placement decisions can be made in a reliable and valid manner in the local testing context. In doing so, we highlight the context, purpose, and use of local tests distinct from large-scale standardized proficiency assessments and illustrate that utilizing local resources to the full extent in a local test is indispensable in strengthening the meaningfulness of the construct and the relevance to the intended use of the local test. The use of local and consequently authentic resources not only supports the validity of the inferences we make from the test results but also renders the test results easily interpretable for local stakeholders. By adopting authentic texts representing the local context as input stimuli in a local listening test, test developers may be able to strike a balance between validity and the need for maximum efficiency in test development (Wagner et al., 2021).
Literature review
Authenticity of listening test input
Authenticity is defined as “the degree of correspondence of the characteristics of a given language test task to the features of a TLU task” (Bachman & Palmer, 1996, p. 23). Authenticity is integral to how we interpret and use test scores. The more reflective of the TLU domain the test is in terms of content, cognitive processes, and tasks, the more accurate and useful the inferences will be. Using unscripted spoken texts as input stimuli is one of several ways to enhance authenticity in listening tests.
Input stimuli in a listening test exist on a continuum of scriptedness (McCarthy & Carter, 1995; Tannen, 1982; Wagner, 2014b). Scripted spoken texts refer to texts written, revised, and edited respectively, recursively, and meticulously by item writers and test developers and recorded by voice actors. Unscripted spoken texts are on the other end of this continuum where composition and utterance occur simultaneously with little to no planning.
While scripted texts constitute some portion of the types of speech that second language speakers of English might encounter, scripted texts are inherently different from unscripted texts in three main areas: organization/discourse, lexico-grammatical, and phonological characteristics (Wagner & Toth, 2014; Wagner & Wagner, 2016). Unscripted spoken texts tend to be less linear, less systematically organized, and less rigorous in their organization in the discourse structure (Buck, 2001; Gilmore, 2007; Wagner & Wagner, 2016). They also tend to be more informal with slang, colloquialism, and less complex sentence structure (e.g., Biber, 1988; Chafe & Tannen, 1987). By the nature of their spontaneity, unscripted spoken texts also differ from scripted texts phonologically in that they feature more connected speech such as linking, assimilation, deletion, epenthesis, and reduction (Celce-Murcia et al., 1994); faster speech rate and more hesitation phenomena such as filled and unfilled pauses, false starts, hesitations, redundancies, and repetitions (Wagner & Toth, 2017).
Using only scripted spoken texts in listening tests thus leads to construct underrepresentation due to their inherent differences from unscripted spoken texts and poses threats to validity in that it is more difficult to make accurate inferences about how test takers would actually perform in the TLU domain (Ockey & Wagner, 2018). The differences between unscripted spoken texts and scripted spoken texts do not stop at the textual level, but also extend to their effects on test performance. The effects of the relative degree of scriptedness (or orality) on test scores have long been noted (e.g., Shohamy & Inbar, 1991). Studies comparing scripted spoken texts to unscripted (or authenticated) spoken texts have found that students performed better on listening comprehension tests using scripted spoken texts than on listening tests using unscripted (or authenticated) spoken texts (Wagner et al., 2021; Wagner & Toth, 2014). Taken together, using scripted spoken texts in a listening test not only leads to issues with construct representation and accurate inferences on test takers’ ability but also may result in overestimation in performance. This, in local tests, may lead to serious consequences such as false positives and concomitantly limited opportunities and access to resources for better academic achievement (Messick, 1994; Wall et al., 1994).
Another way to improve authenticity in listening tests is using videos as input stimuli to simulate the listening tasks in the TLU domain. A vast majority of academic listening includes visual, contextual, and non-verbal cues available for listeners, whereas the audio-only channel deprives listeners of such cues (Wagner, 2014a). Using videos in an academic listening test thus promotes authenticity and strengthens the validity of the score interpretations in terms of construct relevant variance, as the testing context largely mirrors what the test takers will do in a real-life lecture (Ockey & Wagner, 2018).
Despite consistent calls for increased authenticity in listening texts in large-scale assessments, few high-stakes listening tests use authentic, unscripted spoken texts (Wagner, 2016; Wagner & Wagner, 2016). One particular obstacle to using authentic unscripted spoken texts is that it is difficult to adhere strictly to the predetermined test specification in terms of both item writing and stimuli development due to the characteristics of unscripted texts noted above (Ockey & Wagner, 2018; Wagner & Toth, 2017). Because of this, test developers turn to creating scripted texts rather than finding unscripted spoken texts as it is arguably easier, more cost-effective, more practical, and more efficient (Carr, 2011; Wagner, 2014a, 2016; Wagner et al., 2021). This is, however, an argument that is largely applicable to large-scale standardized tests. Local tests that are flexible and adaptable to local needs (Dimova et al., 2020) could be an ideal breeding ground for turning the argument of efficiency on its head due to the distinct nature of local domains where participants and texts representative of the TLU domain can be sourced in large numbers.
Another obstacle is that to our knowledge no empirical study has examined the reliability and the validity of a listening test using authentic unscripted audio-visual texts. Without such documentation, local test developers and practitioners lack both the evidence and data to arrive at and support their decision to move from scripted texts to unscripted texts, which may have contributed to the relatively less pronounced use of unscripted spoken texts in local listening tests. It is thus important to document the steps involved in the development of a local large-scale listening test using authentic unscripted texts so that local test developers and practitioners can make more informed decisions when considering the use of authentic texts.
The onus on the test developers intending to use unscripted texts as stimuli in listening tests extends further than simply making the decision to do so and accruing the materials accordingly. The impact of the decision also seeps into item development and test construction, which may potentially influence the dimensionality of test taker performance in terms of certain design features that are concomitant to the use of unscripted spoken texts. Using authentic, unscripted spoken texts as input stimuli in a listening test poses unique challenges for test developers and item writers, as authentic texts do not contain the same amount of testable information in the same length of time (Rossi & Brunfaut, 2021; Wagner, 2014a). In order to generate as many comprehension questions as possible within a limited length of time, test developers using authentic unscripted spoken texts are tasked with making particular design decisions that are not typically observed in traditional large-scale high-stakes exams: longer texts per topic, fewer topics, and different formats of comprehension questions. These have the potential to invite possible construct-irrelevant variance into the test, namely the format effect (In’nami & Koizumi, 2009) and the topic effect (Jennings et al., 1999).
The format effect mostly refers to differences in the construct or trait being measured and/or in test scores that arise from using different response formats within a test (In’nami & Koizumi, 2009). Listening tests using unscripted spoken texts are more likely to include different response formats, as it is difficult to generate a sufficient number of items using authentic unscripted texts with solely multiple-choice questions due to the relatively redundant nature of the discourse and consequently limited testable information in unscripted texts (Wagner, 2014a). The true/false format may help in item and test development, as its simple and efficient format allows test developers to create items from limited testable information in authentic, unscripted texts (Buck, 2001; Carr, 2011; Downing, 1992). Although the true/false format has been shown to measure the same construct as the multiple-choice format (e.g., Ebel, 1971), it has been noted for its relatively weak performance in terms of reliability and validity while being less discriminating in comparison to the multiple-choice format (e.g., Chaudron & Richards, 1990; Frisbie, 1973; Grosse & Wright, 1985). Including two response formats that have been shown to perform psychometrically in a different manner may create unwanted construct-irrelevant differences in test performance in terms of dimensionality (Buck, 2001; Field, 2009; Fulcher, 1999; In’nami & Koizumi, 2009; Messick, 1989).
The topic effect is inherent in many topic-based tests whereby the topic has the potential to have an effect on the test taker performance through factors such as individual interest and prior knowledge in a way that is irrelevant to the construct (Bachman, 1990; Chiang & Dunkel, 1992; Jennings et al., 1999). The topic effect is suspected to be more pronounced in listening tests especially using authentic unscripted stimuli due to the smaller number of texts to accommodate the increased length of the stimuli (Wagner, 2014a).
Local testing context
Higher education has the responsibility to support enrolled students’ continued development of their language proficiency for their academic success (Read, 2015). Standardized test scores (e.g., SAT, ACT, and TOEFL iBT) are typically used for admissions purposes in higher education, but the use of these test scores does not extend to identifying those in need of language support and placing them into appropriate courses because large-scale standardized tests are not intended to align with any particular curriculum (Fox, 2009). If the standardized test scores were the only barometer by which to make these decisions, those who entered through alternative pathways without a standardized proficiency score (Read, 2016; Wall et al., 1994), or those who do meet the minimum but may continue to face language challenges in higher education (Elder, 2017), would slip through the cracks, missing the opportunity to take advantage of the available language support resources. Standardized tests also target a wider range of overall proficiency than what is immediate to the needs of the local stakeholders, and test takers’ scores may not represent their most current English language abilities (Brown, 2003; Dimova et al., 2020; Kokhan, 2012). All of these factors combined would render using standardized test scores for a local placement purpose inappropriate. Post-entry assessment that is aligned with the local instructional values and curricula can thus be beneficial when making decisions such as identifying students who could benefit from additional language support and placing them into a course that is appropriate for their proficiency levels (Dimova et al., 2020).
A new academic English placement test, the Indiana Academic English Test (IAET), was developed and implemented at Indiana University to assess academic English proficiency of matriculated international students and make decisions about their exemptions from or placements into academic English language support courses. The test development project was launched at the request of instructors, departments, and the Office of the Vice Provost for Undergraduate Education that a computer-delivered test be developed for improved placement decisions to help students achieve further academic potential. The IAET served to both address the concern from empirical research above at the local level and to meet the demands from local stakeholders of better means to assess the needs of international students. Prior to the development of the IAET, a paper-and-pencil test called English Proficiency Exam had been used (hereafter the retired test). The listening section of the retired test was an audio-only test using scripted spoken texts.
Those who are required to take the IAET are mostly newly admitted international undergraduate students whose first language is not English, but international graduate students whose first language is not English may also take the exam at the request of their departments. Students with a TOEFL iBT score below 105 or an IELTS band score below 7.5 may be required to take the test. Those who have completed three full years of secondary school in an English-speaking country are exempt from the exam, according to the university policy.
The IAET is composed of three sections: writing, listening, and oral interview. In this study, only the listening section of the test, hereafter referred to as the Academic Listening Test, is discussed. Indiana University requires the students who do not pass the Academic Listening Test to take a single-level, 2-credit, intensive 8 weeks academic listening course. The academic listening course focuses on developing students’ listening skills needed for comprehending undergraduate-level lectures at the university. The scores on the Academic Listening Test are used solely for placement purposes.
Research questions
Given that the Academic Listening Test results are used to identify the need for additional language support in lectures and place students into or exempt them from an academic listening course, obtaining empirical evidence that the target construct and content of the Academic Listening Test are aligned with the TLU domain and the academic listening course provides support for the validity of the test score interpretation (Fulcher, 1997; Wall et al., 1994). The investigation into whether the placement decisions based on the test results are reasonably dependable is also indispensable in ensuring the validity of the test use (Fulcher, 1997, 1999; Long et al., 2018). Additionally, assessing the dimensionality of the test performance in terms of the response format and the topic was deemed necessary to ensure that methodological choices that inevitably accompany the use of authentic, unscripted spoken texts did not create unwanted separate dimensions within test performance.
In light of the above, we addressed the following research questions in the present study:
How dependable is the Academic Listening Test using authentic unscripted spoken texts for placement decisions?
To what extent is the Academic Listening Test aligned with the local target-language use domain and the academic listening course content?
To what extent do the response format and the topic create construct-irrelevant differences in the Academic Listening Test?
Method
Materials
The academic listening test
The target construct was the ability to understand academic lectures at a university in the United States. The decision to use audio-visual lecture videos as listening input was made to simulate how the listening tasks are carried out in the TLU domain; students typically view and listen to the speaker in academic lectures (Ockey & Wagner, 2018).
Selecting speakers
We obtained real-world lecture videos at Indiana University in two ways: recording videos of actual lectures in the classroom and obtaining lecture videos recorded for online courses. Online lectures were included to acknowledge the wide range of online courses that Indiana University was offering to undergraduate students at the time of the test development. The overarching goal was to obtain samples of lectures that are representative of the TLU domain. The following criteria were used to select candidate lectures:
The course is a general education course (GenEd, the breadth requirements) targeting the first-year undergraduate students.
The topic is not too advanced (lectures from the first 3 weeks).
There is a single speaker giving a lecture.
The speaker is a highly proficient speaker of English.
The lecture consists of an expository spoken text which has sufficient information for creating comprehension questions.
The first two criteria were to make sure that the content of the lectures was easily accessible to newly admitted international undergraduate students so that the target construct of the listening test could be assessed, not their prior knowledge of the subject matter. The third criterion was to reflect the principal genre of instruction in university settings: largely monologue lectures by a single speaker (Lee, 2009; Lynch, 2011). The fourth criterion was to simulate the TLU domain where highly proficient L2 speakers of English, not just native speakers of English, teach undergraduate courses at the university (Wagner, 2014a). The fifth criterion was to exclude the portions of the lectures where the speaker was not giving a lecture (e.g., students filling out a questionnaire, checking attendance).
We first contacted over 30 instructors and professors (hereafter instructors) currently teaching GenEd courses at the university. Many of the targeted courses had a title starting with the phrase “Introduction to.” We explained the need for developing an authentic academic listening test at the local level and asked if they would be willing to either grant us permission to record their lectures or provide video recordings of their online lectures. Four instructors graciously gave us permission to video-record their lectures in the classroom, and six instructors of online courses gave us permission to use their lecture videos. For the in-person lectures, one of the researchers went into the classroom and video-recorded the lecture using a digital camera, microphone, and a tripod. For the online lectures, we viewed over four dozen lecture videos from the course archives and selected videos that met the above-mentioned criteria. The lengths of the lecture videos varied widely from 30 minutes to 3 hours. From this screening process, five instructors’ lectures were selected. Additionally, two professors, as external reviewers, in the field of applied linguistics reviewed samples of the five instructors’ lecture videos in a formal meeting and discussed whether the lectures were comprehensible and whether they required prior knowledge. All five instructors’ lectures were deemed appropriate in this review phase.
Preparing lecture videos
Before preparing the lecture videos, we fully transcribed the lectures. Then, we edited the videos with the following three guidelines in mind: (1) the editing must be done at the discourse level to avoid manipulating the language and disrupting the flow of the lecture; (2) there must be sufficient, meaningful information to generate listening comprehension questions; and (3) the video should be reasonably short. From each of the five instructors’ lectures, we extracted two video clips such that each edited video lasted from 6 to 8 minutes. Since we edited videos at the discourse level, the speakers’ pauses, false starts, repetitions, and self-corrections were kept intact. The texts of the edited videos were separately transcribed for item writing. Each version of the listening test we developed involved four videos. In the version of the test we used in this study, two videos were of a female speaker on the topic of food ethics in an online course. These videos showed the speaker sitting at a desk and the accompanying lecture slides to the left. The other two videos were of a male speaker on the topic of sociology in an in-person course. These videos showed the speaker’s torso and lecture slides projected in the back facing the classroom.
Item writing and piloting
Following the test and item specifications that we developed for this listening test, we wrote items designed to tap into test takers’ ability to identify the main ideas of the lecture, ability to identify important details, ability to recognize the speaker’s opinions, and ability to make inferences from the information stated in the text (Buck, 2001; Song, 2008). We created a total of 30 items for each version of the listening test. The items were either in a four-option multiple-choice format or a true/false format (stylized as “correct” or “incorrect” in the actual test). The choice of these two response formats was to meet both the stakeholders’ requirement for a quick turnaround in scoring and the challenge of generating a sufficient number of items from limited testable information in unscripted spoken texts.
This version of the test was piloted twice over two semesters with a total of 95 examinees. For piloting purposes, the test was computer-delivered using Qualtrics in a computer lab to first-year international undergraduate students. They were either students who were enrolled in an academic listening course because they had not passed the existing listening test or students enrolled in an academic writing course who had passed the existing listening test. After the first pilot, problematic items were flagged and revised based on item difficulty and item discrimination indices for criterion-referenced testing (Carr, 2011). The criteria for flagging an item were item difficulty lower than 0.10 and higher than 0.90 and a B-index below 0.19 (Brown, 2005). The notes that students took during the pilot were used to create or modify underperforming distractors. The same process was repeated after the second pilot. Based on the piloting, 45 minutes was determined as the appropriate time allotment. The cut score was determined using the contrasting groups method, reflecting the local context. An examinee-centered approach to standard setting, the contrasting groups method utilizes expert judgments (in this case, instructors of the academic listening course) on prototypical performance meeting the standards for the class to set the cut score (Van Nijlen & Janssen, 2008).
Questionnaire
A survey questionnaire was constructed to assess the relevance of the test to the academic listening curriculum to ensure that the test can be used for placement purposes, based on previous studies that explored the relevance of (placement) tests using expert judgments (Cumming et al., 2004; Fulcher, 1997; Wall et al., 1994; Weir & Wu, 2006; Wu & Stansfield, 2001). The questionnaire consisted of three parts: background information, the instructors’ judgments on the Academic Listening Test and its stimuli, and their judgments on the English Proficiency Exam (i.e., the retired listening test), and its stimuli. The judgments were elicited on a 5-point scale where 1 indicated “Not well at all” and 5 indicated “Extremely well.” The order in which the tests were shown to the instructors was counterbalanced and randomized.
The questionnaire first presented the listening stimuli and asked whether the listening texts were representative of the TLU domain and the academic listening course in terms of content and language. Then, it presented the items along with the test stimuli and asked how well the test elicited behaviors that aligned with the academic listening course curriculum. For each question on the survey, a comment box was provided for the instructors to further elaborate on their choices. The full questionnaire can be found in Appendix 1.
Participants
Test takers
The data for this study are from the 2017 administration of the test. A total of 518 newly matriculated international students (481 undergraduate and 37 graduate students) from diverse L1 backgrounds took the test. Their TOEFL iBT total scores ranged from 60 to 113 (M = 87.96, SD = 8.86), and their TOEFL iBT listening section scores ranged from 10 to 30 (M = 22.05, SD = 3.81).
Instructors
Five previous instructors of the academic listening course, into which those who fail the Academic Listening Test are placed, participated in the study to provide expert judgments on the alignment of the test to the course objectives and content of the academic listening course. All five were female. One had a PhD in linguistics, two had a master’s degree in TESOL, one had a master’s degree in Chinese linguistics, and one had a master’s degree in Spanish linguistics. All instructors obtained their bachelor’s degrees in universities where English was the primary language of instruction. They were all current or former graduate students at the same institution. Table 1 summarizes their teaching backgrounds.
Instructor profile.
Procedure
The final, student-facing version of the listening test was commissioned to the IT division under the Office of the Vice Provost for Undergraduate Education, consolidating the test-taking, scoring, decision-making, and reporting into a single web interface. The listening test was computer-delivered through this website.
The test was administered in computer labs with a room supervisor and 2–3 proctors in each lab. Each computer station was separated with dividers and equipped with a set of headphones. The students were provided with paper upon the start of the test and were encouraged to take notes, all of which were collected afterwards, but they were informed that their notes would not be evaluated. Everyone began the test at the same time. Students were first presented with a sample lecture video (also an authentic unscripted spoken text taken from an online classroom) with a sample comprehension question to help them familiarize themselves with and anticipate the format of the listening test. Students could view the comprehension questions only after watching each lecture video. They were given 45 minutes to complete the test.
The questionnaire to the instructors was delivered online using Qualtrics. They completed the questionnaire independently, and it took 40–60 minutes for them to finish it. This study was approved by the Institutional Review Board.
Analysis
To answer the first research question, a squared-error loss agreement approach was used to obtain the degree of classification consistency, reflecting the context of criterion-referenced testing (Berk, 1984; Brown, 1990). Specifically, the Φ(λ) dependability index (Brennan, 1980) was calculated for its sensitivity to decision consistency while accounting for the distances from the cut score (Brown & Hudson, 2002). The following formula for dichotomously scored tests was used (Brown, 1990, p. 88):
where λ is the lambda or the cut point expressed by a proportion of raw test scores; k is the number of test items; Mp is the mean of the proportion scores; and Sp is the standard deviation of the proportion scores. The signal-to-noise ratio (S/N[λ]) for Φ(λ) was also calculated as an additional way of interpreting Φ(λ) decision dependability (Brennan, 1984). It estimates the magnitude of a desired signal to the level of other measurement noise in making placement decisions about test takers around the cut score. The higher the S/N(λ) is, the more meaningful information is for classification decisions. The S/N(λ) was calculated based on the following formula (Brown, 2014): S/N(λ) = Φ(λ)/(1 − Φ[λ]).
To address the second research question on the alignment of the test with the TLU domain and the course content, the instructors’ ratings on listening stimuli were collected on each section of the two listening tests, the new listening test (the Academic Listening Test) and the retired listening test (English Proficiency Exam). The Academic Listening Test is a 30-item computer-based listening test and contains two sections with two authentic unscripted lectures followed by true-false and multiple-choice questions with four options. The retired English Proficiency Exam was a 36-item paper-based test with three subsections: (a) short dialogues, (b) long dialogues, and (c) lectures with scripted spoken texts. All items were multiple-choice questions with four options. Both took 45 minutes to administer. The raw count for each judgment category was multiplied by its numeric value (not well at all = 1, slightly well = 2, moderately well = 3, very well = 4, extremely well = 5) and then averaged by the number of sections within each test and the number of respondents, yielding numeric values ranging from 1 to 5.
DETECT (Dimensionality Evaluation to Enumerate Contributing Traits) and DIMTEST procedures from DIMPACK were used to answer the third research question (Jang & Roussos, 2007; Zhang & Stout, 1999). The DETECT procedure is primarily an exploratory technique that gauges the degree of multidimensionality by identifying clusters of items that are dimensionally homogeneous (Svetina & Levy, 2012). DIMTEST is a nonparametric method of assessing dimensionality based on estimated conditional covariances, with hypothesis testing on two subsets of items (referred to as AT, the Assessment Subtest, and PT, the Partitioning Subtest) that have been judged to be possibly dimensionally distinct from one another. The null hypothesis is that AT and PT measure the same ability, and the alternative hypothesis is that AT is dimensionally distinct from PT (Jang & Roussos, 2007). The dataset (n = 518) was first randomly split into a training sample (n = 259) and a cross-validation sample (n = 259). First, an exploratory DETECT analysis was run twice on the training sample, once with the response format as a partition and another with the topic as a partition. Then, confirmatory DIMTEST analyses were conducted on the cross-validation sample to test the assumption of unidimensionality. Two confirmatory tests of DIMTEST were run; one test specified multiple-choice items (k = 13) as AT and true/false items as PT (k = 17), and the other test specified questions pertaining to food ethics as AT (k = 15) and questions pertaining to sociology as PT (k = 15).
Results
Dependability of the test
The Φ(λ) dependability index where n = 518, λ = 0.63, k = 30, Mp = 0.71, and Sp = 0.15, Φ(λ) was calculated as 0.79, a moderately high value, indicating that the Academic Listening Test consistently classified students around the cut score for course placement. Note that the Φ(λ) dependability index of the retired listening test was 0.83 where n = 790, λ = 0.72, k = 36, Mp = 0.81, and Sp = 0.13. The signal-to-noise ratio (S/N[λ]) for Φ(λ) in the Academic Listening Test was estimated at 3.68 (Φ[λ]/1 − Φ[λ] = 0.79/[1 − 0.79] = 3.68), indicating that the signal for placement decisions based on the cut score was almost four times stronger than the background noise.
Instructors’ responses to the questionnaire
Alignment of the test to the target language use domain
The instructors provided ratings on whether the listening stimuli used in each test were representative of topics and the level of content in terms of abstractness and concreteness in first-year university classes. Additionally, they were asked to judge the extent to which each test represented the academic language in first-year university classes. Using mean section ratings as a descriptive summary statistic, Table 2 compares the five instructors’ ratings across the new test and the retired test. As the instructors’ ratings are represented on an underlying ordinal scale, no statistical tests were conducted on these and following comparisons. The instructors judged the stimuli used in the new listening test using authentic audio-visual texts to function “moderately well” and “well” at being representative of first-year university classes.
The average ratings on the listening tests’ representativeness of topic, level of content, and language in first-year university classrooms.
Note: ALT: the listening section of the Academic English Test; EPE: the listening section of the English Proficiency Exam (the retired test).
The five instructors’ ratings indicated that the Academic Listening Test functioned “moderately well” at representing the topics and the level of content in first-year university classes and functioned “well” at representing the language used in first-year university classes. On being asked to comment on what aspects of the Academic Listening Test are effective (or not effective) at reflecting the academic English in first-year university classes, instructors remarked that “the more natural speech is much closer to what is found in first-year classes,” “the language many students are actually exposed to in the first year.” The instructors’ comments highlight the extent to which the language used in the audio-visual listening text of the new Academic Listening Test corresponds to the actual language used in the classroom.
Alignment of the test to the academic listening course
Table 3 shows the average ratings on the alignment of the listening tests to the student learning outcomes (SLOs) in the academic listening course. On average, the instructors rated the Academic Listening Test to function “moderately well,” higher than the listening section of the retired English Proficiency Exam, on its ability to elicit behaviors that accord with SLOs in the academic listening course.
The average ratings on the tests’ alignment to the student learning outcomes in the academic listening course.
Note: ALT: the listening section of the Academic English Test; EPE: the listening section of the English Proficiency Exam (the retired test).
In addition to the tests’ alignment with the SLO, the instructors were asked to rate the extent to which the test stimuli resemble the listening materials used in the course and the extent to which the tests elicit these types of listening skills taught in the academic listening course. Table 4 shows the average ratings on the overall strength of each test on its relevance to the academic listening course as judged by the instructors. The instructors rated the Academic Listening Test to function “moderately well” at demonstrating correspondence to the listening materials in the academic listening course, and “slightly well” at eliciting the listening skills taught in the course, both higher than the retired English Proficiency Exam. Two instructors in their comments reported using YouTube videos, in addition to the videos provided by the textbook, as instructional materials in their classes to promote these skills. One instructor commented on the lack of coverage by the Academic Listening Test on producing lecture outlines, a skill that was reported by the instructors to be taught in class, which may partially explain the lower score on the test’s alignment to SLO tied to note-taking than the first two SLOs in Table 3.
The average ratings of the tests’ relevance to the academic listening course.
Note: ALT: the listening section of the Academic English Test; EPE: the listening section of the English Proficiency Exam (the retired test).
Dimensionality
Table 5 shows the results from the exploratory DETECT analyses. The DETECT indices for both DETECT tests were less than 0.20, indicating unidimensionality (Kim, 1994). The IDN and ratio R indices for both DETECT tests were lower than 0.80, an evidence against multidimensionality in the data (Jang & Roussos, 2007; Kim, 1994). The results from the two DETECT tests do not provide any evidence of multidimensionality in the data due to response format or topic.
Statistics for exploratory analysis of multidimensionality in DETECT.
Note: DETECT: Dimensionality Evaluation to Enumerate Contributing Traits.
Table 6 shows the results from the two confirmatory analyses of DIMTEST. The null hypothesis of unidimensionality was retained for all two tests. The two subsets within the Academic Listening Test according to the response format (one set of multiple-choice items and the other set of true/false items) were not dimensionally distinct from one another, indicating that the response format forms a single dimension. Similarly, the two subsets within the Academic Listening Test specified according to the topic (one set of items on food ethics and the other on sociology) were not significantly different from one another in terms of dimensionality, suggesting that different topics in the listening test did not retain separate dimensions for each topic but rather formed a single dimension in test taker performance.
Results from confirmatory DIMTEST.
Note: AT: Assessment Subtest; PT: Partitioning Subtest.
Discussion
The goal of the study was to document the process of developing a local listening test using authentic unscripted spoken texts and to investigate the extent to which placement decisions can be made in a reliable and valid manner in the local testing context. In this section, we discuss the latter by answering the three research questions that guided this study.
The first research question addressed the dependability of the newly developed academic listening test using authentic unscripted spoken texts for placement decisions. The Φ(λ) dependability index of 0.79 shows that the degree of decision consistency while accounting for the distances from the cut score of 19 out of 30 is 79%, indicating that the decision dependability for placement into the academic listening course was relatively high. Considering that the Φ(λ) dependability index of the retired listening test is 0.83, the Academic Listening Test using unscripted spoken texts achieved comparable dependability (0.79) even with fewer items (30 vs. 36) and a lower number of test takers (518 vs. 790). These results are encouraging in that several design decisions were made in the test development process to accommodate the unique challenges of using authentic unscripted spoken texts, such as long texts with relatively little testable information (Wagner, 2014a). These results suggest that using multimodal unscripted listening stimuli can achieve a satisfactory level of decision consistency while enhancing test authenticity. Future test developers may rest assured that the use of authentic unscripted spoken texts in a listening test is not necessarily a risk to the reliability of test results.
The second research question explored the degree to which the Academic Listening Test was aligned with the TLU domain and the academic listening course content. Overall, the instructors’ responses showed that the Academic Listening Test is appropriately aligned with the TLU domain as well as the curriculum of the course. The instructors’ ratings on the Academic Listening Test were all above 3.0 (i.e., representing the domain “Moderately well” on a scale of 1 to 5), except for the one rating in response to the statement “Elicit types of listening skills taught in the academic listening course” (2.9).
The noticeably higher ratings on the correspondence of the Academic Listening Test to the ESL curriculum than the retired listening test using scripted texts are noteworthy. More specifically, the Academic Listening Test was rated higher than its retired scripted counterpart in its relevance to the academic listening course in terms of the behaviors outlined in the SLOs, course materials, and listening skills taught in class. The high ratings on the Academic Listening Test (all at “Moderately well”) may, in part, be traced back to the audio-visual nature of the input stimuli. The academic listening course makes use of videos as part of the instructional materials included in the textbook to help students develop listening skills for academic purposes. Locally sourced authentic stimuli that is of an audio-visual nature strengthened the relevance of the test to the local context, namely the ESL curriculum. The results lend support to using the Academic Listening Test to make decisions on whether or not a student would benefit from taking the academic listening course. The alignment of the Academic Listening Test to the academic listening course curriculum further alleviates the possibility of students being misplaced into or out of the academic listening course (Fox, 2009; Kokhan, 2013), with additional supporting evidence from the dependability index.
Most importantly, the Academic Listening Test using authentic texts representing the local context seems to align well with the broader TLU domain. The test was consistently rated to function “moderately well” at being representative of classes targeted toward first-year students at Indiana University, including the topic and the level of content of these classes. The characterization of the Academic Listening Test as functioning “well” at representing the types of language being used in first-year university classes speaks further to the benefits of using authentic unscripted spoken texts in listening tests. The high ratings on the test’s correspondence especially to the academic English used in the classrooms can be attributed to the unscripted, audio-visual, and consequently authentic nature of the listening stimuli used in the Academic Listening Test that reflects the language in the TLU domains (Wagner & Toth, 2014). Because the videos were sourced directly from introductory-level classes at Indiana University, they elicited the listening skills used in first-year classes. That is, the topics, the content, and the types of language in the listening stimuli are the ones actually used in classes taken by freshly matriculated students. Thus, the videos are more appropriate than scripted spoken texts because the videos better align with the listening skills needed by the test takers in the near future, once the semester starts. Obtaining the authentic, unscripted test stimuli from local resources instead of relying on scripted texts and further employing the audio-visual mode in stimuli presentation led to improved construct representation, which consequently allows the test developers to make more valid inferences about test takers’ listening ability in the real-world communicative domain (Messick, 1989; Ockey & Wagner, 2018; Wagner & Toth, 2014). Test administrators thus can make more valid placement decisions grounded in improved interpretations in a way that is aligned with both the curriculum and the TLU domain, rendering the test and its consequences beneficial to potential test takers (Fox, 2009; Kokhan, 2012). In addition, using authentic listening texts directly sourced from the TLU domain in a post-entry test presents an opportunity for students to be exposed to and recognize the level of academic performance expected of them in the instructional settings at the university, which is of a practical value in envisioning their ensuing academic work (Dimova et al., 2020).
One interesting trend in the comments we received from the instructors was in reference to the effectiveness of the response format. The academic listening course instructors noted the response format, specifically the true/false response format as being “effective” at eliciting behaviors in accordance with the following SLOs: “indicate knowledge of new vocabulary” and “demonstrate improvement in comprehension of academic lectures.” These comments suggest that the response formats adopted in the Academic Listening Test not only contribute to efficiency in item development and scoring, but also facilitate the correspondence of the Academic Listening Test to the academic curriculum.
One area in which the Academic Listening Test received a relatively low rating was note-taking and producing lecture outlines or summaries. Although students were encouraged to take notes while taking the test, the Academic Listening Test does not score note-taking nor does it require lecture outlines or summaries as part of the response. Even though the instructors emphasized the production of lecture outlines and summaries as a part of their academic listening instruction, the instructors gave the note-taking part of the test a lower rating (“slightly well”) than the test’s other areas when they were asked to consider the effectiveness of the test’s areas on eliciting listening skills taught in the academic listening course. However, we do not necessarily consider the low rating on this area to be a weakness of the test since not every student wants or needs to take notes while listening. Nevertheless, in future test development, additional ways to assess students’ understanding of the organization of the listening texts may be considered. One possible option would be using the drag-and-drop interface with a graphic flow chart to better test for the relationship between ideas (Alderson et al., 2000; Grabe, 2008), another area pointed out by the instructors as needing improvement.
The third research question asked the degree to which the response format and the topic created construct-irrelevant differences in the Academic Listening Test. The dimensionality assessments identified neither the response format nor the topic as sources of multidimensionality. This suggests that the Academic Listening Test consistently measured a single construct irrespective of the response format and the topic. This lack of evidence of psychometric multidimensionality suggests it is not psychometrically problematic to use authentic, unscripted spoken texts, despite the development and construct-coverage challenges associated with them as item components (Rossi & Brunfaut, 2021; Wagner, 2014a), and indeed one might use them especially because they simultaneously retain the validity of test-score use and interpretation. The limited number of topics did not constitute a separate dimension, nor did the two different response formats. Authentic, unscripted audio-visual texts can be used in listening tests without the fear of their potential adverse effect on the test structure and student performance.
Using scripted spoken texts has always been viewed as a necessary evil in order to achieve the maximum efficiency at the test development stage (Wagner et al., 2021). In large-scale standardized tests developed by private organizations, the need for efficiency may trump the need for authentic listening texts because efficiency will reduce costs. However, the concept of efficiency need not be interpreted nor applied in the same way for local tests housed within universities. The present study suggests that using unscripted spoken texts from local resources may in fact be “easier” and “more efficient” for local tests in that input texts can be sourced directly from the TLU domain. Multiple response formats can be adopted in an efficient manner without sacrificing the decision consistency, the interpretation of the construct, and the sufficient amount of testable information. In addition, using authentic, unscripted spoken texts results in a dependable classification system that aligns better with the local standard whereby efficiency is present not only at the test development stage but also in the classification procedure. Altogether, these ultimately allow local stakeholders to make valid inferences on the degree to which matriculated international students have mastered the level of listening proficiency needed for academic studies at an English-medium university while ensuring that the entire procedure stays somewhat manageable for test developers. While some researchers argue that authentic texts are too context-specific to be used in listening assessment with the added difficulty of item generation (Rossi & Brunfaut, 2021), we showed neither to be a problem in a local testing setting. In fact, for a listening test in a specific, local context, opting for authentic listening texts would align with the specific purpose of the local test, thereby rendering the use of both the text and the test more valid.
Developing a local listening test with stimuli from local domains is ripe for many more opportunities to strengthen its link to the local context. For instance, videos that did not pass the screening may instead, with the agreement from the recorded instructors, be developed into instructional materials for the academic listening course, rendering the process of collecting local audio-visual materials fruitful. Another avenue for an improved tie between the curriculum and the test is to share the test results with the academic listening course instructors for diagnostic purposes (Dimova et al., 2020; Harding et al., 2015; Read, 2016).
Limitations and future research
Some limitations of the study include the specific context in which the test was commissioned and administered. One is the limited response format. We had to limit the response format to two selected-response formats due to the specific demands from the stakeholders. It would be worthwhile to explore other response formats to highlight the authenticity of the unscripted spoken texts. Another is the limited range of topics in the listening stimuli. Due to the inherent challenge of developing a listening test using unscripted texts (longer, less dense texts than scripted texts), only lectures from two authentic courses could be included in the test. The construct representation of the TLU domain is thus limited in that the two courses cannot possibly cover the breadth of the GenEd courses. Additional investigation into whether this would incur possible differential item functioning between different test taker groups may be a potential venue to pursue. The study would also benefit from examining other aspects of validity evidence, such as cognitive validity, to investigate the extent to which test takers are engaged in similar processes as they would in the TLU domain during the test (Taylor & Geranpayeh, 2011). Finally, in order to gauge the benefits of the unscripted nature of the listening text, a more controlled design would be helpful whereby the visual nature of the unscripted listening text is excluded and then presented in an audio-only mode to the instructors in comparison to the audio-only listening text of the retired listening test.
The process for language test development is not linear or fixed, but rather iterative (Davidson & Lynch, 2002). By the same token, although all the items in the operational Academic Listening Test were piloted before their implementation, it is necessary to keep monitoring the item qualities in terms of item difficulty and discrimination indices for quality control of the test. For this, it is critical to ensure ongoing collaboration between item writers for item vetting and construction and instructional technology specialists for maintenance of web interface and test delivery.
Conclusion
The present study showed how stronger validity arguments can be made for a local academic listening test by using local, unscripted audio-visual texts. Overall, the Academic Listening Test using authentic unscripted audio-visual texts yielded dependable placement decisions, with higher ratings than its retired, scripted counterpart on its adherence to the TLU domain and the academic listening course in terms of language and content. The use of two response formats, multiple-choice and true/false response formats, was further substantiated in dimensionality assessments, demonstrating that authentic unscripted spoken texts and the concomitant test design choices entailed with their use are not a hindrance to the psychometric structure of the test. Combined together with the evidence of consistency in classification decisions, the results of the study reinforce the valid interpretation and use of the Academic Listening Test with authentic audio-visual texts grounded in the local context.
Local resources, as can be seen in the study, are indispensable in obtaining authentic materials and ultimately strengthening the validity evidence for local tests. The entire process of developing and piloting four different versions of the listening test took us roughly a year, and we hope that this is encouraging for local test developers to start navigating the uncharted waters. Local tests with access to audio(-visual) stimuli in the TLU domain are the perfect site to address the continued call for increased authenticity in listening tests (Wagner, 2016). Local test developers no longer have to compromise for the sake of efficiency, with authentic test stimuli that align with the local TLU domain and the ESL curriculum resulting in dependable placement decisions that pertain to the local context.
Footnotes
Appendix 1
The Survey Questionnaire to Elicit Expert Judgment on the Correspondence of the Listening Tests to the Curriculum and the TLU Domain adapted from Cumming et al. (2004), Wall et al., (1994), and Weir and Wu (2006)
Acknowledgements
We would like to thank the lecturers and professors at Indiana University who graciously allowed us to use their lectures for assessment purposes, instructors at the English Language Improvement Program at the Department of Second Language Studies for their invaluable help with piloting and test quality assurance, and Center for Innovative Teaching and Learning (CITL) and Center for Language Technology (CeLT) for technical help with accessing and preparing the lectures. Special thanks goes to the Office of the Vice Provost for Undergraduate Education (OVPUE) for commissioning and supporting the project, especially the staff at IT Division Clinton McKay, David Waicukauski, Anesu Chaora, G. Martin Berry, and Ben Martin for creating an incredible web interface for the Indiana Academic English Test. Finally, we thank the four anonymous reviewers for their helpful and insightful comments.
Author Note
This publication and research represent the work and views of the author and do not necessarily represent the views or positions of any entities with which she may be currently affiliated.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
