Abstract
This study investigated the optimum task sequence for second language (L2) novice learners of English. One set of task sequences was manipulated using a deductive and theoretical SSARC (simplify–stabilize–automatize–restructure–complexify) model, and two sets of task sequences were manipulated based on a teacher’s inductive classroom observations. A total of 76 undergraduates at a private university in Korea were divided into three groups for the task sequences: task complexity (TC), guided planning with vocabulary (GPV), and guided planning with content (GPC). While the four oral tasks were sequenced according to the resource-directing dimensions [± elements] and [± reasoning] in all three groups, the TC group received pretask planning, the GPV group received teacher-led guided planning with words, and the GPC group received teacher-led guided planning with content for the resource-dispersing dimensions. Pretest and posttest of syntactic complexity, accuracy, and fluency were used as the main data. The analysis showed that the TC group outperformed the GPV and GPC groups significantly in increasing overall syntactic complexity, and the GPV group outperformed the GPC group significantly in improving speed fluency. Both sequencing TC and GPV tasks significantly increased syntactic complexity and speed fluency. Sequencing TC tasks decreased accuracy and increased dysfluency, whereas sequencing GPV tasks increased accuracy and decreased dysfluency. Meanwhile, sequencing GPC tasks did not produce overall positive effects on oral performance compared with the two other groups.
I Introduction
Although task-based language teaching (TBLT) has provided a solid theoretical and pedagogical foundation for language teaching over the last three decades, the TBLT’s applicability and effectiveness, particularly for novice second language (L2) 1 learners, remain controversial among teachers. From a pedagogical perspective, novice L2 learners rely heavily on their first language (L1) and use limited L2; consequently, performing real life tasks is considered extremely demanding for novice L2 learners (Butler, 2011; Cao, 2018; Carless, 2004; Swan, 2005). In addition, many teachers contend that it is not appropriate to apply TBLT to novice learners because they believe novice learners need more lexical and grammatical knowledge before they actually speak and that pedagogical evidence regarding TBLT in real classrooms is lacking. These concerns allow language teachers to avoid task-based syllabi (Kim, 2019) or to provide weak versions of TBLT (Carless, 2004; Swan, 2005). From a theoretical perspective, Long (2016), an advocate of a strong version of TBLT, argued that ‘students in TBLT courses learn as much grammar and vocabulary, even at the beginner’s level, with instruction geared to learners’ processing capacity, and with the additional crucial advantage that course content is relevant to their communicative needs’ (p. 26). Many scholars have pointed out that the units of analysis in any genuine TBLT should be tasks rather than texts from commercial language teaching materials, and a central tenet of TBLT is that tasks should be graded and sequenced according to their inherent complexity even for novice L2 learners. However, most participants in task-based research have been ‘at the roughly intermediate proficiency level’ (Ellis, 2009, p. 491). Only one study (to my knowledge) has conducted task sequence research with roughly set novice-high L2 learners (with the average score 33.7 in terms of the Oxford Placement Test). Therefore, if L2 learners’ proficiency is at the novice level, the feasibility and effectiveness of (a strong version of) TBLT have not only been in doubt due to learners’ linguistic unpreparedness or teachers’ beliefs, but also because these factors have not been extensively examined by task-based researchers.
To actually implement TBLT for novice L2 learners, it is necessary to show how much and what L2 learning benefits novice learners can acquire through task sequences and to determine which task sequences suit these learners. Therefore, the purpose of the present study was to investigate appropriate task sequences for novice L2 learners to improve their L2 speaking in task-based syllabi. Despite the importance of task sequences in task-based syllabi, several theoretical and pedagogical gaps motivated the present study. First, there are no widely agreed-upon criteria for the task sequences (Baralt, 2014; Lambert & Robinson, 2014; Malicka, 2014). Different scholars have suggested different versions of task sequence criteria in task design. Robinson (2001, 2005, 2011) suggested ‘the triadic componential framework’ for the task design, entailing task complexity, task condition, and task difficulty. Ellis (2003) proposed four criteria for grading tasks (input, conditions, processes, and outcomes), and Skehan (1998, 2003, 2009) proposed three criteria for sequencing (code complexity, cognitive complexity, and communicative stress). Robinson’s task complexity forms the sole basis for regarding sequencing tasks in his simplify–stabilize–automatize–restructure–complexify (SSARC) model, whereas Ellis’s and Skehan’s all different task criteria should be considered for the task sequence decision. Recent task sequence studies have tended to compare differently ordered sequences. For instance, Levkina and Gilabert (2014) compared three groups: simple-to-complex, complex-to-simple, and randomized sequences. While Malicka (2014) compared five groups: simple (S) – more complex (C+) – complex (C), CC+S, C+CS, CSC+, and C+SC sequences, Baralt (2014) compared four groups: SSC, CCS, CSC, and SCS sequences. These previous studies manipulated task sequences in relation to resource-directing dimensions, but they did not seem to consider resource-dispersing dimensions: a departure from the SSARC model. They also did not fit into either Ellis’s grading tasks or Skehan’s task difficulty. Therefore, appropriate task sequence criteria with stronger theoretical foundations require investigation. Second, task sequence studies with novice L2 learners have been rare; most recent studies (Baralt, 2014; Lambert & Robinson, 2014; Levkina & Gilabert, 2014; Thompson, 2014) of task sequences have examined L2 learners at intermediate or higher proficiency levels. Therefore, the optimum task sequence for novice L2 learners should be investigated for language teachers and researchers who regard TBLT as inapplicable to novice learners. Finally, most previous studies (Butler, 2011; Cao, 2018; Carless, 2004; Swan, 2005) that have questioned the feasibility and effectiveness of TBLT have focused on teachers’ and/or learners’ perceptions of TBLT by collecting qualitative data including via questionnaires, interviews, class observations, and review papers. Without actual classroom-based research examining novice L2 learners’ real linguistic performance and data, it is difficult to conclude that task-based syllabi are not applicable for beginners. This highlights the need for a classroom-based study of novice L2 learners. Thus, the present study aimed to examine optimum and contextualized task sequences for novice L2 learners following the theoretical and deductive SSARC model and empirical and inductive teacher-led guided planning in a classroom setting.
II Literature review
1 Deductive approach: The SSARC model for task sequence
Robinson (2005, 2011) claimed that the main pedagogic purpose of the triadic componential framework, involving task complexity (cognitive factors), task condition (interactive factors), and task difficulty (learner factors), is to provide a rationale for how to sequence tasks that foster L2 learning. Robinson (2010, 2011) proposed the SSARC model for task sequences based on the triadic componential framework, and this model follows two important instructional-design principles. The first principle is: ‘only the cognitive demands of tasks contributing to their intrinsic conceptual and cognitive processing complexity are sequenced’ (Lambert & Robinson, 2014, p. 210). This means that task sequence decisions are only influenced by the tasks themselves (i.e. task complexity among the above three factors), but not influenced by interaction or learner factors that vary depending on classroom context. Task complexity is subcategorized into resource-directing dimensions and resource-dispersing dimensions. Resource-directing dimensions control cognitive demands in tasks by, for instance, increasing the number of elements or the complexity of reasoning. Increasing task complexity along resource-directing dimensions induces learners to connect their cognitive attentional and memory resources, thereby cultivating the proper context for form-function mapping, which is necessary in L2 development (Robinson, 2011). Resource-dispersing dimensions control procedure demands in tasks by, for instance, eliminating the availability of planning time before performing tasks. Reducing planning time promotes ‘greater control over, and faster access to existing interlanguage systems of knowledge’ (Robinson, 2010, p. 248). The second principle is: ‘increase resource-dispersing dimensions of complexity first (e.g. from + to − planning time), and then increase resource-directing dimensions (e.g. from − to + intentional reasoning)’ (Lambert & Robinson, 2014, pp. 210–211). Figure 1 shows an example SSARC model, which has the following three steps.

Increasing cognitive demands in the simplify–stabilize–automatize–restructure–complexify (SSARC) model.
First, the simple-stable (SS) task on both resource-directing and resource-dispersing dimensions is performed. This first task establishes a cognitively and procedurally simple state (− elements and + planning). Next, while the resource-directing dimensions remain simple (− elements), the complexity of the resource-dispersing dimensions is increased by no planning (from + planning to − planning). This change of complexity in procedural demands promotes faster access to the automatization (A) of the current interlanguage system. Finally, the complexity levels of both the resource-directing and resource-dispersing dimensions are increased the most by improving the resource-directing dimensions (from − elements to + elements) and keeping tasks procedurally complex (− planning). This promotes the restructuring and complexity (RC) of the current interlanguage system. Thus, the SSARC model provides a fixed set of task sequences only manipulated by task complexity, not by learner factors (e.g. proficiency level, motivation, and interest) or teachers’ in-class observations. Therefore, the model represents a theory-driven and deductive approach (Skehan, 2016).
Although this model supplies well-structured and easy-to-follow task sequence criteria, recent task sequence studies (Baralt, 2014; Malicka, 2014) have modified it to test its effectiveness on oral performance. Malicka (2014) manipulated task complexity of simple, complex, and more complex tasks using two resource-directing dimensions, [± elements] and [± reasoning], but apparently did not consider the resource-dispersing dimensions, which are different from the SSARC model. Since the study did not use a pretest–posttest design, three tasks were compared in oral performance (complexity–accuracy–fluency, or CAF). Related to task sequence, the analysis showed no statistical difference in CAF between the simple–complex–more-complex sequence group (n = 5) and the randomized sequence group (n = 5), which meant that task sequence order did produce different learning outcomes. In addition, Baralt (2014) controlled task complexity by [± intentional reasoning] and two simple task (S) and complex task (C) sets were sequenced as follows: SSC, CCS, CSC, and SCS, which is also slightly different from the SSARC model. The study had a pretest–posttest design (oral and written production tests) and it detected language-related episodes (LREs) and the specific L2 target form (Spanish past subjunctive verbs) were detected. The analysis showed that the participants who performed task sequences containing more complex tasks (CCS and CSC) generated the most LREs and produced the target form more than those who performed SSC and SCS in the posttest. However, it found no significant difference between the two groups (CCS and CSC), meaning that task sequence order did not result in different learning outcomes. These recent task sequence studies do not seem to have followed the SSARC model in a straightforward manner; yet Lambert and Robinson’s (2014) study is the exception. They manipulated task sequence using resource-directing dimensions [± elements] and [± reasoning] and resource-dispersing dimensions [± planning], [± prior knowledge], [± steps], and [± multi-tasking]. Two Control and SSARC groups’ written performance – specific syntactic complexity (coordination and subordination), lexical complexity, specific accuracy (i.e. five types of errors on familiar forms per 100 words), explicit reasoning in writing, and summary writing – was compared between the pretest and the posttest. Regarding complexity and accuracy, the study found significant improvement in syntactic complexity in both the Control and SSARC groups, but it found no difference in accuracy in the two groups. It also found no significant difference between the two groups in language learning. Therefore, researchers have not yet conducted task sequence studies for L2 novice learners following the theoretical and deductive SSARC model (to my knowledge); thus, the present study is needed.
2 Inductive approach: Guided planning for task sequence
Skehan (2016) claimed that the inductive approach in task-based research takes ‘a case-by-case view, identifying task features from general theory, previous research, or classroom experience, and then explores whether each feature has a systematic relationship with performance’ (p. 35). This means that task (condition) features (e.g. planning, task repetition, posttask, or task-processing requirements) are carefully considered in a local context to make generalized predictions of L2 learning and performance. Task difficulty (Skehan, 1998) and task planning (Ellis, 2005) are inductive task condition features, which are psycholinguistically rooted in the limited attentional resource capacity and Levelt’s (1989) speech production model. 2 Since learners have limited attentional capacities to simultaneously attend to meaning (at conceptualization) and form (at formulation) when performing tasks, planning time can free up learners’ cognitive space in their attentional resources and thus lead to successful task production (Ellis, 2005). Skehan (2016) argued that these task features have produced more consistent and predictable improvements in CAF than Robinson’s task complexity (2011). For instance, (pretask) planning consistently benefits fluency and there is a trade-off increase between syntactic complexity and accuracy (Ellis, 2009). In addition, learners’ proficiency levels have some impacts on planning. The effects of planning are better for intermediate learners than advanced learners (Kawauchi, 2005; Kim, 2017).
Although the crucial pedagogical purpose of task difficulty is to sequence tasks in terms of their inherent difficulty, the task sequence model in this field lacks a well-structured theoretical model and support from empirical studies because the practical origin of task difficulty research may ‘attest to the inductive nature of the inquiry’ (Skehan, 2016, p. 38). Although Robinson (2011) argued that these task sequence criteria are variant and learner-dependent, implying that it would be hard to make task sequence decisions in the task syllabi due to divergent local contexts, Skehan (2016) argued that these inductive task features tend to be ‘relatively low-theory in nature and often contain developments from previous literature or are connected with pedagogic suggestions’ (p. 40). Thus, teachers and researchers need to consider what existing results and findings in real classrooms can contribute to theory. As a language teacher, I tried to adapt TBLT for novice L2 learners over three separate semesters, and one of the constraints involved in implementing TBLT was novice learners’ heavy dependence on sentence-writing through searching for words in electronic dictionaries whenever they started performing oral tasks (Kim, 2019). Based on my previous study and the subsequent two-month in-class observations, I identified an empirical need for guided pretask planning based on the limited attentional resource capacity for task sequence decisions of novice L2 learners; in this vein, two relevant studies warrant attention.
In Foster and Skehan (1999), 66 intermediate L2 undergraduates were divided into six groups based on focus and source of planning: language by teacher, language by group, content by teacher, content by group, and two control (solitary planning and no planning) groups. The target forms were modals and conditionals, and the first 5 minutes of debate were recorded as data and different effects of guided planning on CAF were detected. The study found that the teacher-led planning improved accuracy significantly, while the solitary planning increased syntactic complexity, fluency, and turn length. No difference was found between language- and content-guided planning. In addition, Thompson (2014) investigated the effects of sequenced guided planning (GP) and sequenced guided–unguided planning (GUP) groups on learning relative clauses among 24 intermediate L2 undergraduates. The participants in both groups took the one-way oral narrative tasks and grammatical judgement tests as the pretest and posttest. The two groups performed five-treatment oral tasks (Tasks A to E), and the five tasks were sequenced by cognitive complexity and obligatory numbers of relative clauses. Task A needed seven times relative clauses, Tasks B and C needed nine times, and Tasks D and E needed 10 times. In the GP group, all five tasks were combined with 10-minute guided planning (Task A), 7-minute guided planning (Tasks B and C), and 4-minute guided planning (Tasks D and E). In the GUP group, Task A was combined with 10-minute guided planning, but Tasks B through E were combined with unguided planning. While the analysis showed that both groups improved syntactic complexity based on an increase in mean score, it showed no difference between the groups (and the results for accuracy, fluency, grammatical judgement tests were not reported). This study was the first (to my knowledge) to combine a task sequence and guided planning, and the manipulation of the task sequence did not fit into the SSARC model but seemed to fit into task difficulty. However, the change in the length of guided/unguided planning (10–7–4 minutes) may have vague theoretical or practical rationales; thus, new task sequence sets combined with guided pretask planning are needed.
3 Research questions
There is no widely agreed upon set of task sequence criteria, and attempts to examine task sequence for novice L2 learners have been extremely scarce. In terms of teachers’ points of view, the SSARC model is fascinating because it makes it easier to design task sequences due to the fact that teachers consider only one invariant factor, task complexity, and its model has been well established. However, Skehan (2016) argued that the inductively motivated task difficulty (from case to theory) tends to show more consistent results than the deductively motivated one (i.e. the SSARC model). Therefore, localized TBLT for L2 novice adult learners is needed in classroom-based settings and task sequence serves an effective way to involve students and gradually introduce them to consecutive simple-to-complex tasks. To identify the optimum task sequence, the present study investigated the effects of three different sets of task sequences on L2 novice learners’ oral performance. One set of task sequences was developed based on the theoretical and deductive SSARC model, and two sets of task sequences combined with guided (pretask) planning were developed based on the teacher’s empirical and inductive classroom observations. The present study addressed the following four research questions:
To what extent do the three task sequence groups differ from each other in syntactic complexity, accuracy, and fluency of English in novice L2 learners’ oral performance?
To what extent does the task sequence manipulated by task complexity affect syntactic complexity, accuracy, and fluency of English in novice L2 learners’ oral performance?
To what extent does the task sequence manipulated by guided planning with vocabulary affect syntactic complexity, accuracy, and fluency of English in novice L2 learners’ oral performance?
To what extent does the task sequence manipulated by guided planning with content affect syntactic complexity, accuracy, and fluency of English in novice L2 learners’ oral performance?
III Methods
1 Participants and settings
A total of 76 undergraduates learning English as a foreign language (EFL) at a private university in Seoul, South Korea participated in this study. Their average age was 20.34, ranging from 18 to 25 years, including 45 females and 31 males, and Korean (n = 70), Japanese (n = 3), Chinese (n = 2), and French (n = 1) nationalities. Based on their speaking and writing scores on the college’s placement test, students were assigned to one of three English proficiency levels: novice (Level 1), intermediate (Level 2), and advanced levels (Level 3). Students could then voluntarily enroll in English classes within their assigned levels. All the participants in the study were found to be at Level 1 proficiency and joined a required college English course, Basic College English, which consisted of two hours every week over a 16-week semester. To select participants with novice to low intermediate levels of English at the time of the task sequence treatment, the oral narrative pretest was rated based on the public version of the American Council on the Teaching of Foreign Languages (ACTFL) proficiency guideline by a researcher (a bilingual teacher and the author) and a native English-speaking rater (Cronbach’s alpha = .970). Participants whose English proficiency on the pretest was found to be above mid-intermediate in terms of the ACTFL were excluded from the main data. Participants who did not join all the task sequence treatment sessions in the classroom due to tardiness or missed classes were also excluded. Thus, a total of 49 participants were included in the focal task sequence data. The participants were divided into the following three experimental groups: task complexity (TC, n = 16), guided planning with vocabulary (GPV, n = 16), and guided planning with content (GPC, n = 17) groups.
The goal of the course was to help students improve their general English skills (especially speaking and listening rather than reading and writing) by applying various educational practices. Throughout the semester, the researcher, who is a bilingual teacher with over 10 years of teaching experience with both traditional and TBLT lessons, made observational field notes. 3 Based on these observations and the application of several tasks in the real classroom before the experiment, three types of task sequences were designed and conducted after the midterm exam period.
2 Task materials
In the present study, three series of pictures were used as task materials. The three sets of pictures for the simple, complex, and pre- and posttasks all conveyed narrative stories, and their resource-directing dimensions were manipulated by [± elements] and [± reasoning]. The number of characters in the stories was for [± elements], and the complexity of the reasoning leading to the problems was for [± reasoning]. First, for the pre- and posttask assessing L2 oral performance, a series of pictures was selected from the EFL materials on the website and the task complexity level between the simple and the complex task was set as average. The first story has two main characters and one minor character. It is about a man waiting for a bus at the bus station who accidentally meets a woman. He falls in love all of sudden and then he goes to the flower shop to buy some flowers for her. When he comes back and tries to give them to her, another man who got off the bus takes the flowers instead of her and the man therefore fails to give her the flowers. Second, for the simple task, a series of pictures of two main characters who provide simpler reasons for creating problems in the story, was adopted from Heaton (1975). The story is about a boy who gets off a bus and accidentally drops one of the boxes he is carrying. The boy goes home without knowing that he lost the little box. A man picks up the box and runs after him to return it. The boy notices that the man is following him so he starts running, but the man catches up with the boy and returns the box. Finally, for the complex task, a series of pictures was adopted from Yule (1997), wherein three main characters and two minor characters provide complex reasons for causing problems in the story. The story is about a woman who goes to the supermarket where she happens to meet her friend and her friend’s little son. The two women are busy talking, so they do not notice when the boy takes a bottle of wine and puts it into his mother’s friend’s bag. After talking and saying goodbye to her friend, the woman pushes her cart to the cash register to pay for her items. When she leaves the store, one of the supermarket clerks stops her because of the bottle in her bag. The clerk reports her to the police and then she is arrested by a police officer for stealing the wine.
3 Procedure and task sequence
a Experimental procedure
The study was carried out during a single semester between September and November 2018, and the task sequence treatment was conducted in early November. Beginning in the second week of the course, the researcher – a bilingual teacher – started writing field notes, which included the above information, focusing specifically on participants’ behavioral features when performing the tasks with the aim of developing empirical support for task sequence decisions. In the second week, the participants’ English background and basic information (major, age, academic year in college) were also obtained, and a needs analysis and oral proficiency test were implemented to inform task design. The class was taught utilizing a combination of presentation-practice-production 4 and TBLT practices, and two (GPV and GPC) task sequences were designed inductively based on the field notes (see note 3) with inductive and contextualized TBLT perspectives to monitor participants’ behavioral features in class: (1) heavy reliance on electronic dictionaries for unknown words, (2) preference for writing complete sentences before doing meaning-based communicative tasks, (3) requesting teacher modeling before doing meaning-based communicative tasks, (4) using their first languages more in two-way tasks, and (5) self-speaking before the tasks. The participants in the three groups – namely, TC, GPV, and GPC – completed the picture-based one-way oral narrative tasks. Figure 2 presents the three different task sequence procedures and treatments that were conducted in intact classes during one lesson.

Task sequence procedure.
Before oral performance, participants received one minute to view the pictures. While telling the story, the participants recorded their oral performance with voice recorders on their cellular phones. After finishing the recordings, they returned the series of pictures to the teacher. After the pretest, the participants in the three groups performed different sequencing tasks, including four simple–complex-designed tasks (Task 1 to Task 4). All tasks were one-way oral narrative tasks based on a series of pictures, and the second and fourth tasks were recorded. After finishing the four sequencing tasks, the participants told the story based on the first series of pictures with no planning conditions, which means that they began performing the oral narrative without preparatory time immediately after receiving the pictures. They recorded the stories they told. After finishing all performances and recordings, participants sent their four voice recording files directly to the teacher through the KakaoTalk 5 instant messaging application for smartphones.
b Three task sequence operations
Task sequences were implemented following Robinson’s (2010) SSARC task sequence (TC) model after two months of inductive real class observation (GPV and GPC). First, for the TC group, four task sequences were designed based on deductive and theoretical suppositions regarding task complexity. The resource-directing dimension was manipulated by [± elements] and [± reasoning], whereas the resource-dispersing dimension was manipulated by [± pretask planning] for 5 minutes. The pretask planning time was set by the previous pilot study. The first task was cognitively simple (− elements and reasoning) and procedurally simple (+ planning), the second task was cognitively simple but procedurally complex (− planning), the third task was cognitively complex (+ elements and reasoning) but procedurally simple (+ planning), and the fourth task was cognitively complex and procedurally complex (− planning).
Second, for the GPV group, four task sequences were designed based on the researcher’s inductive observations with the addition of teacher-led guided (pretask) planning. Four tasks were sequenced based on increasing cognitive complexity (from simple to complex), and the teacher supported guided planning as students viewed the new pictures. For the guided planning, important words (mainly nouns and verbs) were slowly introduced for fewer than 5 minutes while the teacher pointed out actions and objects in the pictures with her finger. She restricted the use of complete sentences when reviewing the words to avoid modeling how to narrate the story. The words (e.g. follow, run, bus station, go shopping, push, or police officer) were introduced before the tasks. The first task was cognitively simple, involving fewer elements and less reasoning, and participants received vocabulary guidance before performing the task. The second task was cognitively simple, but the teacher provided no planning time or guidance. The third task was cognitively complex with more elements and reasoning, but the teacher provided words-related support, whereas the fourth task was cognitively complex without guidance.
Lastly, for the GPC group, four task sequences were designed based on the researcher’s inductive observations, as with the GPV group. Four tasks were sequenced according to increasing cognitive complexity and the teacher supported guided (pretask) planning as students viewed the new pictures. For the guided planning, the teacher told the whole story once in less than 5 minutes. Without focusing on certain words or structures, she told the story shown in the new pictures as slowly as possible. The first task was cognitively simple, involving fewer elements and less reasoning, and participants received the content of the story before performing the task. The second task was cognitively simple, but the teacher provided no planning time or guidance. The third task was cognitively complex, but the teacher provided content-related support. Finally, the fourth task was cognitively complex without guidance. As Figure 2 shows, Tasks 2 and 4 in the three groups had the same conditions, but Tasks 1 and 3 were different. To avoid unequal treatment, the three groups were assigned similar times (fewer than 5 minutes) for either pretask planning or guided planning. To prompt participants’ active involvement in the tasks, they were told they needed to complete four recordings and send their voice files to the teacher after finishing all in-class tasks, and that their final recordings would be graded as part of their class mission score.
4 Measurement and data analysis
Syntactic complexity, accuracy, and fluency were used as measures of oral performance. Although various measures have been used in previous task sequence studies, applications of multiple general and specific measures have been rare. Previous task sequence studies have generally used more specific measures (of target structure) rather than general measures. The occurrence of spatial expression (e.g. near the sofa) with vocabulary test (Levkina & Gilabert, 2014), the number of relative clauses and relative clauses per analysis of speech unit (AS-unit) (Thompson, 2014), and past subjunctive expression in Spanish with language-related episodes (Baralt, 2014) were used as dependent variables. Malicka (2014) is the exception when it comes to using general CAF measures, including words per clause, errors per AS-unit, and speech rate. Thus, when choosing measures, the following aims were regarded as important in the present study: using both general and specific measures in complexity 6 and accuracy, using no redundant measures, and considering participants’ low proficiency in oral performance following the guidelines of Foster, Tonkyn, and Wigglesworth (2000) and Norris and Ortega (2009) (see Table 1).
Eleven measures of complexity–accuracy–fluency (CAF).
Note. ◆ indicates a reversed relation in measures. AS = analysis of speech.
For syntactic complexity, the following six different measures were selected based on research by Norris and Ortega (2009): mean length of T-unit (MLT), mean length of AS-unit (MLU), clauses per AS-unit (C/AS), dependent clauses per clause (DC/C), mean length of clause (MLC), and number of frequent tensed verbs (FreqT). For accuracy, three measures were chosen: errors per AS-unit (E/AS), error-free clauses per clause (EFC/C), and correct verb forms (CVF). For fluency, the selected measures were number of syllables per one minute (Syll/M) and dysfluency ratio (DysFlu) as suggested by Tavakoli and Skehan (2005), wherein the total number of repetitions, restarts, and self-repairs was divided by the total number of seconds and multiplied by 100.
To analyse the data from the 98 oral narrative tasks included in the pretest and posttest, the researcher and a native English-speaking rater coded together. For syntactic complexity and fluency, the numbers of AS-units, T-units, dependent clauses, clauses, words, syllables, and dysfluent words were counted. Twenty percent of the 98 tasks were coded separately, and a one-to-one conference was held to reach agreement. Finally, the researcher coded the rest of the data. For accuracy, both the researcher and the same English-speaking rater found all morphosyntactic errors in the data separately (except article-related errors). Since participants’ English proficiency was novice in the present study, developmentally late occurring article errors might produce subtle changes in their accuracy, so they were excluded. The raters’ reliability based on the number of errors was .990 in the pretest and .977 in the posttest (Cronbach’s alpha). Any differences in judgement of the errors were reconciled through consultation. For statistical significance, four tests were employed in the present study: (1) kurtosis and skewness test on all measures of the pretest and posttest to check normal distribution of the data, (2) one-way ANOVA to check the homogeneity of the three pretests in the TC, GPV, and GPC groups, (3) MANOVA to check for statistical significance in the main effects between the pretests and posttests and the three task sequences, and in the interactional effects between the pretest–posttest and the task sequences, and (4) a series of paired sample t-tests to compare the pretests and posttests in 11 measures of CAF in each group. The alpha for gaining statistical difference was set at .05 using the Statistical Package for Social Sciences Statistics 25.0 version.
IV Results
Before addressing the four research questions, two statistical calculations were conducted. First, all the data in the pretest and posttest followed the approximate normal distribution in terms of skewness and kurtosis; thus, parametric statistics were used in the present study. Second, to detect homogeneity in the TC, GPV, and GPC groups’ pretests, a one-way ANOVA with the Scheffe test was carried out. As Table 2 shows, the analysis found no significant difference in nine measures, and a difference in two measures (E/AS and EFC/C). However, a post-hoc Scheffe test showed no significant difference between the two measures (see Table 3). To sum up, before assigning the task sequence treatments in each group, the analysis showed no significant difference between the participants’ oral performance among the groups.
Differences in the pretest results of the three groups.
Note. *p < .05. Abbreviations: see Table 1 for full forms.
Difference in the pretest of E/AS and EFC/C in terms of the Scheffe test.
Notes. E/AS = errors per analysis-of-speech unit. EFC/C = error-free clauses per clause. TC = task complexity group, GPV = guided planning with vocabulary group; GPC = guided planning with content group.
1 Differences between the three groups
The first research question addressed differences in syntactic complexity, accuracy, and fluency of English in L2 oral performance between the three task sequence groups. To detect cross-group differences between the pretests and posttests of all three groups (Pre vs. Post) and the sum of each group’s pretest and posttest (Task_seq) and their interaction, MANOVA was employed (see Table 4).
Overall differences between the pretest–posttest and between task sequences.
Notes. Pre = pretest. Post = posttest. Task_seq = task sequence. ***p < .001.
The analysis showed significant differences between the pretests and posttests of all three groups (F = 3.636, p < .001, partial η 2 = .350) as well as significant differences between the three different task sequence groups in the sum of their performance (F = 2.949, p < .001, partial η 2 = .304) in terms of Wilks’ Lambda. However, it showed no interaction between pretests-posttests and task sequence (F = .610, p > .05, partial η 2 = .083). This means that the effects of task sequence and the difference between the pretest and posttest were independent of each other.
To test the differences in the posttests between the three groups, another MANOVA was carried out. The analysis showed no significant difference between the posttests (F = 1.627, p > .05, partial η 2 = .358) in all three groups. Compared to the results in Table 4, this finding indicates that although there were significant differences between the pretest and posttest scores in all three groups, the degree of the changes in scores from the pretests to the posttests between the three groups may not have been dramatic. To find the specific differences in the posttests of the three groups, ANOVA with LSD was employed (see Table 5).
Specific differences in the posttests between the three groups.
Notes. Task_seq = task sequence. TC = task complexity group. GPV = guided planning with vocabulary group. GPC = guided planning with content group. *p < .05.
Specific differences in the posttests of the three groups were detected between the TC group and the GPV group and between the TC group and the GPC group in two syntactic complexity measures (MLT and MLU). The oral performance in the TC group (M = 11.63) in overall syntactic complexity (MLT) was significantly higher than in the GPV group (M = 9.74) or the GPC group (M = 9.87). The oral performance in the TC group (M = 11.43) in overall syntactic complexity (MLU) was significantly higher than in the GPV group (M = 9.55) or the GPC group (M = 9.62). In addition, the posttests between the three groups differed in one fluency measure (Syll/M). Oral performance in the GPV group (M = 109.69) in speed fluency was significantly higher than in the GPC group (M = 86.86). Regarding the findings in Table 4 (a significant difference between the pretests and posttests of all three groups), a series of paired sample t-tests was employed to determine how the pretest and posttest CAF scores changed in each group.
2 Gains between tests: Effects of task sequence manipulated by task complexity
The second research question addressed the effects of task sequences manipulated by task complexity on the syntactic complexity, accuracy, and fluency of English in L2 oral performance (see Table 6). Regarding syntactic complexity, sequencing TC tasks improved syntactic complexity significantly in MLT (t = −2.765, p < .05), MLU (t = −2.603, p < .05), C/AS (t = −2.295, p < .05), DC/C (t = −2.435, p < .05), and FreqT (t = −6.611, p < .001) except MLC (t = −1.642, p > .05). Regarding accuracy, sequencing TC tasks did not improve the overall accuracy in E/AS (t = −1.605, p > .05), EFC/C (t = .778, p > .05), and CVF (t = .697, p > .05) significantly, and accuracy in their mean scores even decreased in the posttest compared to the pretest. Regarding fluency, TC tasks improved Syll/M (t = −3.170, p < .05) substantially, whereas the pretest and posttest in DysFlu (t = −1.298, p > .05) were not significantly different and the mean scores of the dysfluency rates increased. The effect size of FreqT was large (over .8), the effect sizes in the five measures (MLT, MLU, DC/C, and Syll/M) were medium (over .5), and the effect size of C/AS was small (less than .5). Therefore, task sequences manipulated by task complexity significantly improved overall syntactic complexity and speed fluency, whereas mean accuracy scores decreased and dysfluency rates increased in the posttest.
Effects of task sequence manipulated by task complexity.
Notes. *p < .05. **p < .01. ***p < .001. Abbreviations: see Table 1 for full forms.
3 Gains between tests: Effects of task sequence manipulated by guided planning with vocabulary
The third research question concerned the effects of task sequences manipulated by guided planning with vocabulary in terms of syntactic complexity, accuracy, and fluency of English in L2 oral performance (see Table 7). Regarding syntactic complexity, sequencing GPV tasks improved overall syntactic complexity significantly in MLT (t = −2.252, p < .05), MLU (t = −2.715, p < .05), C/AS (t = −5.064, p < .001), DC/C (t = −5.211, p < .001), and FreqT (t = −2.915, p < .05) except MLC (t = .651, p > .05). Regarding accuracy, sequencing GPV tasks did not increase overall accuracy in E/AS (t = .540, p > .05), EFC/C (t = −2.075, p > .05), and CVF (t = −1.967, p > .05) significantly, but the accuracy in their mean scores increased in the posttest compared to the pretest. Regarding fluency, GPV tasks improved Syll/M (t = −3.418, p < .01) dramatically. However, DysFlu (t = 1.682, p > .05) did not differ significantly between the pretest and posttest, while the mean dysfluency rate scores decreased. The effect sizes in the two measures (C/AS and DC/C) were large, and the effect sizes in the four measures (MLT, MLU, FreqT, and Syll/M) were medium. Therefore, task sequences manipulated by guided planning with vocabulary improved overall syntactic complexity and speed fluency, improved mean accuracy scores, and decreased dysfluency rates in the posttest.
Effects of task sequence manipulated by guided planning with vocabulary.
Notes. *p < .05. **p < .01. ***p < .001. Abbreviations: see Table 1 for full forms.
4 Gains between tests: Effects of task sequence manipulated by guided planning with content
The fourth research question concerned the effects of task sequences manipulated by guided planning with content on syntactic complexity, accuracy, and fluency of English in L2 oral performance (see Table 8). Regarding syntactic complexity, sequencing GPC tasks improved partial syntactic complexity significantly in C/AS (t = −2.270, p < .05), DC/C (t = −2.758, p < .05), and FreqT (t = −6.898, p < .001), whereas GPC tasks did not change MLT (t = −1.431, p > .05), MLU (t = −1.692, p > .05), and MLC (t = −.091, p > .05). Regarding accuracy, sequencing GPC tasks did not increase overall accuracy in E/AS (t = .470, p > .05), EFC/C (t = −1.281, p > .05), and CVF (t = −.781, p > .05) significantly, but did increase accuracy in their mean scores in the posttest compared to the pretest. Regarding fluency, GPC tasks improved Syll/M (t = −2.375, p < .05) substantially. Interestingly, DysFlu (t = −2.136, p < .05) differed significantly between the pretest and posttest, but the dysfluency rate increased significantly. The effect size in FreqT was large, and the effect sizes in the four measures (C/AS, DC/C, Syll/M, and DysFlu) were small. Therefore, while task sequences manipulated by guided planning with content improved partial syntactic complexity and significantly improved speed fluency, GPC tasks increased dysfluency substantially. Sequencing GPC tasks did not make a significant difference in accuracy although participants’ mean scores increased.
Effects of task sequence manipulated by guided planning with content.
Notes. *p < .05. ***p < .001. Abbreviations: see Table 1 for full forms.
V Discussion and conclusions
The present study attempted to find the optimum task sequence for L2 novice learners in classrooms using deductive and inductive approaches. Three groups’ task sequences were manipulated by task complexity, guided planning with vocabulary, and guided planning with content, and then the CAF in the posttests of the three groups (between-groups) and from the pretests and posttests in each group (within-groups) were compared. Appendix 1 presents a summary of the findings in the present study. In relation to syntactic complexity, the TC group outperformed the GPV and the GPC group in increasing overall syntactic complexity. Both the TC and GPV tasks significantly increased in overall, subordinate, and specific syntactic complexity, while the GPC task significantly increased subordinate and specific syntactic complexity. In relation to accuracy, all three groups did not significantly improve both general and specific accuracy. The analysis showed little increase in mean accuracy in the GPV and GPC groups, but the TC group decreased in mean accuracy. In relation to fluency, the analysis indicated that the GPV group outperformed the GPC group in improving speed fluency although all three groups significantly improved speed fluency after performing the sequencing tasks. However, the dysfluency rate in the GPC group significantly increased, meaning that participants in this group repeated, restarted, and self-repaired their speech substantially more frequently in oral performance after completing the sequencing tasks. Although there was no significant difference between the pretests and posttests, the mean of dysfluency also increased in the TC group, and there was a slight decrease in the GPV group. Overall, for novice L2 learners in TBLT, both the sequencing TC and GPV tasks had the most positive effects in improving CAF (at six significant measures out of 11 measures in CAF) although the GPV tasks tended to be better even in mean accuracy and dysfluency rate scores than the TC tasks. However, the sequencing of the GPC tasks worked well (at four significant measures positively and one dysfluency measure negatively out of 11 measures), but no better than the GPV and TC tasks.
The TC task sequence, which followed the theoretical SSARC model of the task sequence along with resource-directing dimensions (± elements and ± reasoning) and resource-dispersing dimensions (± pretask planning), was helpful in improving syntactic complexity and speed fluency. This set of task sequences was easy to design because only invariant task features were considered without learner-dependent variables. However, the TC sequence may have less positive effects on accuracy and dysfluency rates at the same time because the mean accuracy of the TC group decreased compared to both the sequencing GPV and GPC tasks. The analysis showed a trade-off relationship between syntactic complexity and accuracy in the oral performance of the present study, a result that resembles the findings in Lambert and Robinson (2014). In their study, specific syntactic complexity (coordination and subordination rates) increased significantly, but specific accuracy did not improve between the pretests and posttests in each group and there was no significant difference between the two SSARC and Control groups in written performance. A difference between the present study and Lambert and Robinson’s (2014) study is task modality (sequencing oral task vs. sequencing written task) and L2 learners’ proficiency (novice vs. intermediate level). In Kim (2020), intermediate learners’ written performance showed a trade-off relationship between syntactic complexity and accuracy while advanced learners’ written performance showed a simultaneous improvement in two linguistic forms. As L2 learners’ proficiency levels increases, the attentional resources they allotted to two different forms could expand. However, intermediate (Kim, 2020; Lambert & Robinson, 2014) and particularly novice learners (the present study) had difficulty using resources fully, and thus a trade-off effect can be observed although the sequencing GPV tasks may result in simultaneous increases in syntactic complexity and accuracy despite improvements in mean accuracy in the mean scores.
While novice L2 learners struggled to select the appropriate forms in oral tasks, their dysfluency rates also increased. The example below shows one participant’s (Heejin) parts of speaking in the pretest and posttest in the TC group. [ ] indicates the dysfluent parts including false starts, repetitions, and self-corrections.
Example: Increase in dysfluency rates in the TC group Heejin (pretest): A man saw her and fall in love. (omission in middle) but [uh . . .] strange man steal the flower. Heejin (posttest): [When he]A when he saw the woman, [she falling, she got, uh . . . she falling, uh]B he falling love at a second. (omission in middle) but [the strange]C the another strange guy . . . [get off]D got off the bus and steal the flower.
As shown in the above example, Heejin made many errors and she used only simple sentences in the pretest while she tried to compose a complex sentence and present more detailed information in the posttest. At this point, she repeated what she said (in A) and self-corrected her speaking (in B, C, and D). Since the participants were mainly novice learners of English, they needed a number of trials with repetition, self-correction, or false starts to produce complicated and elaborate oral performances. Previous studies have also showed increases in speed fluency, but the dysfluency rates indicated their efforts to convey complicated and detailed information in which speed fluency was not detected. To sum up, combining the GPV and TC tasks, considering novice learners’ limited attentional resources and unstable interlanguage systems, the SSARC task sequence produced a positive effect. However, the task sequence combined with the guided planning of words performed surprisingly better in all CAF rather than the sequencing TC tasks.
Another interesting finding is that the sequencing TC tasks following the SSARC model outperformed the GPV and GPC groups significantly in improving overall syntactic complexity (MLT and MLU) between three groups although previous task sequence studies (Baralt, 2014; Lambert & Robinson, 2014) did not show significant differences between the groups. One explanation for the group difference may be the use of various measuring indices to detect different parts of syntactic complexity in the present study. The present study used six different syntactic complexity measuring indices, including overall, subordinate, phrasal, and specific ones (see Table 1), whereas previous studies mainly used the following specific syntactic complexity measures: past subjunctive expression (Baralt, 2014), coordination/subordination rates (Lambert & Robinson, 2014), or the occurrence of spatial expression (Levkina & Gilabert, 2014). For instance, Baralt (2014) claimed that the number of past subjunctive in Spanish was difficult to find due to infrequent usage by L2 learners. Thus, as Norris and Ortega (2009) argued, combining the use of both general and specific measuring indices can make investigations of different qualities and parts of linguistic production powerful and predictable.
The GPV task sequence was very helpful in improving syntactic complexity and fluency. Compared to L1 speakers who operate the speaking process only focusing on conceptualization and process both formulation and articulation automatically (Levelt, 1989), L2 speakers, especially those with low proficiency levels, need to execute conceptualization and formulation with more attention (De Bot, 1992), leading to nonparallel and slow processing (Kormos, 2006). Thus, L2 learners should decide ‘what they pay attention to when monitoring and these decisions most frequently involve prioritizing content over form, lexis or grammar, or vice versa’ (Kormos, 2006, p. 173) because novice L2 learners have limited working memory and attentional resources. The support of prioritized forms (i.e. key words), which were provided by the teacher in class, may compensate for their competitive oral performance between meaning (fluency) and forms (syntactic complexity and accuracy). As Figure 2 shows, when resource-directing dimensions increased from the simple tasks (Tasks 1 and 2) to the complex tasks (Tasks 3 and 4), and guided planning with vocabulary was introduced to make tasks procedurally easier in Tasks 1 and 3, they direct their attention toward conceptualizing and producing more syntactically complex and fast spoken performances. This finding aligns with Thompson (2014). His intermediate L2 learners showed an increase in syntactic complexity mean score in both sequenced guided planning and sequencing guided-unguided planning between the pretest and posttest, whereas novice L2 learners in the present study showed significant improvements in diverse aspects of syntactic complexity and speed fluency. His task sequence operation involved increasing ‘content and cognitive complexity’ (p. 131) at three different levels (number of intentional reasoning and obligatory number of relative clauses) and decreasing guided and/or unguided planning time at three different levels (from 10 minutes, 7 minutes, to 4 minutes). However, task sequence in the GPV group was manipulated by resource-directing dimensions (elements and reasoning) at two levels (from simple to complex) and by resource-dispersing dimensions combined with guided planning with vocabulary at two levels (guided planning to no planning). The only difference from the SSARC model was that instead of participants’ pretask planning by learners themselves, teachers provided pretask planning focused on vocabulary. The GPV tasks were modified versions of the SSARC model designed so that the teacher could give novice learners more linguistic support. This sequencing process might be much easier for learners and thus the novice learners might pay more attention to producing complex language instead of learning and familiarizing themselves with new task instructions.
In addition, although their accuracy did not improve statistically from the pretest to the posttest, their means increased only after four consecutive task sequences, which differ from the results for the TC group (decrease of accuracy in mean score). In Foster and Skehan (1999), teacher-led guided planning generated significant accuracy effects. The contrasting results in this study have several potential causes. First, in their study, the teacher explicitly taught the target forms (modals and conditionals) using example sentences, and the participants’ had intermediate proficiency levels. Meanwhile, the participants in the present study were novice learners. Moreover, the guided planning method for giving vocabulary was teacher’s spoken input delivered for a short time (less than 5 minutes), not by writing on the board or giving out a word sheet. For novice learners, the short time for teacher-led spoken input may not have been enough to dramatically increase their accuracy. Third, when looking carefully at the data, participant errors in the posttest were found in all areas, including past tense form (e.g. falled, was fell, fall, He falling love, and He stolen the flower), subject–verb agreement (e.g. cames), singular/plural noun (e.g. a flowers), to-infinitive form (e.g. He bought flowers to gave to her), pronoun (e.g. he for ‘she’ and her for ‘him’), and lexical choice (e.g. He took off the bus and behind for in front of). To increase their general accuracy, novice learners seemed to need more time to develop their current interlanguage systems. Lastly, participants devoted enormous effort to making as many complex sentences as possible by adding conjunctions (in formulation) and detailed information (in conceptualization). The mean length of syllables increased dramatically from 88.79 in the pretest to 105.38 in the posttest. Their limited attention may not have allowed them to pay much attention to accuracy despite the increase in mean score.
Regarding the sequencing GPC tasks, unlike guided planning with words, teacher-led content did not produce dramatically positive effects on overall CAF, although many novice learners requested teacher modeling in terms of the field notes. Before the experiment, I had anticipated that content support might reduce the burden in conceptualization, but this set of task sequences increased dysfluency, leading to more false starts, repetitions, and self-corrections. When listening to all the content, a quick listening itself (compared with reading) may not have been an easy way to acquire a story idea or participants might not have known which area they should pay attention to. Additionally, what they heard was not helpful in producing a positive effect on the different narrative tasks in the posttest. In Foster and Skehan (1999), the guided planning with content and the guided planning with language did not produce a difference in performance, while solitary planning improved complexity, fluency, and turn length. However, the teacher-led content might involve overly inconspicuous information for novice L2 learners. Therefore, content support with task sequences might not produce optimum sets of task sequences.
In terms of task design and TBLT for L2 novice learners, this study demonstrated the possibility of applying empirical methods to adjust and modify theoretical models to make them relevant to specific target learners and highlighted the feasibility of using task sequences following the SSARC model in a classroom setting. While the students had performed the tasks before the experiment, their behavioral characteristics were found to include heavy reliance on lexical items, writing a whole sentence before speaking, preference for teacher modeling for the tasks, using L1 between students, and frequent self-speaking before the tasks. These features signaled a need for additional linguistic support to engage in TBLT actively. In terms of these behavioral and inductive motivations, two modifying versions of task sequences were set and added to the experiment. In spite of the teacher’s observations, sequencing GPV tasks had very positive effects on all CAF, but sequencing GPC tasks did not generate the expected results. Therefore, an inductive model might have two different aspects: an advantage might be that if the teachers are skillful and experienced in TBLT, they could customize and optimize tasks to give learners more opportunities to develop their interlanguage systems; meanwhile, a disadvantage might be that task design is totally dependent on teachers’ decision-making and other factors, including learners and the tasks themselves, so novice or inexperienced teachers could not facilitate relevant TBLT classes. However, task sequences following the SSARC model produced a quite good effect on CAF, and the implicational advantage is that the task sequence is set by only one invariant criterion, task complexity (resource-directing and resource-dispersing dimensions), so it is much easier for most teachers to design and implement task sequences.
The present study had two limitations. First, it included no delayed posttests or qualitative learners’ data. Through these, the results of the study could be supported by the longer effects of task sequence on speaking and learner difference/perception towards task sequence in a classroom setting. Second, although the entire observation and task sequence experiment were conducted for a semester-long period, the actual task sequence was done during a single day of class. If the task sequences were extended over several weeks, it would show the long-term influence on L2 learning in class. To conclude, since the different effects of task sequence in the TBLT have not been sufficiently researched, more varied sequencing tasks should be studied and then implemented in the task-based language teaching instruction regardless of learners’ proficiency or any other implicational barriers.
Footnotes
Appendix
Summary of the findings in the present study.
| Measures | Between-G | Within the group |
|||
|---|---|---|---|---|---|
| TC: t (Sig.) | GPV: t (Sig.) | GPC: t (Sig.) | With-G | ||
| Syntactic complexity: | |||||
| MLT | TC > GPV or GPC | −2.765* (.014) | −2.252* (.040) | −1.431 (.172) | TC, GPV |
| MLU | −2.603* (.020) | −2.715* (.016) | −1.692 (.110) | TC, GPV | |
| C/AS | – | −2.295* (.037) | −5.064*** (.000) | −2.270* (.037) | GPV, TC, GPC |
| DC/C | – | −2.435* (.028) | −5.211*** (.000) | −2.758* (.014) | GPV, TC, GPC |
| MLC | – | −.1.642 (.121) | .651 (.525)▼ | −.091 (.929) | – |
| FreqT | – | −6.611*** (.000) | −2.915* (.011) | −6.898*** (.000) | TC, GPC, GPV |
| Accuracy: | |||||
| E/AS | – | −.1.605 (.129)▼ | .540 (.597) | .470 (.644) | – |
| EFC/C | – | .778 (.449)▼ | −2.075 (.056) | −.1.281 (.218) | – |
| CVF | – | .697 (.497)▼ | −1.967 (.068) | −.781 (.446) | – |
| Fluency: | |||||
| Syll/M | GPV > GPC | −3.170** (.006) | −3.418** (.004) | −2.375* (.030) | GPV, TC, GPC |
| DysFlu | – | −1.298 (.214)▼ | 1.682 (.113) | −2.136* (.049)▼ | GPC |
Notes. Between-G = Significant difference between the groups. With-G: Significant difference between the pre- and posttest within each group. > = The one is significantly higher than the other. − = no difference. ▼ = when comparing the mean score of the pretest and posttest, oral performance decreased. *p < .05. **p < .01. ***p < .001.
Acknowledgements
I would like to show my gratitude to Hossein Nassaji, María del Pilar García Mayo, and the anonymous Language Teaching Research reviewers for their insightful comments on an earlier version of this article. Any remaining errors are, of course, my own.
Correction Note (May 2023):
The Acknowledgement section has been updated with correct Journal Title.
Funding
The author disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Ministry of Education of the Republic of Korea and the National Research Foundation of Korea (NRF-2017S1A5B5A07059927).
