Abstract
Collaborative text reconstruction tasks such as dictogloss have been suggested as effective second language (L2) learning tasks that promote meaningful interaction between learners and their awareness of L2 target grammatical structures. However, it should be noted that the effect of pair interaction on the final product may differ depending on co-participant characteristics and particularly on proficiency disparities between partners. To date, most studies conducted on the effect of the different L2 proficiency of learners on paired performance have focused on the ways in which language learners interact, and the quantity and quality of language-related episodes (LREs) produced (Kim & McDonough, 2008; Leeser, 2004), often sidelining learners’ actual task performance. This study thus aims to investigate the extent to which partner L2 proficiency levels affect tangible language performance, particularly in terms of content accuracy in a dictogloss task. Results show large gains in idea units reproduced between first and second stages of the dictogloss across texts. However, while low-level students paired with high-level partners benefited most, this group also had the largest variation across the board and, overall, proficiency pairing did not systematically affect improvement in idea units. Idea unit analyses indicated that students tended to perform better on idea units from earlier parts of the text, and that some types of idea units were more discriminatory than others.
I Introduction
Pair work and small group work have become increasingly common in second language (L2) classrooms, in part because they allow learners to interact with each other simultaneously (Pica et al., 1993), resulting in peer mediation and negotiation of meaning, which are believed to facilitate second language acquisition (SLA) (Storch, 2002; Swain & Lapkin, 1998). Paired and group assessments are also gaining popularity in L2 testing contexts thanks to research showing that they can be more efficient and better aligned with common classroom practices than traditional language assessment formats (Taylor & Wigglesworth, 2009).
Among the communicative tasks in pairs or small groups in the L2 classrooms, dictogloss (Wajnryb, 1990), a collaborative text reconstruction task, has been suggested as an effective L2 learning task in promoting meaningful interaction between learners and their awareness of L2 target grammatical structures (Kowal & Swain, 1994; Reinders, 2009; Storch, 1999; Sullivan & Caplan, 2004; Swain & Lapkin, 2001). Dictogloss involves several stages of production and revision, including an initial individual listening task, a pair or group reconstruction of the text from the learners’ shared resources, and teacher feedback on their co-constructed work. Although dictogloss originally focused on improving learners’ understanding and use of grammar in the task of text reconstruction (Wajnryb, 1990), it has also been used as a listening comprehension activity to enhance learners’ understanding and retention of L2 aural input (Garcia & Asención, 2001; Prince, 2013; Vasiljevic, 2010; Wilson, 2003).
Dictogloss is argued to enhance listening ability by encouraging both bottom-up and top-down processing (Wilson, 2003). During the first phase of this activity, when learners listen to the spoken L2 input, they are likely to use top-down processing to get overall understanding of a text. In the second stage of dictogloss, they would mostly utilize bottom-up strategies by encoding what they heard into their own representations of the aural text, based on notes taken during the listening stage and in collaboration with their partners. The final phase of dictogloss helps learners to identify their listening difficulties when they compare their reconstructed text to the original aural text. This makes learners aware of their current ability to process aural input (Wilson, 2003).
As can be seen above, dictogloss can be an effective listening comprehension exercise. In addition, it could serve as a useful tool for formative assessment in L2 classroom settings (Vasiljevic, 2010), as recall tasks have been used as listening comprehension tests (Dunkel & Davis, 1994; Henning et al., 1983). Dictogloss can be implemented in multimodal language learning courses for simultaneous assessment and enhancement of students’ grammatical and listening comprehension skills. However, most of the previous studies on dictogloss have focused on the ways in which language learners interact, and on how much and what type of language-related episodes (LREs) have been produced (Kim & McDonough, 2008; Kowal & Swain, 1994; Leeser, 2004; Malmqvist, 2005; Sullivan & Caplan, 2004), often leaving learners’ linguistic products unanalysed. One recent study (Basterrechea & García Mayo, 2013) has investigated the impact of the number of LREs involving the target form generated during interaction on the actual learners’ written text reconstructions. They found that attention to use of the target form, the 3rd singular present tense (-s) morpheme in English, during collaboration led to accurate use of the 3rd singular forms in the reconstructed texts. This study shows that dictogloss tasks enhance both attention to a specific grammatical form and production of accurate linguistic features. Nevertheless, it is not yet known how collaboration during the dictogloss task is associated with the content accuracy produced in the co-reconstructed texts.
From a measurement perspective, it is imperative to document learners’ actual language performance, because test scores and grades are assigned and inferences about students’ language abilities are made not based on the quality of their interaction but rather on how well the students’ actual output matches the task’s criteria. Thus, further research is needed examining each individual learner’s tangible language performance, particularly by content accuracy, to determine dictogloss’ effectiveness as a listening comprehension test.
II Idea units
Content accuracy in a recall task reflects the degree of learners’ reading or listening comprehension skills, and is often measured by the total number of correctly recalled idea units (Carrell, 1985; Chang, 2006; Lee, 2007; Riley & Lee, 1996). Idea units have been shown to be a reliable measure of listening comprehension, with reliability coefficients sometimes exceeding those of multiple-choice tests (Winke & Jung, 2012). However, idea units have been difficult to define and are rarely properly addressed in the literature (Alderson, 2000; Winke & Jung, 2012), despite the fact that they have been widely used by L2 researchers and testers as a measure of comprehension. Carrell (1985) and Lee (2007) viewed idea units grammatical units, defining them as single clauses (main or subordinate, adverbial or relative clauses) and/or infinitival constructions, gerundives, or nominalized verb phrases. On the other hand, Johnson (1970) and Chang (2006) defined idea units as pausal units based on normally paced oral reading. Ellis and Barkhuizen (2005) defined an idea unit as ‘a message segment consisting of a topic and comment that is separated from contiguous units syntactically and/or intentionally’ (p. 154). Their criteria to define an idea unit as a full topic–comment unit focus on propositional or semantic content. Since learners’ reconstructions in a collaborative writing task may not match the structural properties of the source text, and may not be clearly divisible in terms of pausal units, units defined in terms of semantic content appear most appropriate. We thus used Ellis and Barkhuizen’s definition of idea units to analyse our data in detail.
III Position effects in free recall
It is important to understand which idea units are retrieved as well as how many idea units are correctly recalled in using dictogloss tasks as a measure of listening comprehension. Research in the roles of memory in the field of cognitive psychology has shown that items that appear at the beginning and end of a list are likely to be better remembered than those in the middle in free recall tasks (Kellogg, 2002). These are known as primacy and recency effects and have been replicated in memory research (i.e. Glanzer & Cunitz, 1966; Hasher, 1973; Murdock, 1962). The primacy effect refers to the successful retrieval of information located at the beginning, whereas the recency effect is related to a greater chance of recall for items appearing at the end. Recency effects compared to primacy effects are more pronounced for free and longer lists than for serial and shorter ones (Jahnke, 1965). If this holds true for dictogloss task performance, learners may recall more information appearing at the beginning and end of the text. Specifically, they are more likely to retrieve idea units from the end than those at the beginning given that dictogloss can be viewed as a free and long recall activity. However, such a serial position effect in free recall has been based on first language (L1) word lists, and it is not yet clear if they can be applied to L2 recall tasks, particularly those based on aural input. Furthermore, in L2 listening comprehension contexts, various structuring and organizing cues such as connectives and discourse markers can facilitate the comprehension of text (Chaudron & Richards, 1986; Dunkel, Henning, & Chaudron, 1993; Kintsch & Yarbrough, 1982). Regardless of the location of idea units, it is possible that content that contains signaling cues may stand a greater chance of recall.
IV Effects of characteristics of co-participants in L2 pair/group work
Since dictogloss involves pair work, learners’ performance on the task is necessarily affected by the partners’ various characteristics that they bring to the task. It is critical to understand how each partner’s individual performance compares to the paired performance, and what may account for those differences.
Examining the efficacy of paired or group work, the L2 teaching and assessment literature has compared learners’ performance on solo tasks with performance on corresponding tasks conducted in pairs or groups. Previous studies in the L2 oral testing context (Brooks, 2009; Shin, 2007) found that students tend to perform better in dialogic tasks than in monologic ones, as evidenced by their overall test scores. In a L2 writing testing context, Storch (2007) compared the grammatical accuracy of a text-editing task completed individually and in pairs, and analysed the nature of pair interaction. She found that students in pairs did not perform better on a grammar-focused task than those who worked individually, although pair work did seem to offer learners language-learning opportunities by allowing them to use and reflect on language use interactively. On the other hand, her other studies (Storch, 2005; Wigglesworth & Storch, 2009) revealed that pairs produced more grammatically accurate texts than students working alone did, indicating collaboration impacted grammatical accuracy positively in the joint text production tasks. These findings suggest that pair interaction was generally, but not always, facilitative and attracted attention to form (Storch, 1998, 2002).
However, it should be noted that a partner might bring in different factors that may affect a learner’s performance on paired tasks. Previous studies on this issue have shown that paired test performance seems to be qualitatively and quantitatively affected by a number of features of the participants, including their familiarity with their partner, gender, personality, and proficiency level in a paired or group work context (Taylor & Wigglesworth, 2009).
In terms of the effect of familiarity or acquaintanceship between participants on paired performance, Porter (1991) found no effect on scores for Arab learners examined by known and unknown interviewers. Similarly, Ockey, Koyama, and Setoguchi (2013) found that the interlocutor familiarity did not significantly affect group oral discussion test scores. On the other hand, a study by O’Sullivan (2002) on the effect of learner acquaintanceship on oral pair-task performance revealed that grammatical accuracy varied with both familiarity and gender. The results of her study showed that female Japanese learners working with a female friend scored higher than with a male stranger. Additionally, females were more accurate when working with a female stranger compared to a male stranger. Personality of learners also seems to interact with gender, as Berry (2007) found that females scored higher working with an introvert, but the opposite was true for male learners. In addition, personality effects tend to be more pronounced with assertive test takers than non-assertive test takers (Ockey, 2009).
V Effects of partners’ proficiency in L2 pair work
The effects of variability in L2 proficiency levels among participants on test scores have been similarly mixed in the extant research. Csépes (2002) found that proficiency gaps between participants in a paired speaking exam did not affect their rated speaking test scores. Davis (2009) also examined the influence of interlocutor proficiency on speaking performance in a paired oral assessment with a group of 20 Chinese university students who were learning English. Based on the t-tests and Rasch analysis results, he concluded that partners’ differing proficiency levels had no observable effect on speaking test scores, although lower-level learners produced more words when paired with a higher-level learner. Meanwhile, Iwashita (1998) and Shin (2007) observed score increases for lower-level students when they worked with a higher-proficiency partner. Higher proficiency learners have also benefited from working with less proficient partners, in that they achieved higher post-test scores after working with their lower proficiency partners (Watanabe & Swain, 2007).
These contradictory findings are also reflected in the nature of pair interactions in L2 teaching literature. Storch (2001) revealed that language proficiency differences between learners may not determine the degree of collaboration, because in her study the pair with the highest proficiency difference was the most collaborative in executing the task and making decisions about language. By the same token, Watanabe and Swain (2007) showed that proficiency differences did not affect the nature of peer assistance and L2 learning, but found that the core participants benefited more when working with lower proficiency partners than higher proficiency learners. Nonetheless, Kowal and Swain (1994) observed that learners in the more homogeneous pairs in terms of proficiency levels contributed equally to the discussion during a dictogloss task in French, suggesting that successful mediation or scaffolding would be difficult in pair work if the difference in proficiency between learners is too large. Most recently, Storch and Aldosari (2013) found that high proficiency learners tended to focus more on language use when paired with other high proficiency learners than when paired with lower proficiency partners on a joint composition task. However, their overall findings across different proficiency pairing types suggest that dyadic relationships formed during peer interaction confounded the effect of proficiency pairing on the focus on language use and the amount of L2 used as found in previous studies (e.g. Watanabe & Swain, 2007).
We are particularly interested in the influence of partners’ language proficiency on the performance of the pair because different partners’ proficiency levels are likely to affect the amount of mediation that each learner may receive during a dictogloss activity. However, little is known about the effects of L2 proficiency differences in pairs on test performance in an integrated, collaborative paired task, such as dictogloss. Furthermore, idea units have not received much attention in the previous L2 pair/group work research. To date, most prior research on L2 pair or group tasks has compared the performance of individuals to the performance of pairs in terms of their fluency, complexity, and accuracy, instead of examining development within pairs.
To address these points, we examined the effects of proficiency level of partner in a dictogloss task in terms of content accuracy operationalized by the number of correctly reconstructed idea units. An important gap in the current literature on pair work in SLA and testing is that paired performance is examined mostly in oral task contexts alone (Csépes, 2002; Davis, 2009; Iwashita, 1998; Shin, 2007). When research has been conducted on integrated skill tasks such as dictogloss, researchers have focused on the ways in which learners interact with each other in terms of quality and quantity of language-related episodes (LREs) (Storch, 1998; Storch, 2002; Storch & Aldosari, 2013), but not in terms of actual paired production. In addition, many previous studies of pair work have made the comparison of task performance between two different groups of students working in pairs and those working alone (Storch, 2005; Storch, 2007; Wigglesworth & Storch, 2009). Comparing separate groups has the advantage of allowing researchers to rule out practice effects. However, this design will not allow us to capture any systematic improvements as a function of proficiency differences in pairs across multiple stages of a task. Thus, this study used a mixed-method design in which the same students completed the equivalent tasks twice but with learners of varying proficiency levels.
As for our dependent variable, idea unit scores, previous studies have not agreed upon how to operationalize idea units for scoring (Winke & Jung, 2012), and it is not clear how improvement in content accuracy can be quantified reliably through using idea unit analysis. Moreover, little is known about how well learners perform across different types of idea units, and to what extent each idea unit discriminates low-and high-level learners in a collaborative text reconstruction task.
VI Research questions
We therefore devised the following three research questions:
Will students improve between stages of a collaborative text reconstruction task (dictogloss) in terms of content accuracy? (Put differently, will there be positive effects for peer mediation?)
Will the amount of improvement depend on the proficiency level of the partners?
Which idea units are most difficult, and which are most successful in discriminating student ability during a text reconstruction task?
For research question 1, we hypothesized that students in pairs would reproduce more idea units on their second drafts than they did alone on their first ones. Although we are unaware of other studies which have explored the issue of idea units in this type of task, other studies (Storch, 2005; Wigglesworth and Storch, 2009) have found a similar advantage for pairs over individuals in grammatical accuracy.
For research question 2, we hypothesized that lower level students would benefit most when they are paired with higher level partners, in line with the findings of previous studies (Iwashita, 1998; Shin, 2007) in language testing.
For research question 3, we predicted that students would recall the idea units appearing at the end of the text better because they would be more easily retained in their short-term memory, and the idea units followed by transition words such as ‘but’ and ‘however’ would have more recall because rhetorical signaling cues tend to help students process and comprehend incoming information (Kintsch & Yarbrough, 1982).
VII Method
1 Participants
This study involved 38 intermediate ESL learners enrolled in four separate, intact classes in a 7-level Intensive English Program (IEP) at a large Midwestern university in the USA. All participants were taken from level 5 grammar classes. Most students had resided in the USA for about 6 months at the time of the study. They had various first language backgrounds: Arabic (n = 17), Korean (n = 8), Chinese (n = 7), Spanish (n = 3), Portuguese (n = 2), and Japanese (n = 1); 9 were female and 29 were male. They ranged in age from 17 to 40, with an average age of 20.
Proficiency level was estimated using students’ total scores on an in-house IEP exam 1 administered at the end of each 8-week session, comprising reading, grammar, and listening sections with a total of 136 items. The coefficient α reliability of this test is quite high (α = .96), and a previous pilot study revealed that the total score on this exam was a significant predictor of dictogloss task performance. Each student was then classified into high (above the median score) and low (below the median score), and then paired into different groups. Homogeneous pairs were created by having students with consecutive rank paired together, whereas heterogeneous pairs were created by having the top-ranked student from the High group paired with the top-ranked student from the Low group, the second-ranked student from the High and Low groups, and so on. This ensured that the differences between partners remained relatively constant across all pairs. This created four possible conditions on an individual level: High students paired with another High student (HH), Low students paired with another Low student (LL), High students paired with a Low student (HL), and Low students paired with a High student (LH). Importantly, HL and LH pairs represent the same two people in terms of their paired productions, but these were kept separate in our analysis to examine whether one student in the pair benefited more or less than another. The mean difference of test scores between high and low was statistically significant (t = 7.28, p < .01).
2 Texts
Two texts of 120–125 words were constructed for topics relating to the learners’ current course materials (world populations and smoking regulations). The two texts are comparable in terms of reading levels (Flesch–Kincaid readabilities were 11.1 for both texts), number and length of idea units, and grammatical structures used. In order to make the tasks as authentic as possible, texts were created in order to highlight the use of gerunds and infinitives, which were the focus of class instruction at the time. Both texts were recorded at a natural speaking rate (160 words per minute) by a female native speaker of American English. Students completed a practice dictogloss task on a separate topic one week prior to this study to minimize any potential task familiarity effect on participants.
3 Procedure
Both dictogloss texts were administered in four intact IEP level 5 classes, the first text in the fourth week of classes, and the second the following week. A researcher first read out instructions from a script, then briefly introduced any potentially unfamiliar vocabulary words, and reminded students of the target grammar form (gerunds and infinitives). A recording was played once while students listened without taking notes. Recording was purposefully chosen instead of live narration because students were familiar with this type of listening task, and it enabled us to control for the quality and volume of human voice and to enhance comparability. Care was taken to ensure that classroom conditions were comparable and as free from outside noises as possible. Next, the researcher distributed worksheets and pens, and played the recording again, allowing students to take notes. Students were then given 10 minutes for text reconstruction, and were asked to reproduce as many of the ideas from the original text as they could. Once they were done with their individual text reconstruction task, students were given a different color of pen for the next stage of the dictogloss task.
Students were grouped into pairs by the researcher as described above. Each class of students completed two dictogloss tasks, and proficiency of the pairs was counterbalanced. For example, in week 4, classes A & B were paired homogeneously, and classes C & D were paired heterogeneously, and the opposite pairing was made the following week. Note that four classrooms had an odd number of students, resulting in a total of four students who could not be paired. In those instances, we solicited volunteers to perform the task individually, but instead of a partner, they were allowed to use any external resources they wished – grammar textbooks, dictionaries, and electronic devices – while the other students worked in pairs. Pairs were given a new worksheet and 15 more minutes to complete a paired reconstruction together. There were 15 pairs for the first text (5 HH, 4 LL, 6 HL/LH), and 13 pairs for the second text (4 HH, 4 LL, 5 HL/LH), as not every student was present for both sessions. Students (4 students for each text) who worked solo were also given 15 additional minutes to work on their reconstructions and were also permitted to access their grammar books and dictionaries, something that pairs were not permitted to do. Within pairs, both students were engaged in reconstructing the text, but, the final draft of a paired reconstruction was written by one of the partners, which is a common practice in a collaborative writing task in an L2 classroom (Storch, 2005; Wigglesworth & Storch, 2009). The jointly reconstructed text was the result of a combination of each single reconstruction and interaction in pairs. In-class observation and audio recordings confirmed that students typically shared their ideas cooperatively, and that some ideas from each member of the pair made their way into the resultant text. Once the second stage was completed, their drafts were collected. To capitalize on the learning potential of the task, the original text was then displayed using a document camera and compared with students’ paired reconstructions for analysis of content accuracy and grammatical correctness.
We assumed that any familiarity factor between students was to some extent controlled for because students had previously spent at least 20 hours per week for four weeks in class with each other and had been grouped with their classmates many times before. However, other potential factors such as gender and personality potentially affecting their paired performance in a dictogloss were not controlled for. For example, there were an unbalanced number of male and female students in each class. It should also thus be noted that students’ personality might have affected their performance in a dictogloss in ways that could not be identified in this study.
4 Analysis: Idea unit coding
As discussed above, we defined an idea unit as a full topic–comment (Ellis & Barkhuizen, 2005). In a pilot study, it was revealed that dichotomous scoring for idea units was sometimes impossible given the variability of student productions. We thus coded each idea unit on a 5-point scale (0, .25, .5, .75, 1). The semantic features corresponding to each of these 5 points were identified for each idea unit and agreed upon by all the researchers (see Appendix 1). In cases of simpler idea units containing fewer than 4 semantic features, one or more central features were assigned more weight prior to coding. Three researchers each coded 100% of the data, and inter-rater reliability was found to be high (α = .92; absolute agreement at > 80%). On all points of disagreement, consensus was reached through discussion.
In addition, we found that coding for the original text’s idea units did not account for all student production data. Thus, we came up with another variable, the ‘extraneous’ idea unit, referring to ideas and details that were not part of the original source texts. These included details that presumably came from students’ background knowledge, as well as personal commentary or opinions.
We did not categorize our idea units into ‘major’ and ‘minor’ idea units, as done by other reading-to-recall tasks (Carrell, 1985), because our oral texts were too dense and short to be hierarchically organized, and because students in our task were not creating their own original texts but rather reconstructing ideas from a predetermined source. However, we examined any change between solo and paired performance at the individual idea unit level to understand the relationship between the location of each idea unit and participants’ performances.
In sum, there were three dependent variables: total idea unit scores, scores by idea unit, and the number of extraneous idea units. Our independent variable was the partners’ level of proficiency determined by their placement test scores. The data were analysed in SPSS 20 (2011).
VIII Results
With regard to research question 1, as can be seen in Figure 1, paired-samples t-tests revealed that student pairs did produce significantly more idea units in their combined reconstructions than individual students had produced in their first-stage individual reconstructions on both texts (Text 1: t(29) = 4.33, p < .01, d = 0.79; Text 2: t(25) = 3.71, p < .01, d = 0.75).

Total idea units by text and stage.
It should be noted that there were successful repairs and additions when students worked together, as can be seen in a partially successful rephrasal for the same idea unit below (for the full original transcript, see Appendix 1):
Student 5108, Stage 1: ‘Ethiopia become double in the present.’ [Coded 0.25] Student 5105, Stage 1: ‘For example, Ethiopia is hivg [sic] double threaten to people.’ [Coded 0.25] (The students’ spellings are maintained.) Stage 2, Paired Reconstruction: ‘For example, Ethiopia is being doubling population, threaten to people.’ [Coded 0.50]
In brief, the question of whether student text reconstructions improve between stages of the dictogloss as a result of collaboration with a partner can be answered ‘yes’. However, soloists, who had had the same amount of time on task and the aid of dictionaries and electronic devices, showed no statistically significant improvement from their first to second reconstructions on either text (Text 1: t(3) = .52, p = .64; Text 2: t(3) = .73, p = .52).
Research question 2 was whether the amount of improvement will depend on the proficiency level of the partner. The results of a one-way ANOVA with pairing type as the grouping variable and idea unit scores as the dependent variable show that, at least in this experiment, no significant differences existed in the amount of improvement by student on either text (Text 1: F(3, 26) = 1.67, p = .20; Text 2: F(3, 22) = .66, p = .58). 2 Similarly, there was no significant effect of proficiency pairing on the idea unit scores in the jointly reconstructed texts (Text 1: F(3, 26) = .77, p = .52; Text 2: F(3, 22) = .35, p = .79). As shown in Figures 2 and 3, the low-level students paired with a higher partner tended to improve more than students in other pairings, but the amount of improvement observed for those students also varied the most among all pair types. Low-level students who were paired with other low-level students also improved consistently, but not significantly more than other pairs. In sum, student improvement varied significantly between pairs, but it did not vary systematically with respect to the proficiency level of the partner.

Mean differences of idea units (Text 1).

Mean differences of idea units (Text 2).
Regarding research question 3, not all parts of the texts were equally difficult or equally successful in discriminating between successful and unsuccessful reconstructions. By analysing each idea unit from the source text separately, we were able to obtain item difficulty and discrimination statistics for each idea unit (item discrimination was calculated as the point-biserial correlation between idea unit score and total score). Figures 4 and 5 show summaries of these data for each text. As the figures show, there was a general trend of idea units from the beginning of the text being reconstructed more accurately than idea units near the end of the text. This was surprising since it was initially predicted that parts of the text presented at the very end of the recording would be retained the best in students’ short-term memory. As can be seen in Figures 4 and 5, three idea units in each text stood out as particularly difficult: Idea Units 1-2, 1-4, 1-9 from Text 1, and 2-3, 2-6, and 2-10 from Text 2. We analysed these in further detail.

Idea units by stage (Text 1).

Idea units by stage (Text 2).
Idea Units 1-4 and 2-6 were very similar. Both contained large amounts of information and adjectival or other modifiers on almost every noun phrase. They were very dense compared to the other idea units, and most students thus neglected one or more of the details from the source text, lowering their scores. Idea Unit 2-3, on the other hand, had more serious problems. Within Text 2’s overall topic of smoking regulations in the USA, it was not very distinct:
Idea Unit 2-3: Many people in the U.S. believe that smoking should be discouraged.
Most students simply decided not to write this as a distinct phrase. Many students took down the word ‘discourage’ in their notes, but chose to incorporate it at an earlier or later point in their reconstructed texts, where it would be more salient. Overall, the idea unit did not contain a clear, distinct proposition, and as the only item to have a negative discrimination value, we decided that we would either remove or modify this idea unit considerably before reuse. Idea Units 1-2, 1-9, and 2-10 patterned together. When analysing the texts, it became clear that they were quite similar semantically.
Idea Unit 1-2: However, not all populations seem to be increasing. Idea Unit 1-9: However, similar policies have failed to change the birth rate in other countries so far. Idea Unit 2-10: However, some people argue that even these laws aren’t strong enough to prevent teenage tobacco use.
Each of these items qualify or question the central thesis or claims of the text. Many students instead seemed able only to reconstruct the central thesis of the text in its bare form, and the limits on the central claims were not reproduced by the students. The prominence of ‘however’ seemed to have no discernible effect. That said, although all of these yielded lower scores than the other idea units, their discrimination values were high; pairs who successfully reconstructed these idea units tended to be the best performers on the task as a whole. These items thus worked well in enabling us to determine which students were at the top of our group.
Finally, our analysis revealed that, despite dictogloss’ nature as a ‘text reconstruction’, many students – indeed, more than half of our students – produced at least some content that was not present in the original text. We found that these ‘extraneous’ idea units could be divided into two broad categories. First, there were additional details or information, presumably brought in from background knowledge and opinions. For example, on Text 1, one student writes:
Student 5401, Text 1: ‘Increasing population occure [sic] housing trouble, traffic, and so on.’
While the original text does mention the threat of serious problems, it does not elaborate on what those problems might be. Such attempts to make the student’s reconstruction more cohesive, especially through the addition of details, were commonplace, occurring in 34 different first-stage reconstructions. They were also frequently (24/34 instances) left in place by the partner in the second-stage. It is possible that these new details, which could plausibly have come from a text both students were now trying to recall after hearing only twice more than 10 minutes previously, were simply unnoticed, or even mistakenly thought to be from the source text, but it is also possible that the partners simply decided not to bring up the detail for other, social concerns such as maintaining positive face.
The second category of extraneous idea units was much different, however. This consisted of personal commentary or unsolicited opinions about the issue at hand. For example:
Student 5408, Text 1: ‘I think when the popuation [sic] grow up in the world, it is big problem.’
Despite the fact that the instructions clearly indicated that students were to reconstruct the original source text and only the source text, nine separate instances of personal commentary were found in the first-stage reconstructions. However, in eight of those nine instances, the commentary was not present in the second-stage reconstruction. The extraneous content was successfully removed during the partner interactions.
IX Discussion
Regarding research question 1, whether or not having a partner will help a student perform better on the second stage of a dictogloss task, all pairs reproduced significantly more idea units on their second drafts than they did when they worked alone, while most soloists did not. This suggests that the observed improvement is not just a function of time on task, since the soloists who were given the same amount of time and extra resources (which the paired students did not have) were unable to make use of it to improve their texts. This result suggests that neither additional time on task nor access to printed resources was as helpful as having a real human partner, but this finding should be interpreted cautiously, since the sample size for soloists was too small (n = 4) for each text.
In terms of grammatical accuracy, similar results were found in other studies comparing individual and pair work on writing and grammar tasks (Storch, 1999; Storch, 2005; Wigglesworth & Storch, 2009), although little is yet known about how collaboration in pair work would show some advantage for text comprehension. While pairs were not often successful in removing extraneous details taken from background knowledge, they did successfully remove extraneous personal commentary or opinions from their second drafts. One possible interpretation of this result is that personal commentary was more easily identified as coming from outside the source text and therefore off-task, and probably less face-threatening to remove than information or whole propositions that supported the cohesiveness of the text. In sum, these findings seem to suggest that having a pair is helpful when students work on a dictogloss task in terms of reproducing more relevant information from the source text.
Research question 2 concerned whether or not proficiency levels of partners would affect how well students reproduced idea units in a dictogloss task. The overall results show that there was no systematic effect of proficiency level of the partner on gains in idea units. This is consistent with previous studies in a paired oral testing context, which found that proficiency gaps in pairs did not necessarily lead to observable test score differences (Csépes, 2002; Davis, 2009). On the other hand, the general trend was that low-level students benefited more from the collaboration regardless of partner. Particularly, low-level students paired with higher-level partners benefited most but with the largest variation. This aligns with the findings of previous research that examined the effect of partners’ relatively different L2 proficiency levels in a dictogloss task on the number of LREs produced (Kim & McDonough, 2008; Leeser, 2004). Their results revealed that learners produced more LREs when interacting with higher proficiency learners rather than with lower-level ones. Similarly, Davis (2009) found that when working with more advanced peers, lower-level test takers produced more words, although this did not translate to their rated test scores.
Our results indicate that there was no significant difference in students’ gains in idea units across the different pairings formed by their relative L2 proficiency levels. However, we found that there was large variation among students when they worked in pairs. Such variation might be due to other factors such as students’ gender, personality, and attitudes to collaborative activity that we did not control for in this study. This may explain why some high-level students performed even worse when working in pairs. It is also possible that each student may have different attitudes to pair or group work, and students may not embrace collaborative activities as much as teachers might expect (Storch, 2011). Different personalities also may change the nature of pair work, and not all students work in a collaborative manner. This may thus lead to dominant-passive relationships in which not all participants contribute to their pair work equally (Storch, 2002). All these potential factors in a pair work context are likely to be confounded with variability in proficiency levels among students. Unfortunately, these were not identifiable in this study, and should be addressed in future research.
One might also argue that our proficiency gaps between high- and low-level students were not large enough because they were from the same level classes even though their test scores were significantly different. However, we should note that successful collaboration in pair work would not be possible if proficiency differences were too large (Watanabe & Swain, 2007), and we rarely pair students across different class levels in actual classroom contexts.
Research question 3 addressed what type of idea units might help us to best discriminate between students’ ability levels. We found first that idea units towards the beginning of the text seemed to be easier for students to reproduce than idea units at the end. This went against our prediction that recency effects would be stronger than primacy effects on dictogloss performance. While a weak recency effect may be attributed to the fact that students’ reconstructed texts were often based primarily on the notes that they took during the dictogloss activity, this trend was observed even for students who took very few notes in total. Instead, it might be more likely that the real-time, online nature of comprehension of aural texts might leave students little time to follow the subsequent information, which may be related to learners’ limited cognitive processing capacity (Just & Carpenter, 1992; Zwaan & Brown, 1996). Their potential attention failure during perceptual processing might make it difficult for them to retain subsequent information (Goh, 2000).
We also found that idea units that contained attribution or qualification of previous idea units tended to be the most difficult as indicated by extremely low item facility values. The results did not support our prediction that the presence of salient discourse markers such as ‘however’ would facilitate recall for all students. This may echo Dunkel and Davis’s (1994) findings that rhetorically signaling cues on a recall task may not always be helpful. Nevertheless, these units were still useful in that they had high discriminative power; students who correctly recalled these idea units performed best on the task overall. On the other hand, dense idea units containing more adjectival modification and especially idea units that were less distinct from each other were much harder, and did not always discriminate well, indicating their relatively limited usefulness.
The findings of the present study, along with Leeser (2004) and Swain and Lapkin (2008), suggest that dictogloss tasks place high demand on students’ listening comprehension skills; most students either individually or in pairs could not reproduce more than 50% of the entire idea units from the texts. Thus, future research might want to consider a slower speech rate for the recording, and shorter, more cognitively simple texts. While we chose to use expository texts in order to mimic the textbook materials students were accustomed to, a narrative text might be easier for learners to organize and retell.
In conclusion, our study revealed that students working together in pairs assisted each other, thereby recalling more correct idea units from the texts and eliminating extraneous information in their rewrites. However, within the context of this study, the relative L2 proficiency differences in pairs had little influence on improvement in idea-unit recall. It appears that lower-level students tended to benefit more from collaboration, but with considerable variation across the board. Future research is needed to determine why this is the case. Dictogloss can be used as a useful collaborative assessment tool enhancing interaction and an awareness of form-related issues. However, the effects of peer mediation appear not to be systematic according to proficiency, and thus its appropriateness as a summative assessment tool is limited. In terms of research on integrated task performance, our study benefited from an analysis of participants’ recall of idea units on an individual level. Using polytomous coding instead of dichotomous coding for the recall task provided a fine-grained indication of learners’ levels of comprehension of a given text (Winke & Jung, 2012). Our results of the effects of different types of idea units could shed light on how L2 learners process given input for recall not only on a dictogloss task, but also in recall tasks in general.
Footnotes
Appendix
| Idea units original text 1: | Necessary features: | |
|---|---|---|
| 1 | In 2011, the United Nations reported that the world’s population had reached 7 billion people | +2011, +UN report, +world population, +7 billion (+past) |
| 2 | However, not all populations seem to be increasing | +Qualification, +exceptions, +are not increasing (past, if couched in report, also okay), +in that report / appearance |
| 3 | According to the study, most of the population increase is coming from Africa | +following based on study, +most of the increase, +from Africa |
| 4 | In several poor countries, larger populations threaten to cause major problems | +poor countries (pl.) are a group, +growth as source of problem, +threat (i.e. not actually causing now), +cause problems, +severity |
| 5 | For example, Ethiopia’s population is expected to double in the next 50 years | +example of (4), +Ethiopia as TOP, +projection/probability, +population doubling, +in 50 year timeframe |
| 6 | On the other hand, many rich countries have a different problem | +contrast to (3–5), +rich countries as group, +many (but not all) of them, +rich also have problem |
| 7 | In countries like Italy, the population has already started falling | +Italy as example, +population (not economy), +decrease, +started (recently) |
| 8 | To fix this problem, the Italian government plans to increase the number of new births […] and it is thinking about giving extra money to larger families | +Counteraction / to encourage more childbirth, +Italian government as actor, +plan/thought/proposal (not currently enacted), +money to families |
| 9 | However, similar policies have failed to change the birth rate in other countries so far | +Qualification, +other countries have done similarly, +no change in birthrate (+despite money) |
|
|
||
| Idea units original text 2: | Necessary features: | |
|
|
||
| 1 | In the United States, there are many regulations about cigarette advertising | +tobacco regulations, +spec. advertising, +located US, +many |
| 2 | (Because) tobacco has been linked to many health problems | +tobacco/smoking SUB, +connected to problems, +health problems, +completed action with present relevance (+causal relation with (3)) |
| 3 | (so) many people in the U.S. believe that smoking should be discouraged | +Americans as SUB, +many but not all, +discourage smoking, +as belief /conviction/opinion (not actively discouraging themselves) (+result (2)) |
| 4 | In particular, there are strict penalties for tobacco companies | +penalties exist, +severity, +for companies, +especially so for them (must also be in present)* |
| 5 | whose advertisements encourage young people to smoke | +advertisements (not people) as SUB, +smoking encouragement, +condition for penalties, +youth as target |
| 6 | Supporters of these penalties argue that teenagers are too young to make important health decisions | +supporters (and not all) agree with the penalties, +youth as target of the regulations, +health decisions link to tobacco, +too young as reason, +decisions are important |
| 7 | Some scientific studies also suggest that teenagers are more likely to become addicted | +studies support [COMMENT], +young people, +higher probability (but not certainty), +addiction |
| 8 | if they start to smoke at a younger age | +condition for (7), +starting age is key, +young time is bad |
| 9 | As a result, laws were passed to prevent actors from using cigarettes in children’s movies or in commercials on television | +therefore, +legislation or government action, +cigarette ban, +movies and commercials, +spec. children’s media |
| 10 | However, some people argue that even these laws aren’t strong enough to prevent teenage tobacco use | +however, +not all agree, +laws should be stronger, +teenage smoking continues |
Note. *There was only one case where the student used a future expression, but this was marked off.
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
