Abstract
Peer feedback is widely used in second and foreign language writing contexts. While second language (L2) proficiency is likely to be an important factor in determining peers’ ability to give and utilize feedback, its contribution has been relatively under-researched. In the present study, 54 undergraduates in a foreign language writing context gave and received feedback on two different texts. The quantity and type of feedback given and incorporated were analysed, looking at whether these changed or preserved meaning. Generalized linear mixed models were used to assess whether the L2 proficiency of the reviewer (reviewer proficiency) and writer (writer proficiency) in each dyad determined the quantity and type of feedback given and incorporated. Results showed that reviewer proficiency significantly predicted the number of suggestions made. Writer proficiency did not significantly predict the number of suggestions, though lower proficiency writers incorporated significantly fewer meaning-related suggestions into their revised texts than higher proficiency writers. Differences in giving and incorporating suggestions also emerged for different pairings (i.e. matched or mixed proficiency), though these were not significant. The present findings provide further insight into understanding how L2 proficiency modulates the peer feedback process.
I Introduction
Mixed level classes are the norm in many second and foreign language writing classrooms. In such environments, educators aim to maximize learning for all students through differentiation in terms of feedback and support, and by using tasks and activities that can accommodate various levels. Peer feedback activities are widely used in writing classes not only because they have been shown to provide a range of benefits to the learner (Liu & Hansen, 2002), but also because they can be used with students who vary in their language proficiency (e.g. Mendonça & Johnson, 1994; Nelson & Murphy, 1992; Suzuki, 2008). It is therefore somewhat surprising that although second language (L2) proficiency is likely to be an important factor in the process of peer feedback, no studies to date have treated it as the primary variable of interest in investigations concerning the peer revision process. The present study thus addresses this gap in the literature by conducting a quantitative analysis of peer feedback dyads and by examining the impact of reviewer and writer proficiency upon the peer feedback process.
1 Peer feedback
Research into the effectiveness of peer feedback has been overwhelmingly consistent in detailing the validity of the activity in terms of assisting writers in developing their writing (e.g. Hyland and Hyland, 2006; Liu and Hansen, 2002). The main findings from this research are that with appropriate training, peer feedback is likely to lead to the development of social skills, cognitive skills, and meta-cognitive strategies (De Guerrero & Villamil, 2000; Mendonça & Johnson, 1994; Min, 2005; Suzuki, 2008; Villamil & de Guerrero, 1996), improved text quality (Berg, 1999; Suzuki, 2008) and writing ability (Lundstrom & Baker, 2009).
Socio-cultural theory has been used as a framework for investigating peer feedback interactions (Villamil & de Guerrero, 1996) and is particularly appropriate because it stresses the important role of social interaction, language and other signs in the process of individual development. Vygotsky (1978, 1986) proposed the Zone of Proximal Development (ZPD) to conceptualize how novice learners can be facilitated through interaction with an expert. The ZPD is the difference between what a novice learner can achieve by herself and what she can accomplish with the assistance of a knowledgeable other and/or other meditational tool (e.g. Lantolf, 2000, p. 17; Wertsch, 1991). Although much of socio-cultural theory is based on the idea of novices interacting with experts in order to achieve higher learning, research has also used the ZPD construct with peers where there is no obvious expert in the dyad/group (Ohta, 2000). In this case, peers bring with them different knowledge and skills, reflecting their language and learning histories. Research has shown how differential competences of peers can facilitate learning in pairs (Ohta, 1995; Anton & DiCamilla, 1998) and groups (Donato, 1994; see Ohta, 2000). In their article, Villamil and de Guerrero (1996, p. 69) suggested that ‘dyadic peer revision offers an opportunity for bilateral, rather than unilateral, participation and learning; in other words, both peers may give and receive help, both peers may ‘teach’ and learn how to revise’.
Although peer feedback activities may provide benefits for both peers regardless of their individual L2 proficiencies, what is actually ‘given’ and ‘received’ may differ considerably. In other words, while interacting with a peer in a task in a second language is likely to lead to a variety of rewards for both peers, it is unclear exactly what these rewards are and how the participants’ L2 proficiency may influence their distribution. Ferris (2006, 2010) notes that individual variation in responsiveness to feedback has received little attention, even though L2 proficiency is often cited as being potentially highly important for the feedback process (Berg, 1999; Connor & Asenavage, 1994; Liu & Hansen, 2002).
2 Peer feedback and L2 proficiency
Before addressing the issue of how L2 proficiency may interact with the peer feedback process, it is necessary to clarify our definitions. First, we consider the peer feedback process as both the giving of feedback and the revisions that are made on the basis of this. Second, we consider L2 proficiency to cover all knowledge of L2 and skills in the four key areas of speaking, writing, reading and listening. Importantly, peer feedback involves writing then rewriting a text, reading a peer’s text and commenting on various aspects of that text. Thus, when discussing the process of peer feedback a single measure (e.g. writing proficiency) may be less appropriate than a generic measure of L2 proficiency (i.e. that deals with both productive and receptive language knowledge). We return to this issue in Section II when we discuss the measure of L2 proficiency used in the study.
A number of studies have investigated the text-related negotiations that occur between peers during feedback sessions, while controlling for proficiency by using roughly same-level participants (Mendonça & Johnson, 1994; Nelson & Murphy, 1992; Suzuki, 2008). Negotiations are defined as the modification and restructuring of interactions between learners that occur as a result of difficulties in message comprehensibility (Pica, 1994; cf. Swain, 2000). By doing a cross-study comparison, it is possible to assess the frequency of such negotiations and the possible influences of L2 proficiency. Mendonça and Johnson (1994) looked at negotiations made during peer feedback with advanced ESL students. They found that not only did peers engage in text-focused negotiations, but these also led to revisions around half of the time. Thus, while learners were selective about which comments they incorporated, peer feedback was incorporated into their revisions. In another study, Suzuki (2008) recorded intermediate EFL learners’ peer feedback sessions and noted that the peers in her study engaged in fewer negotiations than in Mendonça and Johnson’s (28 vs. 16.7 times per 15-minute session; p. 228). This suggests that more proficient learners make more suggestions regarding the various aspects of peers’ texts. In another study, Nelson and Murphy (1992) studied low-proficiency ESL learners’ peer response behaviours, noting that while the participants successfully engaged in providing feedback, they also required additional support in terms of social intercultural skills, knowledge of writing concepts (e.g. topic sentences) and skills, and language for responding during peer response. Taken together, the results of these three studies suggest an important role of L2 proficiency in the process of peer response, with increased L2 proficiency potentially leading to more negotiations and engagement.
A study by Lundstrom and Baker (2009) investigated whether participants in peer feedback gained from the process in terms of improved writing ability. They conducted a study of 91 ESL students at two proficiency levels who participated in peer feedback. The participants were put in one of two conditions: they either gave feedback on their peers’ writing (and received none on their own writing) or they received feedback (but gave none on their peers’). Writing ability was measured in pre- and post-tests at the beginning and end of the semester using a grading scheme that included organization, development, coherence/cohesion, structure, vocabulary and mechanics. The primary finding was that those in the ‘giving’ group made significantly greater gains in writing ability than those in the ‘receiving’ group. If we consider this result in light of the previous research discussed above, higher proficiency learners should be able to give more feedback, and thus should benefit more from the peer revision process than lower proficiency peers, who are less able to give feedback.
Given the limited knowledge on how L2 proficiency actually impacts the feedback process, we set out with the following primary research question: how does reviewer proficiency (L2 proficiency of the reviewer) and writer proficiency (L2 proficiency of the writer) influence the peer feedback process? The giving of feedback was measured by the number and type of suggestions made by reviewers, and the outcome of this was measured by the number and type of suggestions incorporated by writers. First, the effects of reviewer/writer proficiency were considered as main effects; in other words, did these proficiencies influence the giving and incorporating of feedback regardless of the proficiency of the other member of the dyad? Second, interactions between reviewer and writer proficiency were considered to see whether the members’ proficiencies in dyads that were mixed or matched (e.g. high–low or low–low) further influenced the feedback given/incorporated.
Here, we thus focus not only on the feedback provided by the reviewer but also whether this is incorporated by the writer, and this is done for both peers in all dyads. Unlike the highly controlled study of Lundstrom and Baker (2009), we were primarily interested in how L2 proficiency influences the peer feedback process in actual mixed-level writing classroom situations, that is, with both peers giving and receiving feedback, and with proficiency varying across participants. By analysing the quantity and type of feedback provided and how this is used by learners, we were able to make inferences about how proficiency may impact individual learning. We were also able to make suggestions for pairing students in peer feedback activities (i.e. mixed or matched proficiency).
II Materials and Methods
1 Teaching and learning context
The context of the present study is a compulsory first-year undergraduate English writing course at a high-level Japanese university. The primary aims of the course are to introduce learners to the genre of written academic papers and to register-specific features of language, as well as other important elements of English for academic purposes, such as finding sources, citations and referencing, and using supporting evidence and argumentation in extended pieces of writing.
2 Peer feedback training and procedure
Training in peer revision was provided in the form of an instructional DVD (Middleton et al., 2009) that provided a model of the process of peer feedback, including a variety of suggestions that dealt with issues of content, citations, organization, grammar and academic register. In addition, awareness-raising activities focusing on aspects such as register and errors in formatting, and an exercise in which students critique each other’s research was conducted in previous sessions. These pair and group activities provided practice in cooperative learning and were designed to give students confidence in making and receiving suggestions from peers. In addition, they were taught about the basic structure and features of individual sections of science reports prior to writing their own. Prior to beginning each peer feedback session, students were briefly reminded to focus on the following features: content, structure, language and formatting.
Importantly, peer feedback dyads were self-initiated (i.e. students decided who they would sit next to in class and this formed the dyad) and sessions took around 40 minutes. During this time, peers were instructed to swap papers, read the paper once, then read again and make comments. Peers typically began discussing after 15 minutes. Peers could use either English or Japanese for comments and discussion. Students used either a laptop computer with a touch-screen pen, or an iPad with an annotating application for marking up peers’ papers (PockeySoft, 2013).
3 Participants
Fifty-four first-year undergraduates of the University of Tokyo who submitted all texts participated in the study. In addition, eight students did not submit all necessary texts and were thus excluded as writers from the study; however, they did review texts (of participants who submitted all necessary texts), and thus were included as reviewers only. All participants were native speakers of Japanese and had studied English for an average of six years.
4 Participants’ proficiency
For the present study we needed a generalized measure of L2 proficiency that provided an indication of the participants’ L2 receptive and productive skills, and would allow us to categorize the participants into high and low proficiency groups. To achieve this aim we measured participants’ language proficiency by administering a C-test that had been used in previous classroom-based research (Gilmore, 2007). The C-test is similar to a cloze test but the second half of every second word is deleted. This method of deleting items may be considered more objective than traditional cloze tests, in which the researcher decides which words to delete and in what frequency (Dörnyei & Katona, 1992; Klein-Braley & Raatz, 1984). Past research has shown that the C-test is a valid measure of L2 proficiency (Dörnyei & Katona, 1992; Connelly, 1997; Negishi, 1987) as to successfully complete it, the participant must have a broad collocational and lexico-grammatical knowledge of the patterning of English vocabulary in a variety of contexts. The mean percentage accuracy for the 54 writers was 70.0% (SD = 13.0%), which is roughly an intermediate level of English proficiency. The group was divided into two proficiency groups (high, low) using the median score: high proficiency (29 participants, M = 79.7%, SD = 6.0%) and low proficiency (25 participants, M = 59.4%, SD = 10.0%). The difference between the groups was confirmed using a two samples t-test (t(52) = −9.23, p < .000).
5 Materials
The writing component of the course assessment was a research paper worth 50% of the course grade. This research paper was divided into four sections (Introduction, Method, Results and Discussion) and the first two were used in the present research. The total number of texts collected was 324: 54 initial, 54 annotated (reviewed) and 54 revised drafts for both Introduction and Method sections. Mean individual text length was 95.4 words (SD = 27.5), which while being reasonably short is still representative of the type of writing which L2 learners often receive feedback on in the classroom (i.e. paragraphs or unfinished essay drafts). The reason the sections were quite short is due to the nature of the course: writers wrote one section per week meaning writing time was limited, and the final report was intended to be short (on average around 650 words), because this was for most of the students their first experience writing a full report on a scientific experiment. Topics varied and covered various scientific disciplines including biology, physics, psychology, chemistry and engineering; however, all followed the same format in terms of the Introduction and Method sections.
6 Coding scheme and procedure
Suggested revisions were coded using Faigley and Witte’s (1981) coding scheme. This coding scheme, which has been used in previous studies on the feedback process in L2 writing (Connor & Asenavage, 1994; Paulus, 1999; Phinney & Khouri, 1993), is useful because it distinguishes between two levels: surface changes and text-based changes. The former includes suggested revisions that do not change the meaning of the text and the latter includes those that do. In turn, these two levels are divided into four categories: formal changes (e.g. spelling, tense), meaning-preserving changes (e.g. word choice, active to passive sentence changes) and meaning-related changes (e.g. content and rhetoric-related revisions) at a microstructure level and meaning-related changes at a macrostructure level. We preserved the general distinction between surface and text-related changes, and the distinction between formal and meaning-preserving changes within surface revisions, but we did not distinguish between micro- and macro-structure levels as this level of detail was unnecessary for the present analysis. The distinction between formal and meaning-preserving changes was maintained because formal changes refer solely to grammatical issues, whereas meaning-preserving changes cover a wider variety of suggested revisions that include register-related changes. As the students in the present study had recently learned the basics of formal academic writing style (e.g. preference for however instead of but in sentence initial position and limited use of personal pronouns), we expected a large number of register-related suggestions in the feedback. From experience, we also expected student reviewers to focus on grammatical changes even though they were instructed to focus on multiple features of the text (Table 1; information about annotations and examples of revisions are provided in Appendix 1).
Coding scheme for suggestions.
Two researchers (the present authors) compared initial and revised drafts, identified revisions, and coded them according to the above classifications. Prior to coding the full set of texts, the two researchers practiced coding a number of texts and discussed discrepancies until reaching agreement. Then, the researchers coded all texts individually. After coding, the suggestion counts classified by type were recorded for each researcher, and interrater reliability was calculated. The Kappa statistic was calculated as a measure of ratings similarity while accounting for the possibility of chance agreement, and gave a score of 0.80. All discrepancies were discussed until 100% agreement was attained.
III Results
We first present an analysis of the data in terms of the number of suggestions made and incorporated and the proportion of each type of suggestion. Then, we use statistical measures to assess (1) whether reviewer and writer proficiency predicted the number of suggestions made and incorporated, and (2) whether the interactions between pairings (e.g. high–low) resulted in different patterns of suggestions made and incorporated.
1 Proportion of suggestions made and incorporated
The mean number of suggestions given and incorporated in each category, as well as the proportions, are shown in Table 2. The percentage of suggestions were as follows: 32% formal, 38% meaning-preserving, and 27% meaning related, with 3% being unclassifiable. The combined total of suggestions that were not directly related to meaning (formal plus meaning-preserving) was 70%, indicating that while students did focus on meaning/content related issues, the majority of suggestions were related to surface level issues.
Descriptive data for suggestions and suggestions incorporated.
Notes. The mean (M) number of suggestions given and incorporated and the standard deviation (SD) are provided in parentheses. The percentage of total is provided (%) and %* refers to the percentage of the total number of suggestions that were incorporated in the revised drafts.
The percentage of suggestions that were actually incorporated into the revised drafts was 68% for formal revisions, 57% for meaning-preserving revisions and 56% for meaning-related revisions (Table 2). Suggestions regarding low-level formal revisions (including mechanics, formatting, tense, plurals and subject-verb agreement errors) were incorporated the most. The average across the three types was 61%, revealing that writers acted upon over half of the actual suggestions. This indicates that writers appeared to evaluate suggestions and did not simply incorporate all of their peers’ suggestions.
2 Overview of statistical analyses
For each text we investigated the number of each type of suggestions made and incorporated in each peer feedback dyad, while considering the influence of the writer’s and the reviewer’s proficiency. Because the proficiencies were either high or low, this led to four combinations (high–high, low–low, high–low, low–high; for the descriptive statistics of these groups, see Table 3). When participants were in the same proficiency grouping (i.e. high–high, low–low) we considered them to be roughly ‘matched’ in proficiency, and when participants’ proficiencies differed (i.e. high–low, low–high) they formed ‘mixed’ proficiency pairings. For the statistical analyses, the number of suggestions made and incorporated were the dependent variables and were count (not continuous) data. Three counts were used for each dependent variable, that is for formal, meaning-preserving and meaning-related suggestions. For the independent variables, we included suggestion type as a 3-level factor (formal, meaning-preserving, meaning-related), and reviewer L2 proficiency and writer L2 proficiency each as 2-level factors (low, high). In addition, we included all two-way interactions between these three variables (e.g. reviewer proficiency and writer proficiency). Participants were classed as a random variable, which was included to account for variation that was attributable to individuals. The design was multi-level with suggestion type nested within participants (because for each participant there were counts for each type of suggestion).
Mean number of suggestions given and incorporated by reviewer and writer proficiency respective to peers’ proficiency.
Notes. Standard deviations are given in parentheses.
The distributions of the dependent measures showed that zero counts were frequent in the data. This meant that statistical tests such as Analysis of Variance (ANOVA) that require a normal distribution were not appropriate. Goodness of fit tests, which compare the data to standard distributions including Poisson, binomial and negative binomial distributions, revealed that a negative binomial distribution was the most appropriate for the data. Using this distribution, we selected generalized linear mixed models as these are suitable for designs that include dependent variables that are non-normally distributed, multiple independent variables and random variables (Baayen, 2008; Crawley, 2005). The analyses were conducted using the glmmADMB (Skaug, Fournier, & Nielsen, 2006) package in R open source software (R Development Core Team, 2010). Table 4 provides the following information for each independent variable: an Estimate of the parameter (independent variable or interaction), with larger figures demonstrating greater effects, as well as the Standard Error, Z-score, and P-value for each variable.
Model for suggestions given and incorporated including all main effects and any significant interactions.
Notes. *p < .05, **p < .01, ***p < .001.
3 Statistical analyses
In the analysis of the number of suggestions made, reviewer proficiency was a significant predictor (p < .05), showing that reviewers that were of high proficiency made significantly more suggestions than those who were of low proficiency. Suggestion type and writer proficiency were not significant (p > .05). None of the interactions were significant (p > .05) and are thus not presented. However, though not significant, the interaction between writer and reviewer proficiency on the number of suggestions made is informative. When writer and reviewer proficiencies were matched (high–high and low–low), the number of suggestions made was almost equivalent (mean = 11.1, 10.9, respectively). In contrast, when proficiencies were different (high–low, low–high), the number of suggestions made differed more strikingly: when high proficiency reviewers worked with low proficiency writers’ texts they tended to make the most suggestions (13.9) but when low proficiency reviewers worked with high proficiency writers’ texts, they made the least (9.3).
In the analysis of the number of suggestions incorporated into revised versions, reviewer proficiency was not significant, but more suggestions were incorporated when reviewers were higher proficiency than when they were lower proficiency (6.3 vs. 5.3, respectively). Although writer proficiency and suggestion type were not significant as main effects, the interaction between these variables was significant (p < .05). The interaction shows that while formal and meaning-preserving suggestions were incorporated to a similar degree regardless of writer proficiency, meaning-related revisions, which were the least incorporated type of suggestion, were incorporated more by higher proficiency writers compared to lower proficiency writers (p < .05). This is partly in line with previous research that has shown lower proficiency learners make fewer meaning-related revisions (Berg, 1999; Paulus, 1999).
IV Discussion
Our results show that reviewer proficiency strongly influences the number of suggestions made by reviewers in dyadic peer feedback. Higher proficiency reviewers made more suggestions on peers’ writing than lower proficiency reviewers. This difference was most apparent when higher proficiency reviewers were paired with lower proficiency writers, in which case the most suggestions were made. On the other hand, when lower proficiency reviewers were paired with higher proficiency writers, the fewest suggestions were made. Finally, writer proficiency influenced the number and type of suggestions that were incorporated in the revised drafts, such that lower proficiency writers tended to incorporate fewer of the meaning-related suggestions made by their peers. In the following discussion we consider how these findings may impact learning of writing skills and knowledge. Moreover, we consider how the results may be usefully applied in the classroom when teachers are dealing with multi-level classes.
Socio-cultural theory states that the learner’s ZPD is important for learning and peer feedback can be conceptualized within this framework. Lundstrom and Baker (2009) suggested that when the dyad’s ZPDs differ, learning may vary. Specifically, the reviewer effectively sets the level of the review to his or her ZPD and this may or may not be similar to the level of the writer’s. If the levels are considerably different then the writer may not gain so much from the feedback (either because the feedback is pitched too high or too low for the writer), while the reviewer, in contrast, may still gain from working with the text at their own level. In other words, if the writer’s ZPD is different from that of the reviewer, he or she may not gain feedback that scaffolds learning, which would lead to less learning (Lundstrom and Baker, 2009; Nassaji & Swain, 2000). A case study by Hamp-Lyons (2006) showed a similar result; an L2 learner lacked the capacity to incorporate her teacher’s feedback into her papers, and as a result the feedback process was fruitless.
It should be noted that Lundstrom and Baker’s (2009) study was highly unusual in at least two important aspects: peers did not interact in any dialogic form and did not both get the chance to give and to receive. Numerous studies suggest that it is through the bilateral interaction between participants that learning occurs. Moreover, peer feedback in a real classroom situation involves learners interacting with both peers giving and receiving feedback. In the present study, which was conducted in an authentic classroom context, both peers engaged in giving and receiving feedback, which means they both could benefit from engaging in the giving of feedback and according to Lundstrom and Baker’s (2009) study, could make gains in writing ability. However, the question in the present study concerns the amount of giving that students can achieve at different proficiency levels and when they are paired with peers who are matched or different proficiency. Higher proficiency peers could give more feedback than lower proficiency peers in general and this was most pronounced when the writer was lower in proficiency. According to Lundstrom and Baker, then, higher proficiency reviewers should presumably gain more because they could give more feedback. Lower proficiency reviewers could in contrast gain less because they could give less, especially when they were reviewing higher proficiency peers’ writing.
An alternative view is that giving feedback is not the only way in which learners can benefit from peer feedback activities. Most importantly, when lower proficiency writers work with higher proficiency peers, they have the opportunity to read better examples of writing in a similar genre, which may indeed be useful for learning. In this sense, the better piece of writing serves as a meditational tool that can be used to increase learning for the lower proficiency writer. Thus, lower level learners do have the opportunity for learning during peer feedback with higher proficiency peers, though the present study suggests that at least in terms of giving feedback, they may be disadvantaged in such cases.
In terms of pedagogical applications, mixed proficiency pairings offer different opportunities for giving feedback for the higher and lower peers. In mixed proficiency dyads, low proficiency reviewers are less able to provide feedback on their high proficiency peer’s writing, which we assume would entail minimal learning for the reviewer. The higher proficiency peer, on the other hand, can provide plenty of feedback on their lower proficiency peer’s writing, meaning they can still make learning gains. In contrast, when peers’ L2 proficiencies are matched, both peers can give a similar amount of feedback, which should lead to equivalent gains in long-term learning. As most instructors will aim to improve students’ learning of skills as opposed to perfecting a final product, matching proficiencies may thus be advisable. At the very least, teachers must consider L2 proficiency in addition to other factors, such as classroom dynamics and the presumed topic-knowledge of peers, when assigning pairs in mixed level classrooms. Assigning dyads with peers that differ greatly in terms of their L2 proficiency should perhaps be avoided in the aim of allowing both learners the opportunity to give adequate feedback and thus promote their own learning. Similarly, when peer feedback tasks are assessed in mixed-proficiency classrooms, L2 proficiency should be considered in order to allow all students an equal opportunity to provide sufficient and useful suggestions for improvement.
Our results also demonstrated that over half of suggestions were actually incorporated into drafts. This demonstrates a reasonable contribution of peer feedback to revisions when compared to previous studies such as Connor and Asenavage (1994) who found only a minor contribution of peer comments to revising (around 5% of total revisions). While it is beyond the scope of this article to provide discussion of the different impact of peer versus teacher feedback and the potential reasons for why any difference may exist, it can be noted that studies have typically found that teacher feedback typically results in a greater number of revisions. For example, Yang et al. (2006) found that 90% of teacher feedback was incorporated compared to 67% of peer feedback. In another study, Paulus (1999) found that 87% of teacher feedback led to revision compared to 51% of peer feedback. While there is a considerable variation between individual participants in the amount of feedback incorporated (Tsui & Ng, 2000), our study is in line with the general findings that students are selective about incorporating feedback into their revised drafts, with roughly half (54%) of peer feedback being incorporated. One potential reason why suggestions were not incorporated more often is that some of these were outside the range of the writer’s ZPD.
Another possibility is that the reviewers’ comments differed in terms of quality, with some suggestions leading to improved texts and others actually leading to language errors. To investigate this possibility we counted all suggestions that appeared to be erroneous or could potentially worsen the quality of the text, reaching 100% agreement following discussion. As expected this figure was extremely low (n = 45, 4% of total suggestions), indicating that almost all reviewers’ suggestions were of reasonable quality. Of the problematic suggestions, just under half (n = 17) were actually incorporated, showing that problematic feedback could negatively affect writing quality, though the number of cases of this were minimal. Both high and low proficiency reviewers made these poor suggestions (19 and 20 reviewers, respectively) and they were incorporated by writers of both high and low proficiency (8 and 9 writers, respectively). In sum, very few erroneous or misleading suggestions were made and language proficiency did not influence the likelihood of making such suggestions, or incorporating them into the revised text.
Social dynamics, breakdowns in communication and other aspects of dyadic interactions that are relevant for collaborative learning may explain why some suggestions are incorporated and others are not. Data from peers’ spoken interactions may have may have helped us to explain, for example, why considerably more meaning-related revisions were made than meaning-related suggestions. Storch (2011) states that various factors such as the nature of the task, student attitudes about group work in general, and the mode of communication need to be taken into account when implementing collaborative writing tasks, which are similar in many ways to peer feedback tasks. We are currently replicating the present study with a similar population of participants and are using microgenetic analysis of peer dyads that are of matched or mixed proficiencies. This future work will add an additional level of data to the present research. Nonetheless, the present study adds to a rather sparse research literature looking at interactive writing in the foreign language context and provides extensive data regarding the number and type of suggestions that result from peer interactions.
Finally, because we did not conduct tests of writing before and after the course, this means that we can only hypothesize about how much learning may occur in dyads that are of mixed proficiency or matched proficiency. Though we suggest that matched proficiency pairs may be optimal for achieving similar quantities of suggestions from all students, and thus similar levels of engagement with the task, we are not suggesting that peer feedback has no mutual benefits when proficiencies are different. Studies have shown that regardless of proficiency, learners can learn in a variety of ways (Donato, 1994; Ohta, 1995; see Villamil & de Guerrero, 1996) and experts can assist novices to attain higher learning, while also being assisted or having the opportunity to work on their own language development. Also, having the opportunity to read a better piece of writing may certainly be a valuable learning experience for lower level learners. Indeed, peer feedback is so useful precisely because it can be applied to mixed level classes. However, our study confirms that L2 proficiency is, as many had expected, an important factor that should be considered by teachers when managing peer feedback activities.
Footnotes
Appendix 1
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
