Abstract
Although keystroke logging promises to provide a valuable tool for writing research, it can often be difficult to relate logs to underlying processes. This article describes the procedures and measures that the authors developed to analyze a sample of 80 keystroke logs, with a view to achieving a better alignment between keystroke-logging measures and underlying cognitive processes. They used these measures to analyze pauses, bursts, and revisions and found that (a) burst lengths vary depending on their initiation type as well as their termination type, suggesting that the classification system used in previous research should be elaborated; (b) mixture models fit pause duration data better than unimodal central tendency statistics; and (c) individuals who pause for longer at sentence boundaries produce shorter but more well-formed bursts. A principal components analysis identified three underlying dimensions in these data: planned text production, within-sentence revision, and revision of global text structure.
Keystroke logging promises to be a productive tool for research into writing, and it has grown in popularity over recent years as a research tool (for recent reviews, see Sullivan & Lindgren, 2006; Van Waes, Leijten, Wengelin, & Lindgren, 2012). A major attraction of the method is that it provides an unobtrusive record of the moment-by-moment creation of the text. However, in its raw form, this record provides only information about empty time (when no keys are being pressed) and filled time (when, and which, keys are being pressed). To make sense of this, the researcher needs to relate this to a model of the different types of processes involved in writing. Furthermore, the complex and recursive nature of the writing process means that undifferentiated measures of global properties of keystroke logs are likely to be extremely insensitive measures of underlying writing processes. Indeed, some authors have claimed that this may be a fruitless enterprise. Schilperoord (2001) has suggested that
if the writing process that one wants to study is characterized by intensive problem-solving and massive editing on the part of the writer, with actual language production being but one aspect of these processes, it [may] prove impossible to relate pause data to ongoing processes. (p. 67)
In this article, we describe how we took up this challenge and tried to derive measures from the raw output of keystroke logs that could be related to cognitive models of writing. These logs were collected as part of a larger project investigating the effects of self-monitoring and planning on the development of understanding in writing. We focus exclusively here on the procedures that we used to construct measures from the keystroke logs and on the properties of these measures across the sample of texts as a whole. In future works, we plan to investigate whether these characteristics vary as a function of self-monitoring and type of planning, as well as how they relate to the development of understanding and text quality. Our reason for focusing on the methodological and conceptual issues here is that, although there is a growing body of research using keystroke logs to examine specific aspects of writing (see Sullivan & Lindgren, 2006), there has been relatively little research investigating the writing process as a whole across larger numbers of writers. Furthermore, the research that has been published in this area (e.g. Van Waes & Schellens, 2003) tends to focus on the results of the research rather than the methods used to achieve them. Our aim therefore is to describe the procedures and measures that we have developed. First, we discuss some of the issues that arise in interpreting the basic units of analysis available from keystroke logs—pauses, bursts, and revisions—in terms of cognitive models of writing. We then outline our broad approach to addressing these issues. In the analytic section, we describe the procedures that we have developed for preparing keystroke logs for analysis and examine the properties of the measures that we constructed. Finally, we assess whether these measures can be combined to create global measures that can aid our understanding of writing processes.
Keystrokes and Models of Writing
Flower and Hayes’s (1980) model of the writing process identified three kinds of basic processes and suggested different ways in which they could be coordinated during writing: (a) planning, which included setting goals and generating and organizing ideas to satisfy those goals; (b) translation; and (c) revision. The coordination of these processes was controlled by a monitor, and different configurations of the monitor were assumed to correspond to individual differences in writing strategy (see Table 1). The four configurations differ in whether planning is carried out before, or at the same time as, text production and in whether revision is carried out at the same time or after text production.
Overview of Writing Processes Mapped Onto Stages of Writing
More recently, Hayes (1996) has suggested that revision should not be considered a basic process in its own right but should instead be seen as involving the recursive application of cycles of reading, reflection, and text production. Furthermore, Hayes and Nash (1996) have refined the characterization of planning, distinguishing between process planning (focused on the management of the process itself), abstract planning (concerned with goal setting and content generation), and language planning (concerning the formulation of content in language).
Even this brief sketch of some basic components of a cognitive model of writing makes it clear that there are major problems in tying keystroke units of analysis to specific components of the writing process. Pauses may reflect not just various different levels of planning and reflection but also rereading and text production during revision. Bursts of language may reflect an initial formulation of thought or an attempt to improve previously formulated text. Revisions may reflect semiautomatic correction of errors or a systematic attempt to modify content.
Hayes’s Model of Text Production
In this section, we assess how the basic units of keystroke analysis can be interpreted in terms of the most fully developed current model of text production (see Hayes, 2009, for a review). A sketch of the model, taken from Chenoweth and Hayes (2003) is shown in Figure 1.

Model of the text production process from Chenoweth and Hayes (2003, p. 113)
The model distinguishes among four different processes. The proposer proposes ideas for expression. This component is assumed to include the higher-level processes involved in planning and reflection, as characterized in global models of the writing process, and is responsible for creating an idea package to be formulated in language. Essentially, it is responsible for deciding what to say next. The translator is responsible for converting this message into linguistic strings. The transcriber then converts this linguistic string into written text. Finally, the evaluator/reviser is responsible for monitoring and evaluating word strings and text as they are produced and for revising them when they are found wanting.
In its general form, this model corresponds directly with what has become the standard model of spoken language production (Levelt, 1989, 1999; Vigliocco & Hartsuiker, 2002). We assume, since the output from the translator process in Hayes’s model is a phonological representation, that up to this stage of the process, it is identical to a model of spoken language production. There are then a range of differences between speech and writing. Of these, two seem to us to be particularly important. First, the external representation of language in writing means that the writers have access to the preceding text in planning coherent continuations of what they have already said and are not restricted to the fading trace of speech in short-term memory. This means that pauses between segments of text could reflect rereading of previous text as well as planning of the unit of text itself. Second, there may be important differences in the extent to which internal and external monitoring can be carried out. Vigliocco and Hartsuiker (2002) have suggested that although inner speech can in principle be monitored, this cannot be done at the same time as speaking, arguing that externally articulated speech may overwrite the representation of inner speech in short-term verbal memory. In our view, this may be an important difference between speech and writing. Writing may allow the monitoring of inner speech at the same time as transcription and hence may allow for greater revision prior to transcription. This means that pauses during writing may reflect revision of planned language as well as the mental planning of the next unit of language.
This receives some support from the study carried out by Kaufer, Hayes, and Flower (1986), which compared thinking-aloud protocols produced by expert and novice writers. We describe this in detail to illustrate how bursts in text production are typically analyzed. Kaufer et al. segmented the protocols into language bursts. Bursts that ended in a pause of two or more seconds were classified as P-bursts; bursts that were terminated by an evaluation, revision, or some other grammatical discontinuity were classified as R-bursts (Chenoweth & Hayes, 2001). Figure 2 shows an illustrative example adapted from Kaufer et al. (1986) of the process of producing a sentence, broken down into segments and bursts. We have typed in bold text what we assume was transcribed, though the moment at which it was transcribed is not indicated in the original paper.

Process of the production of a sentence broken down in bursts and segments. Taken from Kaufer, Hayes, and Flower (1986, pp. 125-126).
The writer starts by proposing language, which is then transcribed. Because the part appears in the think-aloud protocol, we take this as representing the output of the translator rather than the transcriber. Since the part ended in a pause, Kaufer et al. (1986) classified this as a P-burst and assume that this represents the full output of the content entered into the translator. The next two segments represent operations involved in creating content and reflect the output of the proposer. In Segment 4, a new sentence part is proposed and transcribed. In Segments 5, 6, and 7, the writer first searches for content and then proposes two alternative sentence parts. Neither of these is written down, so we assume that this is output from the translator, which is rejected by the evaluator. Note, though, that both sentence parts were classified as P-bursts despite the fact that they were not transcribed. By contrast, the proposed sentence part in Segment 9 is classified as an R-burst because it is followed by an explicit evaluation (Segment 10), and this occurs within 2 seconds and so is taken as a possible interruption of the output of the translator. The distinction appears to be that Segments 6 and 7 are rejected after a delay, whereas Segment 9 may have been interrupted before the output of the translator has been completed. The aim of the classification is to distinguish cases where the burst reflects the capacity of the translator from cases where the output of the translator has been interrupted before it is completed.
It is important to note that the observations are based on a spoken protocol. This means that a great many of the bursts that are taken to reflect the output of the translator will not be apparent in writing, since they are rejected before they are written down. Thus, the protocol above would, in a keystroke log, reduce to “
Finally, there is a general conceptual issue with the fact that bursts are defined only by how they are terminated. This means that bursts that follow a revision are treated as equivalent to bursts that follow a pause. It seems possible, however, that bursts that follow a revision involve modifying the previous burst instead of producing new content and hence may differ systematically in length from “pure” P-bursts, which begin and end with a pause. This would be particularly important if writers differ strategically in the extent to which they revise bursts as they are produced. In what follows, we therefore analyze bursts in terms of how they are initiated and how they are terminated, and we test whether this is associated with systematic differences in length.
Aims of the Analysis
In our view, the fundamental issue that arises in trying to use keystrokes as indicators of underlying cognitive processes is one of alignment. As we have seen, if one were to take the raw keystroke log of a writing session and then analyze it as an undifferentiated whole, the characteristics of the pauses, bursts, and revisions that one would measure would potentially reflect a very heterogeneous range of underlying processes. (This is, in fact, not infrequently what is actually done [see Wengelin, 2006, for a review].) Our general approach to the analysis therefore had the overriding aim of trying to improve the mapping of measures onto hypothetical underlying processes as much as possible. This can be broken down into four complementary aims.
First, we used the final product to sort the keystroke log into different kinds of activities. Our main aim here was to isolate text production in its “purest,” most linear form from other activities. In brief summary, this involved the following distinctions: (a) distinguishing the initial draft of text from postdraft revision, on the grounds that text production during postdraft revision may be different in form from text production during the initial draft; (b) within the initial draft, distinguishing between linear and nonlinear transitions between units of text, in order to separate text produced as part of the forward progression of the text from revision activities; (c) then distinguishing within the nonlinear events between revision at the leading edge and revision elsewhere in the text. In the analysis section, we describe the procedures that we used to do this and elaborate on some of the more fine-grained distinctions that emerged in the course of our engagement with the data.
Our next aim was to develop more differentiated measures of pauses and bursts than what have previously been used.
Pauses have typically been measured in one of two ways (Wengelin, 2006). First, by imposing a threshold—typically of 2 seconds—and then defining only the intervals that last longer than this threshold as pauses. Analysis then typically involves comparing the frequency with which these pauses occur at different text locations or the frequency with which they are produced by different writers or under different conditions. The problem is that this restricts analysis to longer pauses, which presumably reflect higher-level processes, and ignores the shorter pauses involved in more linguistic processing. The alternative approach is to use the raw intervals between events and calculate measures of central tendency for different kinds of pause durations. The problem is that the resulting distribution of pause durations is extremely heterogeneous—it mixes the exceptional, longer pauses and the more routine, shorter pauses in the same distribution. This can be reduced by grouping the pauses according to the location at which they occur (e.g., within words; between words, sentences, and paragraphs), but the resulting distributions are still extremely skewed and represent a heterogeneous mixture of different kinds of pauses. In what follows, we use mixture models (McLachlan & Peel, 2000) to identify subcomponents of these distributions and to estimate separate measures of their central tendency. Our aim was to try to achieve a better match between the pause measures and the underlying cognitive processes.
As we have seen, bursts have typically been classified as either P-bursts or R-bursts. The aim is to isolate bursts assumed to reflect the operation of the translator from bursts that may also involve the reviewer. On the basis of the issues we raised in discussing his model, we apply a more differentiated classification with a view to improving the mapping between bursts and different types of process. We assess the value of this by testing whether the resulting burst types differ in length.
Finally, we use principal components analysis to assess the relationships between different measures and the correspondence between the separate components and components of cognitive models of writing.
Method
The data we discuss here were collected as part of a larger study investigating the effects of self-monitoring and planning on writing. As a part of this study, we collected keystroke logs using Inputlog (Leijten & Van Waes, 2006). The texts were written by 80 participants recruited at the University of Groningen. Participants were asked to plan and write an article for the university newspaper discussing whether “our dependence on the computer and the Internet is a good development or not.” The experiment was divided into three phases. During the first and third phases, we administered a collection of measures that investigated the development of understanding through writing (see Baaijen, Galbraith, & de Glopper, 2010, for more information and a full description of the experiment). In the writing phase, participants were given 5 minutes to plan the writing assignment using pen and paper. Half the participants had to write down a single sentence summing up their overall opinion (synthetic planning), and the other half had to construct a structured outline (outline planning). They then had 30 minutes to write the article on a computer. It was stressed that they had to produce a well-structured and complete article in the time available. Participants were allowed to consult their written outlines.
Our aim is to describe the procedures that we used to construct measures from the keystroke logs and to examine the properties of these measures across the sample of texts as a whole. In future work, we plan to investigate whether these characteristics vary as a function of self-monitoring and type of planning, as well as how they relate to the development of understanding and text quality.
Data Preparation and Coding
Our first aim in data preparation was to isolate “pure” text production—linearly produced text—from revision and a number of other types of output.
First, we excluded titles and treated these as a separate category. Although these are clearly text production, they are very different from other examples. They often involve considerably longer pauses and may be produced at the end or in the middle of the writing session and following lengthy rereading episodes.
Second, we categorized text produced as part of explicit planning separately from text production. Some writers in our sample (20%) broke off from producing full text to make a plan on screen. Quite apart from the point that explicit planning is a distinctively different process and therefore should be analyzed separately from text production, these episodes also provide good examples of why the raw output of Inputlog should always be checked closely. When making explicit plans, participants often press ENTER to mark the beginning of a new point. Inputlog automatically codes these as between-paragraph boundaries and does not distinguish these from paragraph boundaries that are intended for inclusion in the text.
Third, we distinguished text produced during an initial draft from text produced during a revision draft. Many writers (65%) show evidence of writing an initial draft and then going back systematically through the text, editing and revising the initial draft. Although this may sometimes involve producing extended chunks of text, we did not include this in the analysis of text production, on the grounds that it may be different in character to text produced as part of the initial draft. We defined end revision as occurring when an individual made revisions outside the final paragraph while writing the final paragraph. In the majority of cases, this amounted to revisions made after the final sentence. But there were some cases where individuals broke off to make revisions and then returned to write a final summary sentence or two. In these cases, these sentences were also excluded from text production analysis.
We then distinguished between linear text production and revision during the initial draft by making a distinction between linear transitions and events. Linear transitions were defined as “empty” thinking episodes before the continuation of text production; events were defined as episodes that include other material or operations before the continuation of text production. These included scrolling and other movements that might indicate rereading or evaluating previously written text or the insertion of text away from the leading edge. This was done, on the basis of the automatic output from Inputlog, by flagging objective locations where other material or movements were executed before the continuation of text production. As well as allowing us to assess pause durations from of the linear transitions alone, this enabled us to quantify the amount of revision at different text levels by calculating the percentage of linear and nonlinear transitions at different pause locations.
Finally, we defined pauses in a way that corresponds to a model of spoken language production rather than in terms of the keystrokes. Typically, a between-word pause in the keystroke log consists of two pauses: one pause before the between-word marker (SPACE) and one pause before the first letter of the next word. These are often classified separately (e.g., Wengelin, 2006). In Inputlog, these separate locations are coded with the same code so that mean between-word pause durations provided in the automatic summary output represent a hybrid of before and after <SPACE> pause durations. We decided to take the sum of these two pauses together on the assumption that space presses between words and sentences are relatively automatic motor activities and that the interval between the end of one word and the beginning of another is a better reflection of the cognitive processes between these units. We applied the same principle to the calculation of pauses at other locations.
Overall, then, we distinguished the following pause locations: within words, between words, between subsentences (indicated by commas, semicolons, and colons), between sentences, and between paragraphs. These correspond to the boundaries used in other research (e.g., Wengelin, 2006) but differ in the way that they are calculated.
Results
Coding Language Bursts
This analysis was carried out on the sections of the keystroke logs corresponding to the initial draft. First, we identified automatically all interruptions longer or equal to 2 seconds and then classified these as either (a) P-boundaries, when they were associated with a linear transition, or (b) R-boundaries, when they were associated with an event. Second, we scanned all other automatically flagged events to identify whether they were associated with revisions of text. A number of important features should be noted.
First, typos are extremely common in keystroke logs. Following conventional practice (see Wengelin, 2006), we excluded these as indicators of revision and included the relevant sections of text as part of a burst. We classified typos as corrections of errors within a word, which leave the word otherwise unchanged and which do not result in a break of more than two seconds before the next keystroke.
Second, another characteristic of keyboarded texts is that writers are much more able to modify previously produced text. In our coding, instances where writers leave the leading edge to insert text elsewhere are automatically flagged as nonlinear transitions. We classified bursts terminated by insertions as a separate category of revision burst (RI;
Third, we classified the insertion-bursts themselves as a separate kind of burst taking place during revision. Since these are produced in the circumscribed context of existing text, we would expect them to be shorter than P- or R-bursts. Overall, we distinguish among three kinds of insertion-bursts: (a) IG-bursts reflect within-sentence revision; (b) IR-bursts reflect end-of-sentence revision; and (c) IB-bursts reflect revision over sentence boundaries (see Table 2 for a detailed definition of bursts). For all bursts, length was calculated by counting all words within the burst, including any partial completions of words.
Definition for Different Language Bursts Distinguished on the Basis of Keystroke-Logging Data
The final distinction was designed to enable us to test whether there is a difference between bursts initiated after a pause and bursts initiated after a revision. We classified pure P-bursts (PP-bursts) as bursts that begin and end with a 2-second pause and distinguished these from RP-bursts, which begin after a revision and terminate with a 2-second pause. We distinguished three types of RP-bursts: RP1-bursts are bursts that consist entirely of new language terminated by a pause; these typically occur following an insertion elsewhere in the text, and we would expect these to be similar in length to PP-bursts. RP2-bursts involve the replacement of preceding text, followed by further text production terminated by a pause. RP3-bursts are bursts that involve only the replacement of the preceding text and then pausing before continuing with text production. We would expect these to be shorter than PP-bursts. In principle, we think that PP-bursts should be separated from RP-bursts if they are to be used as estimates of the language capacities of writers.
Finally, we applied the same principle to R-bursts and classified these separately as PR-bursts and RR-bursts. Table 2 shows a list of the different categories of bursts, along with brief definitions.
Lengths of Different Burst Types
We tested first whether there was an overall difference in length between PP, RP, PR, RR, and I-bursts using a one-way within subjects analysis of variance (ANOVA), with burst type as the independent variable and mean burst length as the dependent variable. Because the data lacked sphericity, degrees of freedom were adjusted using the Greenhouse-Geisser correction. The results showed a highly significant difference in burst length depending on burst type, F(3.49, 272.44) = 144.00, p < .0005. Planned pairwise comparisons, with a Bonferroni adjustment for multiple comparisons, revealed no significant difference in length between PP (M = 5.87 words, SD = 1.87) and RP (M = 5.58, SD = 1.26) bursts, t(78) = 1.71, p = .90, and no significant difference between PR (M = 4.43, SD = 1.25) and RR (M = 4.65, SD = 1.25) bursts, t(78) = 1.32, p = 1. These initial results suggest that whether a burst follows a pause or a revision has no effect on the length of the burst. However, there were highly significant differences between P-, R-, and I-bursts: R-bursts were significantly shorter in length than both P-bursts (p < .0005 for all comparisons), and I-bursts (M = 2.22, SD = 1.45) were significantly shorter than all other burst types (p < .0005 for all comparisons). These results support the assumption that R-bursts are bursts that have been interrupted before they have been fully executed. They also suggest that insertions in earlier text tend to consist of short, partial bursts modifying earlier bursts rather than the production of full bursts of new language.
The next step in the analysis was to investigate subcategories within the RP, PR, and RR categories. A breakdown of the subcategories is shown in Figure 3.

Average lengths for different burst types with standard error bars
First, we used a one-way within subjects ANOVA to compare the burst lengths of the different RP-types with pure P-bursts. This showed a significant overall effect of burst type, F(1.90, 136.75) = 143.65, p < .0005, with the RP3-burst type (M = 1.70, SD = 0.97) being significantly shorter than all the other bursts (p < .0005 for all comparisons). These results suggests that RP3-bursts, which consist simply of the replacement of previous text, are much shorter than bursts where new language is being added and should not be included in estimates of P-burst length.
Second, to test whether revision bursts terminating with revision at the leading edge (RL) differ from revision bursts terminating with an insertion elsewhere (RI) and whether this is affected by how the burst is initiated (following a pause [PR] or a revision [RR]), we carried out a two-way within subjects ANOVA, with termination type and initiation type as independent variables and mean burst length as the dependent variable. We excluded the mixed category (PRLI and RRLI) from this analysis because we wanted to compare pure examples of each type. This showed a significant effect of termination type, F(1, 46) = 6.17, p = .017, but no significant effect of initiation type, p = .36. The results confirm the earlier finding that R-bursts do not vary in length depending on whether they are produced following a pause or a revision. However, the results also suggest that RL-bursts (M = 4.31, SD = 0.15) are shorter than RI-bursts (M = 4.95, SD = 0.24). One possible explanation for this is that RIs occur when evaluation is applied after the burst has been produced, and, hence, a burst is more fully completed than when evaluation is applied during burst production. Note that RI-bursts are not as long as pure P-bursts, suggesting that these bursts are not as fully planned as a normal P-burst.
Pauses
Pausing durations were measured from all the linear continuation pauses at each different location. As can be seen in the histogram of the between-word pause durations for one of the participants (P307) in our data (Figure 4), their distribution is heavily skewed.

Histogram of Participant 307 showing the particularly skewed distribution of between-word pauses
A standard transformation to reduce positive skew is to take logs. The histogram of the resulting distribution is shown in Figure 5.

Histogram of log transformation of between-word pauses for Participant 307
Although this does pull the extreme pause durations in toward the main body, it is clear that the distribution is still extremely skewed. Furthermore, the left hand of the distribution is noticeably bimodal, which implies that the median is a poor estimate of central tendency. This has also been observed by Kirsner, Hird, and Dunn (2005) in between-word pauses for spoken language production. They suggest that, rather than treating this as a single skewed distribution, a better approach is to assume that the distribution is a mixture of several components and to fit mixture models to estimate the parameters of the distributions. We used the R-package EMMIX developed by McLachlan and Peel (2000) to fit a series of models to the data and estimated the goodness of fit of these models using the Bayesian information criterion (BIC). The best-fitting distribution for this participant (P307) for the between word pauses is shown in Figure 6.

Histogram showing three fitted models for Participant 307
This model had a BIC value of 476.33, which was overwhelmingly a better fit than a single normal distribution (BIC = 5,871.08), a single lognormal distribution (725.16), a double lognormal distribution (550.25) and the model with four normal distributions (627.65). (Lower values are superior, and Kass and Raftery [1995, p. 777] suggest that a difference greater than 3.7 can be considered a positive difference, and a difference greater than 20, a strong difference.)
For this particular participant, there appears to be a good case for suggesting that there are three different distributions within the data for the between-word pauses. One might speculate that the left-hand distribution (65% of the pauses) represents word retrieval processes, the middle distribution (26% of the pauses) represents phrase boundary processes, and the right-hand distribution (9% of the pauses) represents higher-level message planning or reflection. Note, though, that some of the more extreme pauses are extremely long (the longest pause for this participant is almost 23 seconds), which suggests that more is involved than a higher level of structural planning. Indeed, it is doubtful whether the right-hand distribution should be treated as a single normal distribution since, as shown in Figure 6, the symmetry assumption implies that it also includes a proportion of extremely short pauses. Our conclusions for this participant are (a) that there are two normally distributed groups of pauses and that these would be better represented by two lognormal means than by a single median and (b) that the more extreme pauses should be treated as a miscellaneous group identified by a threshold determined by the tail of the second lognormal distribution. In this case, three standard deviations above the mean, the threshold would be log 7.43, which is equivalent to a pause duration of 1,686 milliseconds.
To assess the generality of these findings, we fitted the same series of models to the pause durations for all participants and assessed the goodness of fit of the models using the BIC values. This analysis showed that for 58% of the participants the three-distribution model was the best fitting, with 19% of the participants being best fit by a two-distribution model, and the remaining 23% being indeterminate between two and three distributions.
We repeated this analysis for the other pause locations. For the pauses between sentences and subsentences, single lognormal distributions generally had the best fit. It was harder to determine which distribution was the best overall fit for the within-word pauses. Generally, three distributions were a better fit than four distributions (54% of the participants), and therefore, we selected these lognormal scores for the analysis of the within-word pauses.
To assess the differences in pause duration for different text locations, a one-way within subjects ANOVA was conducted comparing the means for the left-hand distributions at each text location. There was a significant main effect of pause duration, Greenhouse Geiser F(2.56, 77) = 1,378.68, p < .001. Planned comparisons, with a Bonferroni adjustment for multiple comparisons, showed that pause durations at all locations were significantly different from one another (p < .0005 in all cases). These results (see Table 3 for an overview of pause times at different locations) confirm previous findings suggesting that pause length increases for higher-level locations within the text (Wengelin, 2006).
Means and Log Scores for Different Pause Locations
Further analysis of the subdivisions within specific pause locations showed that the means from these distributions were still significantly lower (p < .0005) than the pause times for the next level up. This suggests that these subdivisions correspond to subcomponents of the distributions at particular locations rather than to overlaps with higher-level text locations.
We then assessed the relationship between pause duration and burst length. A natural assumption is that the longer one pauses, the longer the burst length. First, we correlated the mean (log) pause durations for each individual at between-word, between-subsentence, and between-sentence locations with their mean PP-burst and mean R-burst lengths. As can be seen in Table 4, these were negative in all cases and very significantly so for between-subsentence and between-sentence locations. This suggests that the average length of time that an individual pauses at grammatical boundaries is, in fact, negatively associated with the average length of the language bursts they produce. To assess this within individuals, we then calculated the correlations at between-word and between-sentence levels for all an individual’s pauses and the associated burst lengths in their text production samples. Although a few individuals showed significant correlations, these were not consistently positive or negative, and the majority of individuals (85% for the subsentence level and 96% for the sentence level) showed no significant relationship between pause duration and pause length.
Pearson Product - Moment Correlations Between Mean Pause Length and Burst Characteristics
p < .001 (two-tailed).
An alternative possibility is that longer pause durations are associated with producing “better bursts,” perhaps through mental evaluation and revision prior to production. To assess this, we calculated the correlation between mean (log) pause duration and the percentage of P-bursts for across participants. As can be seen in the table, there was a highly significant positive correlation for the between-sentence-level pauses, with longer pauses being associated with a higher percentage of P-bursts.
Revision
We assessed revision in four ways. First, we calculated a number of global indicators of the extent to which the text was modified in the course of creating the final product. These included (a) the text modification index, which consists of the ratio of the total number of words in the process log to the total number of words in the final product, with higher numbers indicating greater amounts of revision, and (b) measures of the total amount of text deleted as a percentage of the total number of process words.
Second, to assess changes made during postdraft revision, we calculated the number of words that were modified during postdraft revision (as a percentage of the total number of words) and the number of words that were added during postdraft revision (again as a percentage of the total number of words).
Third, to distinguish revision at the leading edge from revision elsewhere, we assessed these separately. To assess revision at the leading edge, we counted the total number of words modified at the leading edge as a percentage of the total number of process words produced during the initial draft and the mean amount of deletion carried out in each of these instances. To assess nonlocal revision, we counted the total number of words added during insertions (see earlier burst analysis section) as a percentage of the total number of process words produced during the initial draft and the mean size of these insertions.
Fourth, we used the distinction between linear transitions and event-filled transitions to assess the extent to which the linear progression of text production was disrupted at different levels of text structure. To do this, we calculated the percentage of linear transitions compared to nonlinear transitions (events) at each different text level. These measures provide a fine-grained measure of the forward movement of the text and—by measuring this separately for word, sentence, and paragraph transitions—enabled us to assess different levels independently. Note that these measures capture linearity in processes as well as product. Thus, a nonlinear transition between words can, in principle, involve anything that is not a straightforward continuation to the next word: It could be a matter of scrolling elsewhere in the text to reread and then returning to the leading edge to produce text, or it could involve revising at the leading edge or at another location within the text. To assess linearity at a more global level and in a way that captured product changes alone, we also calculated a sentence linearity index. This involved comparing the order of sentences in the final product with the order that they were produced in. Each sentence in the final product was numbered by when it was produced during text production, and the percentage of nonlinear sentences was calculated.
Interrelationships Between Pauses, Bursts, and Revision During Text Production
To assess the relationships between these different measures, we carried out a principal component analysis (PCA) on 16 variables selected to cover the range of different activities and because they satisfied the assumptions required for PCA. The set of measures had a Kaiser-Meyer-Oklin sampling adequacy of .71, and Barlett’s test of sphericity was highly significant (p < .001), indicating that PCA is appropriate for these data (Field, 2005). Varimax rotation was used to extract orthogonal components from the data.
We selected the five components with eigenvalues over Kaiser’s criterion of 1 and which, taken together, explained 73.6% of the variance. Table 5 below shows the loadings of each variable on these components after rotation, as well as the amount of variance that each component accounts for and Cronbach’s alpha for each component.
Principal Component Analysis With Varimax Rotation for Five-Factor Solution
Note: Factor loadings over .40 appear in bold. Component 1 = planned sentence production; Component 2 = within-sentence revision; Component 3 = revision of global text structure; Component 4 = postdraft revision; Component 5 = careful word choice.
The components can be described in the following ways:
Component 1—Planned sentence production: This component consists of a combination of longer pauses at grammatical boundaries (sentences and subsentences), as well as the exceptionally long pauses between words, along with short bursts.
Component 2—Within-sentence revision: This component consists of high levels of revision at the leading edge of the text combined with nonlinear transitions at the word and burst levels. This suggests that it reflects the extent to which different writers revise sentences as they are produced.
Component 3—Revision of global text structure: This component represents the extent to which the global plan of the text is revised as captured by the between-sentence linearity. Note that it captures a revision at a higher level than the preceding component.
The remaining two components are much lower in reliability, with few item loadings, making them harder to interpret. Their main interest is in suggesting further independent features of the process for which additional measures could be developed. We have tentatively labeled these as
Component 4—Postdraft revision: reflecting the extent to which revision is postponed until after the initial draft; and
Component 5—Careful word choice: reflecting perhaps the extent to which words are chosen carefully.
Component 5 is mainly of interest in demonstrating that different aspects of between-word pauses load on different components, supporting the suggestion that these should be measured separately.
Discussion
We set out on this analysis with four main aims, which we discuss in turn.
Our first aim was to establish a set of procedures for separating out different components of the writing process to be better able to relate keystroke measures more directly to specific cognitive processes. Overall, we think that the procedures that we employed did enable us to separate “pure” text production from other processes and to distinguish among different forms of revision. This is an important initial step in aligning keystroke units of analysis more directly with the cognitive components of writing process models. We would not wish to claim, however, that the procedures we have used are necessarily the best way of doing this, nor that they are the only procedures that need to be applied. We have focused almost entirely on procedures for isolating “pure” text production from other activities and on procedures that can be carried out automatically or quasi-automatically. Our main point is to emphasize the need to carry out these kinds of procedures and the general principle that the procedures should be designed to isolate components of the keystroke log that can be mapped on to components of process models of writing. Our impression from reading reviews of research in this area (e.g., Sullivan & Lindgren, 2006) is that it is not uncommon for researchers to use the automated output from a software analysis program to analyze undifferentiated logs of the whole text.
Perhaps the most important of these procedures was the separation of linear transitions from event transitions, which had a number of important benefits. First, it enabled us to separate “pure” text production within the initial draft from other processes and, hence, to base our analysis of pauses solely on those associated with the forward progression of the text. Although these may in part reflect rereading, the fact that they are associated directly with the following unit of text makes it much more plausible to interpret them in terms of the planning of text production. In addition, this enabled us to distinguish between revision at the leading edge of the text and revision elsewhere in the text, and—in the form of percentages of linear continuations at different text levels—it provided us with global measures of the extent of revision at different levels of text production.
Finally, our experience of coding the logs reinforced the need (stressed by the developers themselves) to be wary of the automated output provided by the software. Logs and associated automatic statistical analyses really do need to be closely screened before analysis begins.
Pauses
Our second aim was to use mixture modeling to estimate the parameters of pauses at different locations. Our results suggest that this has a range of advantages, and we would recommend that this become standard practice. First, it extends the general principle of differentiating different components of the writing process down to the pause level of analysis. Modeling the pauses between words, for example, enabled us to distinguish between routine transitions between words, a further distribution centered on a longer pause duration, and exceptional cases where some nonroutine and presumably higher-level reflective activity is taking place. We have suggested that these might reflect lexical retrieval, phrase structure processing, and higher-level message planning, but we do not want to make strong claims about this. Identifying the source of these differences should clearly be an important topic for future research. The key point is that, without modeling the distribution of pauses, these distinctions would remain completely unobservable. A second benefit of doing this is that it enables one to assess effects on the different components of the distributions separately. We found, for example, that the group of exceptionally long pauses loaded on the component in the PCA corresponding to planned sentence production along with other measures of higher-level sentence processing, whereas the more routine group of pauses did not. This approach also enables one to take a more principled and contextualized approach to establishing thresholds for identifying nonroutine pauses at different locations. Furthermore, thresholds can be established relative to the particular characteristics of individual writers rather than against a general criterion. Finally, we believe that this approach can be extended to pauses at a higher level than just within-word and between-word pauses. Although we found that a single lognormal distribution fitted the data best for between-sentence pauses, this is probably a reflection of the relatively low number of such pauses present in the relatively brief texts that we analyzed. We would expect analysis of much longer texts to reveal the same kinds of multimodal distributions for between-sentence pauses as we have found for within- and between-word pauses.
Bursts
The principal aim of this analysis was to establish how best to divide up the bursts of language occurring during text production. However, in the course of doing this, we also observed that writers who pause for longer on average before writing sentences tend to produce shorter rather than longer PP-bursts and that, instead, they tend to produce a higher percentage of clean p-bursts in their text. It seems to us that if, as Hayes assumes, PP-bursts are a direct reflection of the capacity of the translator component of text production, then one would expect writers who typically pause for longer at grammatical junctures to also produce longer PP-bursts. The fact that they do not and instead produce a higher percentage of clean bursts suggests that the bursts that appear in written output reflect not only planning of content but also the extent to which the writer is concerned with producing well-formed text and, hence, the operation of the reviewer component of the model. This could be tested by assessing whether writers demonstrating this feature in keystroke logs also show more evidence of revision prior to transcription in think-aloud protocols. More generally, these results imply that research into the relationship between bursts and the capacity of the translator may be more valid if they are carried out on spoken output or on relatively spontaneous kinds of writing where the use of strategies is minimized.
This conceptual issue aside, our aim with this analysis was to establish whether a more refined classification system is required. We evaluated this by examining how the burst types within this classification system varied in length. Our first finding was that insertion bursts were much shorter than normal P- or R-bursts. This provides strong support for our general assumption that text produced as part of the forward progression of a draft should be analyzed separately from text production occurring during revision. We should stress, however, that we do not regard this necessarily as an intrinsic feature of insertion bursts. Our texts were collected from writers producing a completed text in half an hour. Under other circumstances, writers might engage in much more extensive revision and produce larger I-bursts. Rather, our point is that for analytic purposes these bursts should be classified separately; future research is needed to establish how these bursts vary for different writers and under different conditions.
Our second main finding was that RP3-bursts, which involve only the modification of previous output, were much shorter than other bursts. These bursts that terminate with a pause but are initiated following a revision would be classified as P-bursts using the standard P- and R-burst classification (see Chenoweth & Hayes, 2003, for an example) and hence would seriously contaminate the collection of P-bursts. For analytic purposes we therefore recommend that bursts be classified in terms of how they are both initiated and terminated.
Finally, we found that revision bursts terminated by revision at the leading edge were significantly shorter than revision bursts terminated by an insertion. It is possible that the relative frequency of these two types of revision bursts may vary among writers employing different drafting strategies, with some writers deliberately postponing revision to the end of sentences and others evaluating and revising bursts as they are produced. Future research should assess these differences.
Revision
The approach to revision that we have taken has focused less on different types of revisions (e.g., Lindgren & Sullivan, 2006; Stevenson, Schoonen, & de Glopper, 2006; Witte & Cherry, 1986) and more on when they occur and on simple quantitative measures of how much they occur. In part, this was a deliberate decision to focus on measures that might in principle be automated. But it was also a side effect of the fact that, in focusing on isolating “pure” text production from text production produced during revision and then on developing more fine-grained measure of pauses and bursts during “pure” text production, we paid less attention to the revision component of writing. We have, for example, not actively examined the role that pauses play in revision. Clearly, this is an area that should be developed further in future research.
That said, there are two features of the measures that we used that should be noted. First, in line with the general principle underlying our approach, we have distinguished between the contexts in which revision occurs—after a draft has been completed, at the leading edge of text production, and as nonlinear departures from the leading edge—rather than treating them all as examples of an undifferentiated revision process. The fact that these three forms of revision load on different components in the PCA suggests that they are independent activities. Amount of revision at the leading edge is not, for example, inversely related to amount of postdraft revision. This suggests that there is not a simple trade-off between revision at the same time as producing text and revision after a draft has been completed, as is implied by the different configurations of the monitors described in the original Hayes and Flower model. Rather these may involve different kinds of revision and may need to be analyzed separately.
Second, the revision measures that we have used capture complementary features of the process. The measures that we have just discussed capture the amount of revision at the leading edge or elsewhere; the linearity measures (percentage of linear transitions at word, sentence, and paragraph levels) capture the extent to which this happens at different levels of the text. Furthermore, since nonlinear events included simple scrolling, these measures capture the minimal component of revision involved in rereading previous text. In combination, these measures provide information about both the amount of revision at different places in the text and the point at which this is initiated. The fact that linearity of sentence transitions loads on the same component in the PCA as amount of revision insertion, whereas linearity of word transitions loads on the same component as revision at the leading edge, suggests that revision elsewhere in the text is more likely to take place at sentence boundaries than within sentences.
Interrelationships Between Pauses, Bursts, and Revision
Our aim with the PCA was to identify groupings of interrelated measures, considering both how different measures were related to one another, and the range of independent characteristics of text production that they captured. As we have seen already in the discussion, this played a valuable role in helping to interpret measures in terms of potential underlying processes. We should stress, however, that this is not a full factor analysis. Given the small sample of texts, the results of the PCA should be taken only as an economical summary of the relationships between the measures within our sample. An important goal for future research is to collect data using a much larger sample of texts with a view to identifying a factor structure and scales that can be generalized across studies. One consideration that we have had in mind in the course of developing the different measures described here has been to consider how readily they can be automatically assessed and hence would be practical to use in collecting large samples of data.
The three main components that we have identified correspond fairly directly to different components of the writing process: Planned sentence production and within-sentence revision appear to capture the extent to which different writers plan and revise at the sentence level; revision of global structure appears to capture the extent to which different writers move back and forth in the text, altering the linear ordering of sentences. An important feature of these components is that they are orthogonal to one another. In other words, writers who score highly on planned sentence production do not necessarily score higher or lower than other writers on the other components. These measures should therefore enable us to identify writers who combine planned/unplanned text production with immediate revision or not and assess how this relates to other variables. At the same time, the combination of complementary measures in higher-order dimensions helps consolidate our interpretations in a way that would not be possible with single measures alone.
Conclusion
We hope that we have demonstrated that the procedures and analytic techniques that we have described here help to improve the alignment between measures derived from keystrokes and potential underlying processes. We do not claim to have exhausted all the possible ways in which this could be done. In particular, we think that a range of more fine-grained analyses of revision could be carried out. However, we do think that a productive guiding principle is the general strategy of sorting the keystroke log according to actions in relation to the final product and then classifying measures of pauses, bursts, and revisions either statistically or conceptually—a principle that goes some way to addressing Schilperoord’s doubts about the possibility of analyzing text production in a context replete with intensive problem solving and massive editing.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: We gratefully acknowledge the European Union office for Cooperation in Science and Technology (COST) who provided funding through COST Action IS0703 - European Research Network for Learning to Write Effectively – to support research exchanges between the authors.
