Abstract
Web sites increasingly encourage users to provide comments on the quality of the content by clicking on a feedback button and filling out a feedback form. Little is known about users’ abilities to provide such feedback. To guide the development of evaluation tools, this study examines to what extent users with various background characteristics are able to provide useful comments on informational Web sites. Results show that it is important to keep the feedback tools both simple and attractive so that users will be able and willing to provide useful feedback on Web site pages.
Organizations increasingly recognize the importance of giving the user a voice, and many Web sites contain a feedback option that invites users to comment on various aspects of the Web site. In this article, we focus on evaluation methods that enable users to give feedback on specific pages of a Web site. Such methods have received little attention in the literature about Web site evaluation methods, and there is no generally accepted term for this type of method yet. These methods, however, can be categorized as self-reported metrics (Tullis & Albert, 2008) because, as in surveys, users are asked about their experiences on the Web site. But unlike surveys, the methods we focus on ask for page-level feedback and allow for open comments, sometimes combined with an overall rating or some scale questions. We propose to call these methods user page reviews because they invite users to review a Web site by clicking on a button that appears on selected pages. In such reviews, users evaluate a Web site in much the same way as experts evaluate a Web site (Welle Donker-Kuijer, De Jong, & Lentz, 2008). Although users cannot be expected to have professional expertise about Web design, they can provide feedback from their own perspective about their own attitudes and experiences with the Web site.
Tools that can be used for gathering user feedback on Web site pages include Opinionlab, Kampyle, Usabilla, and Infocus. These instruments enable Web site visitors to share their opinions on everything they consider important. Selected Web site pages (or sometimes all pages) contain a button that users can click on if they want to react to something. These buttons can be small icons (e.g., a thumb or a plus–minus icon with the word feedback) or longer text links that invite reactions to the page. Users click on the link to open a screen on which they can provide their comments. Users can give open-ended comments, but the instruments often also ask users to choose a feedback category, provide ratings, or answer questions about page-specific topics. Figures 1-4 show screen shots of feedback forms from Opinionlab, Kampyle, Infocus, and Usabilla, respectively.

Opinionlab feedback form.

Kampyle feedback form.

Infocus feedback form.

Usabilla add-note option and invitation to users to click on elements they like.
The Opinionlab form in Figure 1 looks rather dense and asks users to complete several tasks: to choose a topic from predefined categories, enter an open comment, rate the page on three aspects as well as overall, enter an e-mail address (which is optional), and indicate whether or not their comment is about the Web site. The Kampyle form in Figure 2 uses icons that users select to express their feelings and categorize their feedback. It asks users to rate the site by choosing an emoticon, to select a feedback topic (under the topics that are visible in the figure are rows with subtopics), and to fill in an open comment. The Infocus form (see Figure 3) includes a list of predefined categories, a place to formulate comments, and three options for marking specific elements or segments: Users can underline, point an arrow at, or draw a frame around a relevant section that they want to comment on. This marking function is different from most other tools, in which users have to describe the exact location of the object of their feedback. Only Usabilla (see Figure 4) offers a form of marking that enables users to add a note on the Web page. But it is not possible to mark the exact size of the selection or to use other more precise markings. Besides these four feedback tools, many examples of feedback buttons can be found on Web sites.
Even though such methods for obtaining user feedback on Web site pages have become more and more popular in practice, little is known about the merits and limitations of these methods. The literature on methods of user-focused Web site evaluation has focused strongly on the use of think-aloud usability testing and surveys and neglected the possibilities of asking users to review a Web site (e.g., Cunliffe, 2000; Tullis & Albert, 2008). The user page reviews strongly depend on users’ skills in providing feedback and on their willingness to do so. Users may consider giving feedback as an extra task in addition to what they are doing on the Web site. In this article, we focus on the skills that are needed to provide feedback, examining the extent to which users are able to provide comments on a Web site. More knowledge about the users’ skills may provide useful information about the value of the output these review methods yield. Furthermore, this research may contribute to our knowledge about how to design effective user feedback tools. Before discussing our research questions, we address related work on the skills that users need to adequately provide comments on Web site pages.
Related Work
Although it may be important to give users the opportunity to provide feedback, they must be able to express that feedback. Users need some critical capacity to signal the problems they are experiencing with a Web site before they can report these problems. They need to be able to monitor their thinking processes and behavior. This critical skill is closely related to the concept of metacognition, which has been described as “one’s knowledge and beliefs about one’s own cognitive processes and one’s resulting attempts to regulate those cognitive processes to maximize learning and memory” (Ormrod, 2006). Literature on metacognition and comprehension monitoring (e.g., Baker, 1989) shows that readers with higher verbal abilities have greater awareness and control of their own cognitive activities while reading than do readers with lower verbal abilities. Readers with higher verbal abilities also appear to use a greater diversity of evaluation standards in their assessment of text quality.
Once users are able to signal their own problems with a Web page, the next step is that they attribute these problems to the Web site. The self-serving bias predicts that users tend to blame the Web site when they fail to reach their goals (Moon, 2003; Serenko, 2007). But several studies have shown that users sometimes blame themselves for problems they encounter. Schriver (1997) concluded that users of manuals for complicated home electronics predominantly blamed themselves for their problems with the described products. In a follow-up study, Jansen and Balijon (2002) found that a high percentage of the respondents (more than 50%) reported blaming themselves and not the manual when something went wrong. And Serenko (2007), who studied “interface agents” (software systems that support users on the computer), found that users may attribute their success to an interface agent and hold themselves responsible for task failure. Thus, users who experience problems and signal these problems may not attribute them to the Web site and therefore may not report them with the feedback option.
The final step in the feedback process is to submit the feedback by formulating the problem on a feedback form. Besides describing their feedback, users are often asked for additional information, such as the category of the comment or answers to scaled questions. Asking users to formulate feedback is not common. Some usability specialists, such as Nielsen (2001) and Spyridakis, Wei, Barrick, Cuddihy, and Maust (2005), have argued that users’ self-reports may be unreliable. Others, such as Sauro (2010), considered users’ self-reports to be an effective evaluation method in addition to traditional user testing or heuristic evaluations. In certain contexts, user feedback might be very useful. Nichols, McKay, and Twidale (2003) advocated empowering end users to proactively contribute to usability activities with end-user reporting tools. Karahasanovíc, Nyhamar Hinkel, Sjø´berg, and Thomas (2009) described a feedback-collection method that asked for written feedback at different times during an experiment. They compared this method with concurrent and retrospective think-aloud protocols and concluded that the feedback method revealed more difficulties than did the concurrent and retrospective think-aloud protocols and that the feedback method seemed better at identifying difficulties that prevented progress or caused significant delay. Castillo, Hartson, and Hix (1998) found that users are able to identify and report their own critical incidents, but these studies involved highly educated participants. Earlier research on the evaluation of documents has shown that highly educated people provide more comments and more diverse feedback than do people who are less educated (De Jong & Schellens, 2001). If less educated or inexperienced users are indeed less able to formulate their feedback, user page reviews will probably fail to uncover some relevant problems.
All these steps in the feedback process require that users perform two tasks at the same time: finding and using the information while reviewing a Web page. Switching between these tasks is difficult for users because it requires a configuration of mental resources. Monsell (2003) provided an overview of studies on task switching, which, for example, show that task switching leads to longer response times and a higher error rate. Users are inclined to concentrate on one task and to neglect other tasks, a phenomenon that is called “cognitive lockup” (Neerincx, Lindenberg, & Pemberton, 2001). Older people especially seem to have trouble switching between tasks (Kramer, Hahn, & Gopher, 1999). A study by Barkas-Avila, Oberholzer, Schmutz, De Vito, and Opwis (2007) of error messages for users who fill out a form online shows practical consequences of problems with task switching. Their results show that error feedback should be provided only after users have completed the whole form because providing error feedback immediately negatively affects users’ performance. This study also underlines that trying to perform additional tasks while completing the primary task may cause users to expend too many additional cognitive resources, resulting in cognitive overload. This difficulty of having to switch between two different tasks is likely present in all user page reviews. The same difficulty has been reported in think-aloud usability testing in which users may stop verbalizing thoughts when they encounter difficulties during their task completion (Boren & Ramey, 2000). Also, verbalizing thoughts during task performance can lead to reactivity: Participants perform worse in concurrent think-aloud conditions than in conditions in which they think aloud afterward (Elling, Lentz, & De Jong, 2011; Van den Haak, De Jong, & Schellens, 2003, 2004).
Research Questions
In this study, we focus on the extent to which users are able to productively formulate review comments on Web sites. We conducted the study in a controlled context in which users with various background characteristics provided feedback. Our research questions focused on (1) the number and characteristics of the users’ comments, (2) the consistency of the users’ comments in relation to the opinions they provided in a questionnaire, and (3) the consistency between the users’ comments.
What Are the Numbers and Characteristics of Users’ Comments?
Our first research question comprises several subquestions. We first looked at the number and types of user problems that our study generated:
1a. To what extent are participants able to provide comments on the relevant issues for informational Web sites (navigation, textual and visual content, and design)?
Then, we investigated whether differences in the participants’ backgrounds (age, education level) correlated with differences in the characteristics of their feedback:
1b. Are there any age- or education-related differences between participants’ user page reviews?
To get more detailed information about the characteristics of the comments, we looked at the participants’ descriptions of their comments, their use of the marking options included in the tool, and their categorization of the comments. Thus, we addressed the following three subquestions:
1c. How do participants formulate their comments on a Web site?
1d. To what extent do participants use the problem-marking function?
1e. How well are participants able to categorize their feedback?
The main purpose of Web site evaluation methods is to provide insight into the way users interact with Web sites and to detect and diagnose problems that they encounter. We examined, then, the extent to which participants’ comments pointed to usability problems on the Web site and the severity of these problems. In other words, we addressed the following subquestion:
1f. How useful is the feedback that the participants provide?
Are Users’ Comments Consistent With the Opinions They Provided on the Questionnaire?
Our second research question focuses on users’ overall consistency in their judgments about a Web site. To address this question, we analyzed the correspondence between the participants’ user page review comments and the results of a questionnaire that we asked them to complete. We expected that users who reported many problems on the Web site would express stronger negative opinions in the questionnaire. Thus, we investigated the following question:
1. To what extent do participants’ comments on specific Web pages relate to their overall opinions about the Web site?
Do Users Report the Same Problems?
User page reviews ask for local and open-ended feedback and can therefore generate very diverse user comments because all users have their own opinions about different aspects of the Web site. Moreover, an evaluation on a local level generates more specific user problems than does a global evaluation. These characteristics of user page reviews may make it difficult to realize a high overlap in comments between participants. Nevertheless, when all the users review the same pages and use the same scenarios, we can expect some consistency in the results. Our last research question, then, is this:
2. To what extent do participants report the same problems on a Web site?
Method
In this study, 93 participants provided feedback on three Web sites. Two Web sites were evaluated by 30 participants each, and one Web site was evaluated by 33 participants. We obtained our institutions’ approval to conduct human-subjects research. The participants, who received financial compensation for taking part in the study, were recruited by a specialized agency that has an extensive database of potential research participants. All the participants indicated that they use the Internet at least once a week. Men and women were almost equally distributed in the groups, with 43 men and 50 women participating in the study. Participants’ ages ranged from 18 to 71 (with an average age of 44). The participants were divided into four different age categories (18- 29, 30- 39, 40- 54, and 55 and older) and three education levels (low, medium, and high), based on the highest level of education that they had completed. Participants in the low-level educational group ranged from having an elementary school education to a junior general secondary education. Those in the medium-level educational group had an intermediate vocational education, a senior general secondary education, or a preuniversity education. Those in the high-level educational group had a higher vocational or university education. All groups were almost equally represented in the three Web site evaluations. Characteristics were mixed in such a way that, for example, all age groups consisted of nearly equal numbers of men and women of different educational levels. Table 1 provides an overview of the participants’ background characteristics.
Background Characteristics of the Participants in Each of the Three User Page Review Studies
Web sites
The participants evaluated three Web sites of medium to large Dutch municipalities: www.apeldoorn.nl (Web site 1), www.dordrecht.nl (Web site 2), and www.nijmegen.nl (Web site 3). Municipal Web sites provide information and services to citizens and to other interested users, such as tourists and businesses. These Web sites invariably contain a variety of information because they are designed to satisfy the informational needs of a broad target group. The three Web sites cover similar types of information, but this information is structured and presented in different ways.
Scenario Tasks
The participants evaluated the Web sites based on two scenario tasks that they were given. The content of the tasks differed per Web site, but they resembled each other in difficulty and length of navigation path. Moreover, all the tasks (a) covered realistic activities that correspond to those that users usually perform on municipal Web sites; (b) included searching for information, as well as reading, understanding, and applying the relevant information to the described scenario; and (c) applied to different domains of the Web site. For example, a task pertaining to information on a government subsidy supporting people who are buying a house for the first time asked participants to search for answers to three questions: (1) What is the name of this subsidy? (2) Do you meet the requirements for this subsidy? and (3) What should you do to make a request for this subsidy? The shortest navigation path to this information consisted of six links, the last of which was a link to a pdf file containing information about all kinds of subsidies available for buying or renovating a house.
Tool for User Page Reviews
In our study, we used the Infocus tool 1 for gathering user feedback. This feedback tool is an extension of the plus–minus method (De Jong, 1998; Sienot, 1997) and Focus (De Jong & Lentz, 2001), both of which were developed to evaluate paper documents. With the Infocus tool, users can surf through a Web site and click on a comments button whenever they want to give positive or negative feedback. The tool enables participants to select and comment on any element on a Web screen, ranging from specific words, sentences, paragraphs, chapters, illustrations, and navigational aids to the entire Web page. After clicking the comments button, participants see a screen shot of the Web page. They have several options for marking the specific element that they want to comment on: They can point to it by drawing an arrow, underline it, or add a frame around it. On the left side of the screen, participants can choose one of the predefined comment categories and type in their feedback. Users’ feedback is saved in a database together with images and URLs of the marked Web screen and some log data such as the navigational path and a time outline.
Like many other tools for user page reviews, Infocus is able to ask users to categorize their feedback. In this study, we used four categories: (a) navigation (easy or difficult to find), (b) content (clear or unclear information), (c) design (does or does not look good), and (d) other. This categorization has three potential benefits. First, the availability of comment categories may remind participants of their reviewing task. The comment categories emphasize that we are asking participants to provide feedback on the Web site and not, for example, answers on tasks. Second, the specific categories chosen may guide the kinds of feedback participants provide. By explaining the categories and including them on the feedback screen, we are underlining which aspects of Web site quality we consider important to evaluate. Third, the comment categories may serve as clues to help us to interpret unclear comments when we analyze the data.
Posttest Questionnaire
After the review process, participants filled out a questionnaire. The first part asked users for demographic information and about their experiences reviewing the Web site. The second part asked users for their opinions about the Web site. This Web Site Evaluation Questionnaire (WEQ), which we developed specifically to evaluate informational Web sites such as municipal Web sites, focuses on three dimensions: navigation, content, and design (Elling, Lentz & De Jong, 2007, 2012). The WEQ consists of 25 questions to which participants provide their responses on a 5-point Likert-type scale.
Procedure
We conducted our study in a laboratory setting using Infocus. We chose this controlled environment in order to get more insight into the users’ abilities to provide feedback and to compare the results of different groups of users who reviewed the same pages of a Web site. The four evaluation sessions took place in a computer lab with workstations for small groups (seven to nine participants). Each session lasted 90 minutes. To reduce potential work tempo differences, we invited participants to work in rather homogeneous group settings. One group consisted of participants who were older than 55 years, and the other three groups were distinguished by participants’ level of education (low, medium, and high). We made this division because we expected that older and less educated participants would need more time to complete the evaluation.
To reduce the participants’ cognitive load, we split the tasks of navigating and reviewing. We had two solid reasons for doing so: First, research has shown that users are not good at combining different tasks. Second, in a pilot study, we observed that participants appeared to struggle with trying to carry out scenario tasks and review the Web site at the same time. They concentrated on finding the right answers to the scenario tasks and often forgot to provide feedback on the problems they encountered. Our observations are in line with “cognitive lockup,” the idea that users concentrate on one task and neglect other tasks (Neerincx et al., 2001). In our pilot study, participants who had trouble finding or understanding the information, in particular, seemed to have no cognitive energy left to comment on their experiences. To ensure that participants could use all their cognitive energy in reviewing the Web site, we prioritized the reviewer role and reduced the requirements for participants to perform tasks on the Web site. Of course, this somewhat artificial sequential procedure deviates from the online way of reviewing, but it enables all the users to provide feedback and answer our main research question about users’ abilities to provide feedback on a Web site.
The review session started with a brief introduction in which participants were told that we investigated municipal Web sites to improve them and that we needed their help to find the problems that users might experience. Then, we illustrated the functions of Infocus and explained the four feedback categories (navigation, content, design, and other). The facilitator described how to formulate a comment, mark something on a screen shot, choose a feedback category, and save the comment while the participants followed each step on their screens. Participants were then given the opportunity to practice on a tourist Web site and to write and save some comments. After this practice session, the facilitator asked participants to open the Web site to be evaluated and to surf freely on the Web site for 10 minutes while providing comments with Infocus. After this, the guided review procedure started in which participants received two task scenarios that reflected the kinds of tasks users perform on municipal Web sites. After giving them the task scenarios, the facilitator asked them to open the home page of the Web site and decide for themselves which link they would choose on the home page given the first task description. The facilitator then showed the group the link that led to the required information and asked the participants to give feedback about the home page. Participants formulated their comments individually on their computers.
When everyone had completed their feedback, the participants were asked to click on the link and to look at the new page to decide what the next step should be. Then, the facilitator again showed the link that led to the information and asked each of them to provide feedback on this page. If there were alternative options to reach the information, the facilitator explained these options before continuing on one of the possible routes. When the participants reached the final information section, they used that section to think about the answer to the main question in the task description though they were not required to write down the answer. After a while, the facilitator gave them the answer to the scenario question and asked them to provide feedback on that page. The participants were allowed 5 minutes to complete each step of the task, including searching, thinking, and formulating a comment. When they had finished the task, the participants were asked to provide feedback on the whole process of task completion. After finishing both tasks, the participants filled out the questionnaire.
Analysis
To answer our research subquestion 1a, we determined the number of comments made per participant, the proportion of positive versus negative comments, and the number of unclear comments. We also categorized the comments using the four feedback categories (navigation, content, design, and other). The first author and an independent rater coded all the comments, achieving a satisfactory Cohen’s κ (.81).
Next, we analyzed the negative feedback. As Hornbæk (2010) argued, the matching of comments into a list of user problems is not straightforward and can be done in different ways. In our matching procedure, we used a four-step analysis. First, we gathered the comments that were made on the same page; second, we divided the feedback into comments on different parts of the page (menu, illustration, text block, etc.); third, we analyzed these groups of comments using the categories; and fourth, we analyzed these comments using the participants’ exact description. The description of the cause of the problem was decisive for merging two comments into one user problem. If, for example, two comments were about the bad legibility of the text, we would consider them to be two different problems if one comment criticized the small font and the other the color of the text. This rather strict and fine-grained method of matching resulted in a large list of different user problems with relatively little overlap.
To answer subquestion 1b about the influence of user characteristics such as age and education level, we analyzed the comments for differences between the four age categories and the three education levels using analysis of variance. We used four dependent variables: overall number of comments, number of positive comments, number of negative comments, and number of unclear comments. To answer subquestions 1c (about formulation), 1d (about marking), and 1e (about feedback categorization), we used a detailed analysis of the comments as they were collected in the database. We qualitatively analyzed the ways in which participants formulated their comments. Then we analyzed the degree to which participants used the various marking options and determined the quality of their categorizations by comparing the participants’ categorizations with those of the two expert coders.
To answer subquestions 1f, concerning the severity of the problems, two independent experts rated the severity of the comments on the two Web pages with most comments of each of the three Web sites. In total, they rated 88 unique comments (24% of the whole set of 373), 18 comments from Web site 1, 26 from Web site 2, and 44 from Web site 3. Many studies have shown that experts are not good at rating the seriousness of problems and that their severity judgments are highly personal with little agreement between different experts who rate the same set of problems (e.g., Hassenzahl, 2000; Hertzum & Jacobsen, 2003; Hertzum, Jacobsen, & Molich, 2002; Molich & Dumas, 2008; Nielsen, 1995; Uldall-Espersen, Frøkjær, & Hornbæk, 2008). We should therefore interpret the experts’ severity ratings with caution. We asked the experts to determine the seriousness of the problems described in the comments, looking at their impact on the users’ task performance. Following the user-focused part of the rating that Uldall-Espersen, Frøkjær, and Hornbæk (2008) used, we asked the experts to rate the problems according to three severity levels: level 1, cosmetic comments about problems that do not interfere with users’ task performance; level 2, critical comments that point to problems that annoy users and can in any way disturb their task performance; and level 3, catastrophic comments that may cause users to give up their task. The experts received information about the tasks the participants performed, and they had screen shots of the pages to which the comments referred. After the experts rated the comments on the first Web site, we discussed the comments that were rated differently by both experts and strengthened the agreement about the interpretation of the three severity levels for rating the remaining comments. Then the experts independently rated the comments on the other two Web sites. The Cohen’s κ score of the two independent experts was .66, indicating that they had substantial agreement (76% of the comments were rated identically by both experts). The experts discussed which severity score to choose for those comments that they rated differently.
To answer question 2, concerning the relationship between the participants’ page-review comments and their overall opinions about the Web site, we analyzed the correlation between the number of negative comments per dimension (i.e., navigation, content, and design) and the participants’ score on the corresponding dimension in the WEQ. For this analysis, we used the experts’ categorization because this categorization is most adequately related to the questionnaire. The reliability of the participants’ WEQ scores was measured using a Cronbach’s α. The navigation dimension included 13 items with an α of .93, the content dimension had 9 items with an α of .82, and the design dimension included 3 items with an α of .88.
To answer research question 3, concerning the consistency between users’ comments, we first analyzed the relationship between the sample size and the number of problems detected. This relationship has been investigated for several other evaluation methods, in particular, heuristic evaluation and think-aloud usability testing. Some of these previous studies reported the encouraging outcome that five or six participants would suffice to detect the majority of the problems (Nielsen, 1994a, 1994b; Virzi, 1992) whereas other studies showed that larger samples were needed to obtain more or less exhaustive results (Faulkner, 2003; Lewis, 1994; Spool & Schroeder, 2001). Of course, this relationship varies depending on the type of evaluation method. User page reviews are likely to generate many different comments, making it difficult, if not impossible, to find a complete set of problems. For this analysis, we used a Monte Carlo procedure (Lewis, 1994; Nielsen, 1994a; Virzi, 1992). For each sample size, ranging from 1 to N-1, we took 1,000 random subsamples (with replacements) and computed the mean number of different problems detected. To facilitate our comparison between the three Web sites, we expressed the results in terms of percentages of the total number of problems detected by a sample of 30 participants. To further explore the stability of the results, we used an adaptation of the Monte Carlo procedure. For all possible sample sizes (ranging from 1 to N-2), we randomly selected 1,000 sets of two independent subsamples (with replacements). We computed the percentage of agreement between two subsamples by dividing the number of shared problems in both subsamples by the total list of problems in either subsample.
Results
The review sessions took place in a positive atmosphere. Participants seemed to feel motivated to evaluate the Web site. As we expected, the sessions with participants who were 55 years of age and older took more time because these older participants needed more help and instructions about how to make comments than did younger participants. Contrary to our expectations, less educated participants sometimes indicated that the sessions took too much time: They did not need all the time given to make comments because “everything was clear” to them. Overall, the participants judged their experiences and their review abilities positively. They reported that the review task was easy to do (mean score on a 5-point scale: 4.03, SD = 0.80) and that they enjoyed the process of giving comments on Web pages (mean score: 4.15, SD = 0.82).
Number and Characteristics of Users’ Comments
To answer our first research question we addressed subquestions on the extent to which the participants were able to provide comments on relevant issues, whether there were age- or education-related differences between participants’ page reviews, how participants formulated their comments on a web page, the extent to which they used the marking function, how well they categorized their feedback, and how useful their feedback was.
To What Extent Are Participants Able to Provide Comments on the Relevant Issues for Informational Web sites?
On average, each of the participants provided 16.1 comments during a session (SD = 6.1). Of these comments, an average of 0.8 (SD = 1.3) were judged to be unclear (5%). Of the clear comments, an average of 9.5 (SD = 5.6) were negative (62%). If we consider the negative comments to be the core of the evaluation, the net result of the user page-review evaluation is 9.5 potential user problems per participant. Because many of the negative comments may point to the same potential user problem, we also looked at the number of different potential problems per Web site. This analysis resulted in, respectively, 107, 110, and 156 potential user problems for the three Web sites.
Table 2 gives an overview of the categories of the total set of negative comments on the three municipal Web sites (based on the two coders’ categorization). Navigation appears to be an important aspect of the participants’ feedback, but many comments were also made about content and design. Whereas the proportion of navigation-related comments was nearly stable across the three Web sites, the proportion of content and design comments varied between the three sites, with relatively more comments about content in Web sites 1 and 3 and design in Web site 2. Web site 2, for example, had a lot of comments on the inconvenient arrangement of the home page and on poor legibility due to the size and color of the fonts. The other Web sites had more content comments, such as long texts with complicated and vaguely formulated information. The category other included feedback on technical problems (e.g., a picture that did not appear correctly), redundant information (e.g., an item on the home page with detailed information about clearing away leaves in autumn), or missing functionalities (e.g., a form that was not available online).
Number of Negative Comments per Category for Each of the Three Web sites
In all, we can conclude that users were able to report many comments and that they succeeded in addressing navigation, content, and design issues.
Are There Any Age- or Education-Related Differences Between Participants’ User Page Reviews?
Differences between the four age categories and the three education levels were analyzed in regard to four dependent variables: total number of comments, number of positive comments, number of negative comments, and number of unclear comments. The results are shown in Tables 3 and 4 . Regarding the total number of comments made, we found no significant differences between the four age groups and between the three education levels, for age: F(3, 89) = .443, p = .723; for educational level: F(2, 90) = 2.839, p = .064. Also, no significant interaction effect was found. The same applied to the number of positive comments, for age: F(3, 89) = .484, p = .694; for educational level: F(2, 90) = .372, p = .690. For the other two dependent variables, however, significant education-related differences were found, for the number of negative comments: F(2, 90) = 6.818, p < .01; for the number of unclear comments: F(2, 90) = 5.725, p < .01. Post hoc analyses showed that participants with a high education level produced more negative comments than did participants with a low or medium education level and that the participants with a higher level of education produced fewer unclear comments than did participants with a lower level of education.
Mean Number of All Comments, Positive Comments, Negative Comments, and Unclear Comments for the Three Educational Levels (SD)
* p < .05.
Mean Number of All Comments, Positive Comments, Negative Comments, and Unclear Comments for the Four Age Groups (SD)
In all, the user page review appeared to be usable for all age groups and all education levels included in our study. When we focus on quality of feedback, however, higher educated participants were better able to provide feedback, producing more negative comments (which could serve as clues to improve the Web site) and fewer unclear comments.
How Do Participants Formulate Their Comments on a Web site?
Looking at the participants’ comments in more detail, we made some interesting observations about the content of these comments. First, participants differed in the ways they formulated their comments, which could be written in the form of reporting a problem’s cause, effect, or solution. For example, on one of the Web pages, participants had difficulty navigating back to the home page due to the small size of the home icon. Several participants reported this problem by describing its effect (“I can’t find my way back to the home page”). Others reported its cause (“The icon of the homepage is too small and inconspicuous”) or a possible solution (“This icon should get more attention”). Individual participants sometimes reported both the cause and the effect and occasionally a solution. These differences in style of reporting sometimes made it difficult to categorize the comments. In the preceding example, the cause of the reported problem was a design issue, but the effect concerns navigation: not finding the way back to the home page. The experts systematically categorized on the cause of the problem unless the cause was not mentioned by the participant.
Second, the participants sometimes used the comment button to provide an answer on the scenario task. Participants who used the feedback option to comment on the scenario or to formulate answers to scenario questions did not understand, or forgot, the reviewing task at hand. We did not find differences in the number of scenario-related comments between the education levels, F(2, 90) = 2.215, p = .115). On average, participants in all the groups entered between 0.4 and 1.0 scenario-related comment. There was, however, a significant difference between the age groups, F(3, 89) = 2.794, p < .05: Older participants made significantly more scenario-related comments (1.0, SD = 1.4) than did younger participants (0.29, SD = 0.7). Older participants perhaps experience more trouble understanding the task of reviewing a Web site and separating this task from searching for information and answering scenario questions.
Third, some participants repeated the same comment several times. For example, one participant reported that the Web site did not fill the whole screen. He repeated this comment in nearly the same wording on several Web pages. Perhaps, the structure of the guided review procedure pressures participants to comment even when they are uncertain of what to say. Repeating the same comment could be a reaction to this pressure. Other participants wrote down a positive remark (e.g., “This is clear to me”) when they did not see any problem.
Fourth, participants sometimes made more than one comment in one feedback screen. In the instructions, we asked participants to make a new annotation for every comment they wanted to report. Nevertheless, they sometimes combined two comments, often a positive and a negative statement (e.g., “The home page looks up to date, but it is difficult to find what you look for on this page”). In doing so, participants may have been trying to soften their negative feedback by saying something nice about the Web page as well. Giving negative feedback can be seen as a “face threatening act” (Brown & Levinson, 1987) that people want to compensate for by showing appreciation and respect.
Finally, the older participants often referred to using another medium in their feedback (e.g., “I would have used the telephone by now, or It is a lot easier to go and ask for the information”). They seemed to feel less familiar with the Internet than the younger participants did and, as a result, seemed to prefer other media, such as the telephone or face-to-face contact. This observation corresponds with the research finding that people need time to get accustomed to new media (Van Dijk, Pieterson, Van Deursen, & Ebbers, 2007).
To What Extent Do Participants Use the Problem-Marking Function?
The marking function was introduced to help users adequately indicate the object of their feedback. A verbal description of the location of an object on the screen is far less specific and effective than a picture of the specific page with, for example, an arrow pointing at that object. Of all the participants’ negative comments, 51% were accompanied by some form of marking—a reasonably good result, considering that most of the participants were introduced to this functionality for the first time. So participants seem to be able to mark the object of their feedback, which makes this marking functionality a useful innovation for the user page reviews.
Of course, for some comments, the absence of a marking could be explained by the nature of the comment. In three situations, marking an object was impossible. First, some comments were about the whole Web page (e.g., It’s strange that the design of this page deviates from the design of the other parts of the Web site). Second, some comments were about items that were missing on the page or that could not be found (e.g., I don’t know what link I should choose here). And third, some comments did not even pertain to any object on the Web page (e.g., I had to click too much to reach this information).
How Well Are Participants Able to Categorize Their Feedback?
Participants categorized all their feedback into the four categories (i.e., content, navigation, design, and other). Also, one of the authors and an independent coder categorized all the participants’ feedback. This categorization resulted in a Cohen’s κ of .81, indicating a high agreement between the two expert coders. But the agreement between the participants and the expert coders was considerably lower, with a κ of .26. The expert coders categorized 51% of the comments differently than the participants had.
Content was the most problematic category. Participants tended to put not only comments about content in this category but also comments about unclear labels for links and other navigational problems. The expert coders transferred as many as 191 of these comments into the navigation category. These differences in coding choices between the participants and the experts do not mean that the participants did not thoughtfully categorize their problems. The categories had rather open and broad formulations, so many different comments fit into these categories. This result raises questions about users’ ability to categorize their feedback. In our study, we explained the categories to the participants, providing examples, and they could ask questions if they did not understand the categories. In an online remote setting, users would not have these opportunities, which would probably result in users’ having even more difficulty categorizing comments than what they would have in a laboratory setting.
How Useful Is the Feedback That the Participants Provide?
A sample of 88 negative comments were rated on severity by two independent experts. Of these comments, 21 (24%) were rated as cosmetic problems, 48 comments (55%) were rated as critical problems, and 19 comments (22%) were rated as catastrophic. In other words, 77% of the negative comments pointed to problems that disturbed participants or even prevented them from reaching their goals on the Web site. Examples of cosmetic comments are “This home page is boring,” and “The design of this page is not inviting.” Examples of critical comments are “It is not easy to find the relevant link on this page,” and “The navigation path to the information is too laborious.” Examples of catastrophic comments are “The information about school holidays cannot be found because of the misleading link label,” and “Crucial information about the procedure is missing in this text.” Thus, participants were able to provide useful feedback, with the majority of their comments referring to critical or catastrophic usability problems on the Web site that were related to the scenario tasks.
To What Extent Do Participants’ Comments on Specific Web Pages Relate to Their Overall Opinions About the Web site?
If participants report many problems in a certain category, we would expect that the overall score for the corresponding dimension in the WEQ would be low. Table 5 displays the Pearson correlations between the number of problems detected in three Infocus categories (we did not include the category other) and the scores for the corresponding WEQ dimensions. On all three aspects, we found significant but weak negative correlations between the number of negative comments and the scores in the questionnaire: The more negative comments that participants made in a category, the lower the scores for the corresponding WEQ dimension. Although the numbers of content and design comments correlated significantly with their corresponding WEQ dimensions, the number of navigation comments correlated significantly with each of the three WEQ dimensions. This finding may be attributed to the fact that navigation problems may be caused by unclear link names or visual cues. In all, these correlations provide some initial support for the consistency of the judgments underlying the user page review comments. The two evaluation methods show similar tendencies in the participants’ opinions about the quality of the Web site. The participants’ detailed comments in the user page review corresponded to their more global evaluations about using the Web site. This result means that participants were able to provide stable feedback.
Correlations Between the Number of Problems Detected in Three Infocus Categories and the Overall Scores for the WEQ Dimensions
Note. WEQ = Web site Evaluation Questionnaire.
*p < .05, **p < .01.
To What Extent Do Participants Report the Same Problems on a Web site?
Using a Monte Carlo analysis, we explored the relationship between the sample size and the exhaustiveness of the list of user problems. Figure 5 displays a graph of this relationship. The graph represents the mean percentage of problems detected (of the total number of problems detected for that site by 30 participants) for sample sizes ranging from 1 to 30. All three Web sites show a remarkable similarity. The patterns show little dependence on the particular characteristics of the Web site evaluated. None of the three patterns shows a clear curve toward a finite set of comments. Based on the graph, then, we must conclude that additional participants would likely detect new problems. Our analysis of the agreement between two independent samples shows that two samples of 15 participants have an agreement between 35% and 48%. To obtain results with at least 60% agreement between two samples, a sample size of between 20 and 25 participants would be needed. In sum, to a certain extent, the participants were able to point to the same problems, but they also showed a clear diversity in the comments they reported.
Relationship between the number of participants and the percentage of problems detected in the three municipal Web sites (N = 30).
Discussion
In this article, we have argued that it is important to enable users to give feedback on specific pages of Web sites. We introduced user page reviews (e.g., Opinionlab, Usabilla, Kampyle, and Infocus) as a category of remote evaluation methods. In our study, we examined the extent to which users are capable of giving adequate written feedback in open-ended comments, focusing on the numbers and characteristics of users’ comments, the consistency of users’ feedback and their questionnaire opinions, and the correspondence between users’ comments.
Numbers and Characteristics of Users’ Comments
The results of our study are encouraging. Participants reported that they liked to produce the user page reviews, and they did not have any substantial problems with providing feedback. Overall, the number of unclear comments was small (5%). Highly educated participants produced more adequate feedback (in terms of both potential clues for revision of the Web site and clarity) than did less educated participants. This finding is in line with research that shows that higher educated people are better able to monitor their own cognitive activities (Baker, 1989) and provide more comments and more diverse feedback (De Jong & Schellens, 2001). This result also corresponds with studies by Karahasanovíc et al. (2009) and Castillo et al. (1998) that reported positive findings on providing feedback by higher educated users. In our study, we also found that older participants and participants with lower education levels were capable of providing feedback on relevant dimensions.
The participants provided substantial feedback on the content and design of these informational Web sites. Additionally, they detected many problems that involved navigational issues. The large number of navigation comments may have been a result of the step-by-step procedure we used, which could have triggered attention to the quality of specific links. But navigation may just generally be an important aspect of users’ perception of Web site quality. Future research combining user page reviews with free surfing tasks might clarify this issue.
The marking function was used in 51% of the comments, so many participants were able to mark the object of their comments and thus easily identify the cause of their problem. But not all the comments could be related to a specific object on the page. And some participants perhaps did not mark the objects of their comments because they needed to get used to a new functionality before they could optimally use it. That may especially have been the case with our sample that included many older and less educated people. In an unpublished follow-up evaluation study that we conducted with younger, more highly educated users, the marking function was used substantially more often.
In both our pilot study and our main study, participants had difficulty categorizing their own feedback. This corresponds with Bruun, Gull, Hofmeister, and Stage’s finding(2009) that users had trouble categorizing the severity of their comments. These findings raise the question of whether it is feasible to ask lay people to categorize their comments. They may not be able to reflect on categories when they formulate feedback because the combined task of formulating and categorizing feedback—in addition to performing tasks on the Web site—may demand too much cognitive energy. Despite this difficulty, there are potential benefits to categorization, as we described earlier. First, the categories prompted the participants to perform the role of reviewer although the effectiveness of this prompt is difficult to determine. Second, the categories guided participants to the kind of feedback we desired. Most of the comments did indeed fit into the three categories of content, navigation, and design. But the third function, guiding the evaluator in the interpretation of the comments, was not achieved. The categories chosen by participants did not help us to interpret unclear comments. Moreover, in more than 50% of the comments, we did not accept the category that the participant had chosen because it did not match our definition of the category. This finding should be a warning for all the user page-review tools that ask users to categorize their comments. These user categorizations should be considered carefully. More research is needed on users’ ability to categorize their comments and on the types of categories that are best to use.
The severity rating of a selection of the reported problems shows that participants were able to provide useful feedback. Most (77%) of the negative comments pointed to problems that might disturb or prevent users’ adequate task performance. The expert rating provides a first indication of the usefulness of the comments that are reported in a user page review. Future studies could compare other evaluation approaches, such as think-aloud usability testing, to get more insight into the extent to which the comments correspond with the problems users experience when performing tasks.
Consistency of Users’ Feedback and Opinions
The correlations we found between the participants’ feedback and their overall scores on the questionnaire indicate that the detailed comments participants provided reflected their general attitude toward the Web site. Thus, participants were able to provide feedback that reflected their opinions.
Correspondence Between Users’ Comments
The Monte Carlo analysis showed that participants were somewhat able to point to the same problems. Due to the enormous diversity of problems that participants may mention, relatively large sample sizes may be necessary to collect a stable and more or less exhaustive list of problems, in line with studies by Lewis (1994), Spool and Schroeder (2001), and Faulkner (2003). Because the user page review asks for open comments on a local level, we could expect that this feature would result in a broad range of different comments. And our rather fine-grained way of matching increased this effect because we were cautious in considering two comments as referring to the same problem.
Conclusions and Future Research
This study shows that users are able to provide useful feedback in situations in which they can devote themselves entirely to the review task. However, a possible limitation of this study is that it shows only one way of evaluating a Web site with the user page review method: a step-by-step laboratory procedure with the Infocus tool. We chose this procedure because research shows that it is difficult for users to conduct tasks and provide feedback at the same time (Monsell, 2003; Neerincx et al., 2001). Older and less educated users, in particular, may have trouble switching between task processing and reviewing (Kramer et al., 1999). In our procedure, then, we chose to reduce the participants’ cognitive load by allowing them to focus on one task: reporting problems with the feedback option. In this way, we could optimally study the abilities of users to provide feedback. But the results of our study cannot be generalized to contexts in which users conduct more than one task. Further research is needed about producing feedback under different extents of cognitive load.
Also, the evaluation in this study does not resemble evaluations that are done in practice. Our goal was to study users’ abilities to provide feedback and not the evaluation method as a whole. Future research should test the user page reviews in more natural circumstances and compare the outcomes of the evaluation to those of observational methods, such as the think-aloud method, in order to obtain more information about the merits and restrictions of user page reviews.
The literature about problems with task switching raises questions about the online feedback tools. To what extent are online users able to provide adequate feedback? What are the characteristics of the users who provide feedback with online tools? Perhaps especially highly educated and experienced users are triggered to provide their feedback while other users need all their cognitive energy just to realize their primary goals on the Web site. It would be useful, then, to have more knowledge about the personal characteristics of the online respondents. Also, comparisons should be made between the feedback that users give with the online tools and in the laboratory: To what extent do these comments correspond with each other? We plan to address these questions in future research.
In this study, we have gained more knowledge about users’ abilities to provide feedback on Web site pages in a controlled context. Although our study showed that most users have the ability to provide feedback, evaluators in an online context should carefully interpret the results of the user page-review tools. Less educated users have more trouble reporting their problems and are possibly less inclined to share their feedback. Tools should have a clear and uncomplicated design. Also, to stimulate users to share their feedback, these tools need to be eye-catching, attractive, fast, and easy to use. Extra options, such as a categorization, scales, or ratings, should be used with reserve.
Footnotes
Notes
Acknowledgments
This article is based on a research project financed by the Dutch Organization for Scientific Research (NWO). It is part of the research program Evaluation of Municipal Web sites. The authors would like to thank master’s students Floris Baan and Philip Henssen, who contributed to the studies in this article. The authors would also like to thank the anonymous reviewers for their valuable comments on earlier versions of this article.
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The authors disclosed receipt of the financial support for the research, authorship, and/or publication of this article from the Dutch Organization for Scientific Research (NWO).
