Abstract
Recent literature on organizational category spanning demonstrates that organizations that span multiple categories on average suffer social and economic disadvantages in markets. While multiple mechanisms have been proposed to explain this finding, most studies do not test directly nor contrast these mechanisms. In this article, we contrast two of the main mechanisms proposed in the literature: the audience-side typicality-based explanation (category spanners are atypical of each categories spanned) and the producer-side quality-based explanation (category spanners produce lower quality output because they cannot develop expertise in any of the categories spanned). We find evidence for both mechanisms. Furthermore, we argue that quality and typicality interact such that high-quality organizations can benefit from being atypical. Finally, we contrast two kinds of spanning, “fusion” and “food court,” and argue that their effects are different depending on the overall quality of the organization. Our empirical setting is the restaurant domain, and we analyze menus and reviews of 474 restaurants located in San Francisco.
Introduction
An active line of contemporary research on organizations investigates the effects of category spanning on organizational outcomes. A now accepted empirical pattern shows that organizations that span multiple categories on average suffer from social and economic disadvantages in markets. Category spanners receive less attention and legitimacy and they have lower chances of success (Hsu et al., 2009; Zuckerman, 1999; Zuckerman et al., 2003). Organizations assigned to multiple categories tend to be either ignored (Zuckerman, 1999) or devalued (Hsu, 2006; Hsu et al., 2009; Kovács and Hannan, 2010, 2011; Rao et al., 2003). This pattern is shown to hold in domains such as stock recommendation (Zuckerman, 1999), films (Hsu, 2006; Zuckerman et al., 2003), wine producers (Negro et al., 2010, 2011), online auctions (Hsu et al., 2009), and restaurant reviews (Kovács and Hannan, 2010).
Although researchers studying category spanning have put forward an impressive set of empirical findings over the last decades, the theoretical underpinnings of these findings are not clarified, and a number of different theoretical processes have been advanced to explain the consequences of category spanning. The explanations proposed by organizational theorists mainly fall in two areas, which we term the “quality-based” and the “socio-cognitive-based” explanations. The quality-based explanations relate to the producer-side consequences of category spanning and explore the effects of spanning on the quality of the offerings of organizations. The socio-cognitive explanations focus on the audience-side consequences of spanning. One socio-cognitive explanation argues that category spanners violate institutionalized expectations and thus are viewed illegitimate (Zuckerman, 1999). Another approach focuses on the confusion and ambiguity stemming from category spanning: Category spanning confuses audience members who do not know what to expect from these organizations (Hsu, 2006; Hsu et al., 2009; Kovács and Hannan, 2010, 2011). A mechanism emphasized here is typicality: Organizations that span multiple categories are atypical to each of the categories spanned; thus, audience members cannot rely on category schemas to form clear expectations about the offerings of the organizations (Hsu et al., 2009; Kovács and Hannan, 2011; Negro et al., 2010).
While researchers are aware that quality-based and socio-cognitive-based explanations could be confounded (Hsu, 2006; Hsu et al., 2009; Zuckerman, 1999; Zuckerman et al., 2003), there has been little attempt to test the mechanisms directly, 1 nor have researchers put forward an empirical analysis in which multiple alternative mechanisms are contrasted. There is a need to open the “black box” of the consequences of spanning in order to identify the mechanisms that cause them. In most of the above-mentioned articles, one (or more) of the mechanisms is emphasized to derive macro-level consequences, and the tests operate on the macro-level consequences. In this article, we contrast the typicality-based and the quality-based mechanisms. Getting at the mechanisms is important not merely for theoretical purposes but because different mechanisms have distinct implications. For example, the quality-based and the typicality-based explanations provide different predictions for quality-category organizations: The quality argument would not differentiate among single-category organizations, while the typicality-based argument predicts that single-category organizations that are atypical to their category have lower value than single-category organizations that are typical to the category they are in. The different mechanisms also provide diverging recommendations for organizational action: If the locus of punishment lies in the audience side (Hsu, 2006; Zuckerman, 1999), then organizations need to focus on audiences’ perceptions to dodge the negative consequences of spanning. If the mechanism is on the producer side and involves reduction in skills and capabilities, then organizations need to assess the relatedness of skills and technologies before they decide to span categories.
Our empirical setting is restaurants and restaurant reviewing. In short, we investigate whether restaurants that span categories (such as “Japanese” and “Mexican”) receive lower ratings from reviewers. The review data come from the online review website Yelp.com. We collected reviews between January 2010 and October 2011. The sample consists of 59,605 reviews written about 474 San Francisco-based restaurants by 32,624 unique reviewers. These data have been used previously in Kovács and Hannan (2010) and Kovács and Hannan (2011), who demonstrate that restaurants that span categories receive lower ratings. However, the two articles by Kovács and Hannan, as most other articles in the field, suffer from two shortcomings: They do not measure the typicality of restaurants but assume that multiple-category restaurants are less typical than single-category restaurants, and they do not control for the alternative, skill- or quality-based mechanisms. To further their results, in this article we collect and analyze two additional data sources: To assess typicality, we collected the menus of the restaurants in our sample. To assess quality, we collected the “food” quality scores of the Zagat Guide, which has previously been used as a measure of restaurant quality (Roberts et al., 2013). Combining these two additional data sources with the review data allows us to directly test and disentangle typicality-based and quality-based explanations for the negative consequences of category spanning.
A main novelty of this article is a direct instrument of typicality. Previous research assessed typicality indirectly and assumed that the more categories an organization populates, the lower its typicality in each of the categories populated (Hsu et al., 2009; Kovács and Hannan, 2010; Negro et al., 2010). For example, a restaurant that is labeled both “Japanese” and “Mexican” is likely to be atypical of both categories. While we are sympathetic to this approach, here we argue that a more direct instrumentation of typicality is needed. First, approaches that use multiple-category membership, in lack of better evidence, have to assume that single-category organizations are all typical to their category. This is clearly an oversimplification. For example, not all single-category “Italian” restaurants are 100% typical Italian. As single-category organizations are prevalent (e.g. they constitute almost half of our sample), dealing with them is crucial to test the theory. 2 Second, not all multiple-category organizations violate the “categorical imperative” to the same extent: For restaurants, an “Indian” and “Pakistani” combination is likely to be less detrimental than an “Indian” and “Japanese” combination. To account for such cases, ideally one needs to examine the actual offering of the organizations. 3
Our empirical strategy to assess the typicality of restaurants is to contrast the offering of the restaurant (items on the menu) with its labels (Italian, Japanese, etc.) and assess the extent to which the menu fits the label(s) the restaurant claims. We use a commonly utilized computational linguistics approach, word-category co-location mapping (Manning and Schütze, 1999), to explore the schemas of organizational categories. We establish the typicality of the restaurants in the categories by comparing the restaurant’s offering to the schemas of the categories. We assert that the category mismatch of a restaurant is high when the offerings are atypical of the label(s) to which it is assigned. To our knowledge, ours is the first article in the category-spanning literature that actually measures the offerings of the organizations and assesses typicality in such a way.
Besides contrasting the two main mechanisms for the effects of category spanning on the average effect of category spanning, we also contribute to current literature by theorizing about situations in which category spanning can be beneficial. We argue that the above-specified mechanisms, typicality and quality, interact in a way that typicality is advantageous for low- and mid-quality organizations, but high-quality organizations can benefit from being atypical.
Finally, we distinguish between two kinds of category spanning. One, where the items of different categories appear side by side on the organization’s profile, we call it “food court” type of spanning with respect to our empirical setting. The other type of spanning we label “fusion,” the elements of different categories combined within the organizational items. We argue that “fusion” type of category spanning is more beneficial of high-quality organizations, while “food court” type of category spanning is more beneficial to lower quality organizations.
Theoretical background and hypotheses
The interest in the implications of category spanning has a long tradition in organizational research, resulting in a wide range of seemingly contradictory findings. Many researchers emphasize the negative average effect of spanning. As we discuss in detail below, these researchers evoke explanations such as the negative average effect of spanning on quality or typicality (e.g. Hsu, 2006; Hsu et al., 2009; Kovács and Hannan, 2010; Negro and Leung, 2013; Zuckerman, 1999).
Other studies document cases in which category spanning is beneficial. The literature on related diversification, for example, discusses cases in which category spanning promotes the evolution of new organizational capabilities (Markides and Williamson, 1994). Alvarez et al. (2005) develop a micro theory of creative action by examining how distinctive artists shield their idiosyncratic styles from the isomorphic pressures of a field. They show that auteur directors receive both critical and public acclaim because they span multiple movie genres. Importantly, most of the innovation literature emphasizes the positive effect of spanning (“recombination”) on innovative output (Fleming, 2001; Schumpeter, 1934). Baker (1992) documents cases in which diversification can increase value. Villalonga (2004) argues that after taking selection effects into account, no overall negative effect of “diversification discount” can be found. Dobrev et al. (2001) finds that larger niche width decreases the hazards for exit and disbanding.
While we do not aim here at fully reconciling all these findings, we would like to note a few points that help situate our theorizing and delineate the scope conditions of our theory. First, the literature that emphasizes the negative consequences of spanning typically focuses on the average effect, while the creativity and innovation literatures focus on the upper tails of the distribution by showcasing that highly successful organizations or innovators tend to be category spanners. Therefore, it might be the case that spanning leads to higher variance (as Fleming, 2001 shows), but at the same time, spanning is detrimental on average. Second, as we demonstrate later, it can be the case that spanning is beneficial if the organization is of high quality or status (Phillips and Zuckerman, 2001) but not otherwise, thus not controlling for quality or status might result in contradictory findings. Third, if the typicality argument is right, it might be the case that category spanning has a negative effect on organizational outcome only in settings where audience perceptions play a significant role. Fourth, the effect of spanning might differ by the outcome variable used. Fifth, the consequences of category spanning might depend on environmental conditions. As Freeman and Hannan (1983) assert, specialists are preferred in stable environments, but generalists have higher survival rates in uncertain environment.
Given these considerations, we emphasize that our theorizing below refers to a setting in which audience perceptions are important and that our main outcome variable is rating of the organizations’ [products] by audiences. This is also a setting in which the environment is quite stable, at least in the few years our observational window encompasses. In the first three hypotheses, we focus on the average effect of spanning, while in the last two hypotheses, we explore how the effect of spanning might differ by organizations.
The producer side: how category spanning affects average quality
Category spanning or, as it used to be referred to, the “generalists vs. specialists” issue, was a central topic in early population ecology (Dobrev et al., 2001; Freeman and Hannan, 1983; Hannan and Freeman, 1977, 1989). Building on the niche theory of Levins (1968), researchers argued that because an organization’s level of resource, budget, and attention is finite, the more markets the organization engages (ceteris paribus), the fewer resources it can spend on developing products, to attend to specific audiences, and to cumulate expertise in each of the categories spanned (Hannan et al., 2007; Hannan and Freeman, 1989). For example, a film actor has to decide whether to specialize in comedy or to prepare for both comedic and dramatic roles.
Because of scale advantage, such a dispersion of attention and budget implies that organizations that span multiple categories cannot excel in each of the categories spanned. Organizations that attempt to develop skills and invest in quality in multiple categories run the risk of becoming a “jack of all trades but master of none” (Hannan and Freeman, 1989; Hsu, 2006). Building on this argument, researchers predicted that category spanners have lower value to audiences. This prediction has been confirmed in various settings such as restaurants (Freeman and Hannan, 1983) and feature films (Hsu, 2006). 4 (It is important to note that Freeman and Hannan (1983) argue that generalists have higher survival chances in highly uncertain environments, where the benefits of hedging the risk against specializing in the wrong niches can outweigh the negative effects spanning has on average quality.)
The diversification literature in strategy also investigates the consequences of category spanning. For example, Lang and Stulz (1994) and Berger and Ofek (1995) find that diversified firms trade at a discount relative to single-segment firms. While the validity of these finding are questioned (Villalonga, 2004), their argumentation mostly concerns producer-side consequences. For example, Berger and Ofek (1995) cite as potential benefits of diversification “greater operating efficiency, less incentive to forego positive net present value projects, greater debt capacity, and lower taxes;” as possible negative consequences of diversification, they list “the use of increased discretionary resources to undertake value-decreasing investments, cross-subsidies that allow poor segments to drain resources from better-performing segments, and misalignment of incentives between central and divisional managers” (Berger and Ofek, 1995: 40).
A common shortcoming of the above research streams is that they rarely distinguish empirically the producer-side and audience-side effects of spanning and diversification. We think that aggregate-level tests are not adequate to settle whether there exists a producer-side advantage or disadvantage of spanning, as audience-side effects might counteract producer-side effects. Indeed, this might possible explain the inconsistent findings in the diversification literature. As an important exception, Negro et al. (2010; Negro and Leung, 2013) recently analyzed the effect of spanning on quality using a unique empirical data set: blind tasting of wines. Using blind taste ratings of wines to compare the quality of the wines produced by specialist and category-spanning wine producers, they find that wines of category-spanning wine producers on average get lower ratings. This, they argue, is clear evidence of the quality effect, assuming that in blind tasting experiments the raters are unaware of the identity of the producers. Negro et al. (2010) also compare the blind wine tasting results with non-blind wine tasting results, and hypothesize that the negative effects of category spanning should be stronger for the non-blind tasting setting because besides the quality-based effects, socio-cognitive effects are present as well.
We extend previous research by using a mediation framework to investigate whether quality mediates the negative consequences of category spanning. We aim to demonstrate three relationships. First, we test whether category spanners on average indeed confer lower quality offerings. Second, we test whether low-quality offerings lead to diminished value for audiences. Third, we investigate whether quality mediates the effect of category spanning. Figure 1 demonstrates these relationships visually.
Hypothesis 1. Organizations that span multiple categories on average receive lower value ratings from audiences than single-category organizations.
Hypothesis 2a. Organizations that span multiple categories provide on average lower quality offerings than single-category organizations.
Hypothesis 2b. Organizations that provide lower quality offerings receive lower value ratings from audiences.

Visual representation of the first three hypotheses.
The audience side: category spanning, typicality, and audience evaluations
A distinct group of mechanisms used to understand the consequences of category spanning focus on the social and cognitive consequences of category spanning. Researchers in this tradition emphasize the role audiences play in shaping organizational outcomes. For example, Hsu and Hannan (2005), when discussing organizational identity, assert that “[organizational identity is] not simply a list of observable properties,” but resides in the “perceptions, beliefs, and actions of contemporaneous audiences.” (p. 474)
Central to the perception of organizations by audiences are labels and categories used by organizations. Categories provide an interface between organizations and audiences (Ruef and Patterson, 2009; White, 1992). Labels evoke expectations in audience members, assisting them to navigate the organizational space. From this aspect, category spanning is detrimental because category spanners confuse audience members. Zuckerman (1999) demonstrates that when financial analysts specialize, they tend to ignore firms that span categories. Hsu (2006) argues that category spanning confuses audience members about the offerings of the organization, which hinder audience members in affiliating to organizations that fit their preferences. A main mechanism used to justify the negative effect of spanning on cognitive confusion is typicality. Because category-spanning organizations are atypical of each of the categories spanned (Hsu et al., 2009; Kovács and Hannan, 2011), audience members will not be able to make sense of them. This is especially the case when the categories spanned are distinct (Kovács and Hannan, 2011).
The hypothesis that typicality has a positive effect on value builds on extant behavioral research. An illustrative study is presented in Fiske et al. (1987), who investigate how the consistency between labels and attributes (i.e. typicality) influences affect. They presented experiment participants with descriptions of hypothetical persons. Each description consisted of a label that describes the occupation of the person (such as doctor and artist) and personality attribute lists (e.g. obedient, productive, and greedy). Fiske et al. (1987) demonstrate that subjects automatically judge the consistency between the occupation schemas and the listed attributes (e.g. the doctor-reliable pair is consistent, while the artist-reliable pair is less consistent) and show that the more typical the presented person is to the category schema, the more likely that the affect toward the category will apply to the presented person. For example, the higher the typicality of the presented doctor is to the “doctor” schema, the higher the affective rating will be. Thus, typicality in positively valued categories increases value.
Demonstrations of the positive effects of typicality on average value can be found in marketing as well. Ward et al. (1992) study how the typicality of the physical design of fast-food restaurants affects customer liking and demonstrate that the more typical a fast-food restaurant is in terms of physical design (i.e. the more similar it is to the prototype, McDonald’s), the more customers like it. Babin et al. (2004) show that the fit of the physical environment of a department store to the category prototype increases customers’ affect, quality perceptions, and shopping value. These results are all congruent with the hypothesis that lowered typicality results in lower value.
The chain of explanation goes as follows: category spanning decreases typicality, and lowered typicality results in lowered value to the audience members. Because organizational research has not yet tested this mechanism directly, we state these two steps as the following two hypotheses:
Hypothesis 3a. Organizations that span multiple categories are less typical to the categories they populate than single-category organizations.
Hypothesis 3b. Typicality on average has a positive effect on value ratings by audiences.
The interaction between category spanning and quality: how high-quality organizations can benefit from category spanning
The above arguments theorized about the average effect of category spanning. But could certain types of organizations benefit from category spanning? Here we build on the innovation literature and the literature on status and conformity (e.g. Phillips and Zuckerman, 2001) to explore the interaction of quality, typicality, and spanning. We argue that organizations that are of high status or high quality might benefit from engaging in category-spanning activities. Our argument combines two theoretical streams.
On the one hand, the innovation literature argues that innovation and problem solving are enhanced when previously unrelated categories of knowledge or technology are brought together and integrated. According to this view, spanning distinct and taken-for-granted categories provides actors with a broader variety of information and perspectives, which prompts unexplored mental models (Holyoak and Thagard, 1995), leads to unusual insights (VanLehn and Jones, 1993) and, thereby, initiates innovation (Burt, 2004; Fleming, 2001; Rosenkopf and Nerkar, 2001). The innovation-enhancing effect of category spanning has been systematically documented across levels of analyses, including individual inventors (Nerkar and Paruchuri, 2005), innovation projects (Fleming, 2001), business units (Rosenkopf and Nerkar, 2001), teams (Reagans and Zuckerman, 2001), and firms (Hargadon and Sutton, 1997). Fleming (2001) shows that category spanning (recombination) leads to higher variance in performance: recombinative patents are more likely to be a breakthrough success but also more likely to be a failure.
On the other hand, audiences’ reaction to category spanners depends on the status of the organization that engages in category spanning. Phillips and Zuckerman (2001) distinguish between three strata of status: low, middle, and high status, and argue that the consequences of “categorical imperative” differ across these levels. Low-status actors are deviants who do not conform to a minimal set of requirements to be considered as legitimate “players” in the market (Phillips and Zuckerman, 2001: 385). Middle-status actors are only considered legitimate to the extent they conform to the rules and expectations of the relevant audience; therefore, one could expect the highest level of conformity (typicality) for middle-status actors. High-status actors, however, are considered legitimate members of the market and thus do not risk losing legitimacy by innovating. These actors, thus, are more likely to engage in behaviors and practices that are less typical to their category and to use these behaviors and practices to further differentiate themselves from their middle-status peers.
Taking these two sets of arguments and applying them to the food facilities setting, we predict that category spanning and atypicality are more prevalent in high- and low-status food facilities than in middle-status food facilities. As Phillips and Zuckerman (2001) would predict, low-status food facilities such as hot dog stands are not considered legitimate members of the restaurant domain. Our data set does not contain very low-status food facilities that have no chance to become legitimate member of the restaurant domain (we sampled on facilities that are listed as restaurants). Thus, the lower and middle-status and quality 5 restaurants in our sample have higher need to conform to their categorical schemas and thus stay typical; only restaurants that are of high status and quality can successfully use atypicality as a way to differentiate themselves from other restaurants in their categories:
Hypothesis 4. Typicality is mostly beneficial for mid- and low-quality restaurants. Atypicality can be beneficial for restaurants that are of high quality.
Distinguishing the “food court” and the “fusion” type of category spanning
Category spanning can be of two kinds: one in which the elements of the spanned categories coexist side by side but are not combined and one in which the elements of the spanned categories are combined into the same products. 6 In the case of restaurants, we call this “food court” and “fusion” type of spanning (Baron, 2004). That is, a Mexican–French restaurant can be either such that on one side of its menu it offers Mexican dishes, while on the other side it offers French dishes or it might mostly serve dishes that fuse elements of the two cuisines. These are qualitatively distinct cases of category spanning. Here, we argue that these restaurants would attract different audiences. We build this argument around sociological theory of omnivorousness of audiences (Peterson and Kern, 1996). Traditionally, as Bourdieu (1984) documented, high-status individuals tended to differentiate themselves from others whom they viewed as lower status by engaging in cultural consumption patterns that were exclusive to them (such as going to the opera; see Bourdieu, 1984). Bourdieu argued that there exists a homology between social stratification and cultural consumption, whereby each social group consumes a set of cultural products that are typical to them. In the last decades, however, this cultural consumption pattern has undergone a substantial shift, and empirical research in the sociology of consumption has documented a shift toward omnivorousness in modern societies (e.g. Peterson and Kern, 1996; Vander Stichele and Laermans, 2006). High-status individuals have opened up to a wide range of cultural preferences and consume cultural products from a wide variety of genres, be it traditionally defined high or low brow (Peterson and Kern, 1996; Vander Stichele and Laermans, 2006; Warde, 2005). Not only are omnivores open to a wide variety of genres but they are also more tolerant to, and are often in search of an interesting combination of categories and innovation. Thus, we predict that omnivores are more likely to appreciate the “fusion” type of spanning.
As omnivores tend to come from more well-to-do socioeconomic backgrounds (Warde, 2005), we predict that they can afford visiting better quality and more expensive restaurants. Therefore, we predict that
Hypothesis 5. “Fusion” type of category spanning is more beneficial for high-quality restaurants.
Empirical setting, data sources, and data operationalization
Our empirical setting is the restaurant domain of San Francisco. The restaurant domain provides an apt setting to study the consequences of category spanning. First, although the restaurant domain contains a variety of organizations, these organizations are similar enough to be compared meaningfully. Namely, they have easily comparable product structure (menus), and the same notion of quality applies to them. This would not be the case if we were to use a multi-domain setting: For example, comparing the product structure of a car manufacturer and a fast-food chain is unobvious, and so is comparing the quality of the offerings of such organizations. Second, analyzing the restaurant domain allows us to build on previous research in restaurants and categories (Carroll and Wheaton, 2009; Freeman and Hannan, 1983; Kovács and Hannan, 2010, 2011; Rao et al., 2003, 2005). Third, in the restaurant domain, detailed and comprehensive records are available on restaurants, their menus, and customer evaluation. Fourth, restaurants are mostly of comparable (and small) size, which rules out the alternative quality-based explanation that larger organizations could develop expertise in multiple categories.
Below we discuss the three data sources we used: Yelp.com, MenuPages.com, and the 2011 edition of the Zagat Guide. Our sample contains the 474 restaurants that were covered in all three data sources.
Restaurant ratings
We collected the reviews on the restaurants from the website Yelp.com. The website generates its reviews through a volunteer process in which customers can go online and write a review. Each review captures four pieces of information that are linked to the restaurant: (1) a unique identifier of the reviewer; (2) a star rating, ranging from 1 to 5 as an integer number; (3) a text review; and (4) the date of the review. We use the reviewer IDs to control for reviewer-specific effects. The average of the ratings is 3.8; the median is 4 stars. In this article, we do not utilize the text of the reviews. Note that Yelp.com encompasses a broader audience than many food and gourmet magazines and media outlets (cf. Johnston and Baumann, 2007). We downloaded the reviews for the restaurants starting from 1 January 2010 to 1 November 2011. 7 The emerging sample contains 59,605 reviews.
Restaurant menus
To compile a data set on the menus of San Francisco restaurants, we used the website MenuPages.com. We downloaded the menus from MenuPages.com on 18 October 2011. The menus are sent by the restaurants to the management of the website (either by mail, fax, or by uploading electronically to MenuPages.com), where the menus are checked and formatted in a standard format. All font styles, pictures, colors, and other stylistic items are removed, and all that appears on the MenuPages.com website are the names of the items offered in the restaurant, their descriptions, and the prices. Figure 2 illustrates this standard format, showing a snippet of the menu of “Andale,” a Mexican restaurant in San Francisco. Note that in our analyses, we use all items on the menu (i.e. not only food items) because we believe that drinks can be part of the schemas as well. For example, a typical Japanese restaurant is supposed to carry sake. As we explain later, our method of analysis ensures that items that appear indiscriminately on most menus (such as “coke” or “juice”) drop out from the schemas.

Illustration of the format of restaurant menus on MenuPages.com. This figure shows the first few items on the menu of “Andale,” a Mexican restaurant in San Francisco.
The restaurants self-categorize themselves each into one or more cuisine categories. They can choose from 91 labels. 8 Most restaurants are in one category (44%), others are in two categories (40%), and some are in three or more categories (16%). The most popular categories are “Sandwiches,” “Chinese,” “Italian,” and “Japanese.”
The menus vary in length. The shortest menu only lists seven items, while the restaurant with the longest menu offers 293 items (this is “Cheesecake factory”). The average number of menu items is 92.21. Not surprisingly, there is a significant positive correlation of 0.19 between the number of items on the menu and the number of labels assigned to the restaurant.
Using the Zagat Guide to assess quality
We obtained quality scores for the restaurants from a third source, the 2011 online edition of The Zagat Guide. The Zagat Guide, similar to Yelp.com, bases its restaurant ratings on the experience and satisfaction of restaurant goers, who voluntarily submit their ratings and reviews to Zagat. The scores along each of these dimensions can range from 0 (lowest quality) to 30 (highest quality). There are three major differences between Yelp.com and the Zagat Guide. First, and the reason for choosing the Zagat Guide, Zagat reviewers are asked to score the restaurant along three specific dimensions: food, décor, and service. That is, Zagat’s reviewers supposedly do not take typicality into account when evaluating the restaurants (Schkade and Kahneman, 1998, see more about this later). Second, Zagat compiles (averages) the individual scores and only publishes the aggregated scores along these three dimensions but not the individual scores.
Third, Zagat provides lower coverage than Yelp.com and MenuPages.com, covering 474 San Francisco restaurants in its 2011 edition. Zagat tends to cover “better” restaurants, so this is not a random subsample. The average Yelp.com star rating of restaurants covered in Zagat is 4.1, as opposed to the 3.72 average in the full sample. This might bias our estimates, but we do not see this as a primary source of concern as there is still much variance in the sample, both along the quality ratings and the Yelp star ratings.
Quality, typicality, and price as dimensions of value
The ratings reviewers give to restaurants are a function of the perceived value of their visit to the restaurant. The value of a product or service to a consumer can be regarded as the person’s overall assessment of the utility of a product, brand, service, or experience (Zeithaml, 1988: 14). This overall assessment often involves emotional, social, quality, and price dimensions (see Sweeney and Soutar, 2001). The main research question addressed here is how quality and typicality of a restaurant affect its perceived value and thus influence the choices consumers make. Clearly, our finding that quality increases value is not that surprising, and so our main contribution in this sense is examining whether typicality confers value. One issue is worth discussing here. It is rather hard to obtain “objective” quality ratings, especially in the restaurant domain: many would argue that restaurant and food quality are inherently perceptual and subjective. What we mean by “external quality assessment” is to separate the quality-related components of perceived value from the typicality-related components of perceived value. Thus, by using “food quality” scores from the Zagat Guide as a measure of quality, and also controlling for price, we separate various dimensions of value. We believe that Zagat, by asking to rate specific dimensions, is successful in priming the evaluation of specific quality criteria and holds typicality ratings in the (cognitive) background. In other words, while the typicality affects the perception of quality, elicitations of quality ratings in a dimension-specific way reduce the importance of typicality that it has in the overall evaluation question (which Yelp asks). This assumption builds on results in psychology. For example, research shows that by priming specific dimensions of evaluation, subjects overweigh the primed dimension in their overall evaluation of options (this effect is known as the focusing effect, see Schkade and Kahneman, 1998). Schwarz (1996) reports an experiment in which subjects were asked two questions: one about the number of dates they had recently, and one about their general happiness. When the dating question was asked first, the correlation between the answers was 0.66; when the general happiness question was asked first, the correlation dropped to 0.12, proving that when a specific dimension is primed, subjects focus on that dimension of evaluation.
Assessing restaurant typicality
Hannan et al. (2007) assert that established organizational categories build up of two main components: a label that denotes the category and a corresponding schema that describes the category. A schema usually consists of typical attributes that characterize the category. For example, the schema of the “Italian” category could include words such as “pizza, pasta, mascarpone, tiramisu, and cannelloni,” or the “Japanese” schema could include “sushi, teriyaki, udon, seaweed, nigiri.” We note two properties of schemas: first, a schema is not simply a list of attributes, but these attributes can vary in importance to the schema, some attributes being more core than others (Murphy, 2002); second, a schema could contain negative elements as well, that is, elements that should not be in the attribute list, such as having “sushi” on the menu decreases the fit to the “Italian” schema.
The first step in our empirical strategy is to map the schemas underlying the restaurant categories. As Hannan et al. (2007) write, a schema is a set of attributes that defines a category or, more precisely, that describes the central tendency of a category (Hannan et al., 2007; Murphy, 2002). The typicality of an item in the category is a function of the number of attributes the object shares with the prototype. In our empirical setting, the attributes are words in restaurant menus, and the schema of a restaurant category is a (weighted) set of these words.
Instead of defining what the category schemas are, we take a constructivist stance and learn the category schemas from the menus. We want to identify the words that tend to appear on menus of certain categories but not others, and recreate the category schemas from these word occurrences. In the parlance of computational linguistics, this means learning the category schemas from word-category associations (Church and Hanks, 1990; Manning and Schütze, 1999).
Note that the approach we take here is a combination of the exemplar view and the prototype view of categories (Murphy, 2002). On the one hand, we follow the exemplar view because we take instances of the category and their descriptor, and we learn the category schemas from these instances. Thus, the category schemas we map from the data describe the current schemas in San Francisco. 9 One the other hand, we follow the prototype approach because we identify categories with their central tendency and do not compare the restaurant to all other restaurants in that category individually to assess typicality.
Our approach to map category schemas is as follows. First, we calculate the typicality of each of the menu words in each of the 91 categories. We calculate the typicality of the word in the category by calculating the Jaccard similarity of the word to the category. Formally, if #(wordi&categoryJ) denotes the number of times the word i appears on menus in category J, #(wordi) denotes the total number of times the word i appears on the menus of the restaurants, and #(categoryJ) denotes the total number of items in category J, then
See Appendix 1 for a detailed illustration of how word-category typicality is calculated.
Jaccard similarity is a commonly used similarity measure (Batagelj and Bren, 1995), and it satisfies three desiderata for the typicality measure: first, typicality of an item in a category increases with the number of co-occurrences; second, typicality of an item in a category decreases with the number of times the category appears. These two criteria ensure that items that tend to appear with a category have high typicality (e.g. the word “brie” tends to appear on menus of French restaurants). The third property of Jaccard similarity (typicality of an item in a category decreases with the number of times the item appears) discounts items that appear on most menus indiscriminately (such as “juice” or “coke”).
For computational simplicity, to assess category schemas, we only include words that are mentioned at least five times in the whole data set. We exclude all prepositions, conjunctions, and interjections. This leaves us with 12,323 unique words. We calculate the typicality of these 12,323 words in all the 91 categories.
In the next step, we calculate the typicality of each restaurant in each category by aggregating the individual word-category associations by taking the average of the category-word typicality for all words in the menu. Thus, each restaurant is assigned a typicality score in each category, where the typicality score ranges from 0 to 1; 0 denoting the lowest typicality and 1 denoting the highest typicality. Because the typicality values are low in absolute number (due to the division by the count of words in the Jaccard formula), for better interpretability we rescale the typicality values so that the maximum observed value of typicality will be 1. Note that as such a multiplicative rescaling only changes the unit of measurement. Before proceeding with the analysis, we provide three validations for our typicality measure. Figure 3, which illustrates category schema for three categories, provides the first validation: Italian, French, and Greek (for readability, we only show a select set of words; the schemas contain many more words). In the figure, the centers of the categories are denoted by the triangle, and the distance of words from the category centers are inversely related to their typicality score. This graph was calculated with correspondence analysis, a commonly used method to plot dual item-category distance data (Greenacre, 1984). The figure shows that some words are close to the Italian category but not to other: for example, “pasta,” “prosciutto,” “ricotta,” or “mozzarella.” Some words are close to the Italian category but are also close to the French category, such as “onion,” “balsamic,” or “wine.” These are words that are typical to both the French and Italian schemas. Finally, some words such as “bread,” “olive,” and “salad” are at equidistance from the three categories, indicating that they are present in all three categories but are not highly distinctive of each of the categories. The schemas on this figure do correspond to our expectation, providing face validity to our word-category typicality measure.

Category codes of three selected categories (Italian, French, and Greek). Correspondence analysis on a selected set of words.
For a second validation, we build on previous results in cognitive psychology demonstrating that categorization is a positive function of typicality (Hampton, 1998). We expect that the more typical a restaurant is to a category, the more likely the restaurant is self-categorized in that category. To test this proposition, we ran logistic regressions using typicality to predict whether a restaurant categorizes itself in the category. Results (not shown here) demonstrate a strong correspondence between typicality and categorization: the coefficient is positive and significant, and typicality explains 62% of the variation in categorization. This relationship holds even when we do not use the same restaurants to estimate typicality and predict categorization: our analyses show that the typicality values have a strong predictive power even on restaurants outside the sample (taking a holdout sample approach, see Stone, 1974). In Figure 4, we plot the probability of categorization as a function of typicality and find a positive relationship (note that the shape of this curve is similar to that in Figure 2 of Hampton, 1998).

Typicality and the probability that the restaurant will belong to that category.
Figure 5 provides a third validation, demonstrating that using menu words to explore the category structure of restaurants captures the macro category structure well. We calculated the Jaccard similarity between the categories (Batagelj and Bren, 1995) in terms of menu word overlaps: two categories are similar if words that appear on menus of category A tend to appear on menus of category B and vice versa. The specific measure we use is

Hierarchical clustering of restaurant cuisines with more than five instances. For the calculation we used the Jaccard similarity scores of the cuisines, based on the overlap in the menu words.
In Figure 5, we use these similarity values to create a hierarchical clustering of the restaurant categories (to avoid cluttering on the figure, we only plot the clustering for categories with at least five restaurants). As the figure shows, the category structure recovered by this method has high face validity, confirming the validity of mapping categories based on menu word occurrences.
Finally, to measure the typicality of the restaurants in the labels it claims, for each restaurant, we aggregate the typicality of words it contains on its menu (calculations are shown in Appendix 1). Figure 6 shows the distribution of the typicality values for restaurants in our sample. The figure shows a large variance in typicality of organizations in the categories they are in. We exploit this variance to understand the effects of typicality on audience value.

Distribution of the rescaled typicality values of restaurants in the categories they are classified in.
Our measure of typicality has an important assumption: we assume that schemas are shared by all actors. While it has been argued that category systems are more useful if there is a consensus about the category schemas (Hannan et al., 2007), there is undoubtedly some variance in the schemas audience members use to evaluate organizations. As we do not have access to these schemas, we assume that the schemas are shared.
Distinguishing “food court” and “fusion” type of category spanning
As we discussed above, category spanning can be of two kinds: one in which the elements of the spanned categories coexist side by side but are not combined (“food court”) and one in which the elements of the spanned categories are combined into the same products (“fusion”). To differentiate these two cases, we calculate the typicality of each of the dishes in the menus to each of the 91 categories. That is, we follow the same procedure as for the menus above, but now at the dish level. For food courts, we expect that each of the dishes will be only typical to a single cuisine, but that there will be dishes from multiple cuisines on the menu. For example, one half of the menu is Japanese, and the other half is Italian. For fusion restaurants, however, we expect that the individual dishes will themselves belong to multiple categories. For example, the dishes on the menu combine ingredients and techniques from Japanese and Italian cuisines.
We measure the extent of “fusion” kind of category spanning by calculating the Herfindahl index of the grade of memberships of each dish in each cuisine and then average these values for each restaurant. For example, if a menu contains two dishes, Dish 1 is 80% Italian and 20% Japanese, Dish 2 is 50% Italian and 50% Japanese, then the measure for fusion is calculated as ((.8^2 + .2^2) + (.5^2 + .5^2))/2 = .59. Note that this fusion measure falls between 1/91 (pure fusion) and 1 (pure food court), controlling for the overall extent of spanning.
Results
Variables
Our main outcome variable is the rating provided by Yelp.com reviewers. Our first explanatory variable is a continuous measure of category spanning: organizational niche width. Traditionally, organizational niche width has been measured with the count of market segments the organization targets (e.g. Hsu et al., 2009). Recently, Kovács and Hannan (2011) argued that such a simple counting of categories is not optimal because it does not take the similarity structure of categories into account. For example, while both in two categories, a “Mexican” and “Indian” restaurant has a wider niche than an “American (New)” and “Californian” restaurant. Kovács and Hannan (2011) propose a novel measure of niche width that captures the similarity structure of the categories spanned.
where catnum denotes the number of categories the organization claims, and
As we believe that this niche width measure is superior to simple category counting, this is the measure we use in the results presented throughout the article. However, we note that additional analyses (not shown in the article) revealed that the results hold even if we use the number of categories as a measure of niche width.
The other two main explanatory variables that have been described above are (1) restaurant quality measured with Zagat food quality scores downloaded from the online version of the 2011 Zagat Guide and (2) restaurant typicality.
We include a number of control variables. First, we control for the price level of restaurants. Yelp.com uses four categories to classify the price level of restaurant, indicating “the approximate cost per person for a meal, including one drink, tax, and tips.” (Yelp.com) A “$” restaurant denotes “cheap, under US$10,” “$$” denotes “moderate, US$11–US$30,” “$$$” denotes “spendy, US$30–US$61,” and “$$$$” denotes “splurge, above US$61.” In our sample, 32% of the restaurants are in the lowest price range, 49% are in the “$$” range, 15% are in the “$$$” range, and the remaining 4% are in the “$$$$” category. To allow for the nonlinear effect of price, we included dummy variables for all levels of the Yelp price rating. Second, we control for the overall popularity of the restaurants. We measure the popularity of a restaurant with the number of reviews it has received prior to the focal review. Third, we control for reviewers’ activism, as previous studies show that activist reviewers tend to give lower ratings (Kovács and Hannan, 2010). We measure reviewers’ activism with the number of reviews they have written prior to the focal review (Kovács and Hannan, 2010). To take into account the diminishing effects of popularity and activism, we used the log of the count of reviews written about the restaurant and the count of reviews written by the reviewer. Fourth, we control for the restaurant’s age at the review, which we measure with the number of months since the first review of the restaurant. Fifth, to control for cuisine-specific effects, we include dummy variables for the categories the restaurant is in. Finally, to control for geographical heterogeneity among restaurants, we include dummy variables for the zip code of the restaurant. See Table 1 for descriptive statistics and correlations for the main variables.
Descriptive statistics and Pearson correlations for the main variables.
SD: standard deviation.
N = 59,605.
The main effect of category spanning
We start by investigating the main effect of category spanning on value. Because of the ordinal and bounded nature of the outcome variable (the rating is from one star to five stars), we analyze the effect of category spanning on reviews using an ordered logit modeling framework. This framework has been used previously to analyze restaurant ratings (Kovács and Hannan, 2010, 2011). 10 As some control variables are review specific, we conduct the analyses at the review level. To account for possible heterogeneity among reviewers and restaurants, we present robust standard errors, but the results we present hold with alternative approaches to standard error calculations as well, such as clustering on reviewers and clustering on restaurants.
Table 2 presents the results. Model 1 shows the main effect of niche width on value. Model 2 includes the control variables. Both models confirm the hypothesis that restaurants with wider niches receive lower ratings. Table 2 shows that the effect of price is not linear: Cheap ($) and very expensive ($$$$) restaurants are more highly rated than medium-price restaurants. The control variables have the expected effect: more popular restaurants get higher ratings, and consistent with the previous results of Kovács and Hannan (2010), activists tend to give lower ratings.
Ordered logit regressions on Yelp ratings: the effect of category spanning (measured with niche width).
p < 0.01, **p < 0.05, *p < 0.1.
Robust standard errors are given in parentheses; N = 59,605.
Quality as a mediator of the negative effect of category spanning
Tables 3 and 4 analyze the mediating effect of quality. Table 3 contains two models investigating the effect of category spanning on quality. The results confirm Hypothesis 2a: category-spanning restaurants receive significantly lower food quality scores.
Linear regressions on the Zagat food quality scores: The effect of category spanning (measured with niche width).
p < 0.01, **p < 0.05, *p < 0.1.
Robust standard errors are given in parentheses; N = 59,605.
To test the second step of the mediation, we investigate how the food quality scores affect Yelp ratings. Table 4 shows the results. As expected, Zagat scores have a positive effect on Yelp ratings: the better the quality of the food, the higher the average star rating on Yelp, even after including the control variables.
The relationship between Zagat food quality scores and Yelp ratings (ordered logit regressions).
p < 0.01, **p < 0.05, *p < 0.1.
Robust standard errors are given in parentheses; N = 59,605.
The results of Tables 3 and 4 confirm the mediating effect of quality in the negative relationship between category spanning and value. The comparison of Models 3 and 4 in Table 4 shows that the mediation is complete: including the quality-related variables decreases the magnitude of the negative effect of category spanning to a level at which it becomes insignificant.
Typicality as a mediator of the negative effect of category spanning
In the next analyses, we investigate whether typicality serves as a mediator in the relationship between category spanning and value ratings. Table 5 shows the first step in the mediation analysis: the effect of category spanning on typicality. Because typicality is a continuous variable, we run linear regressions. Both models in the table confirm the hypothesis that category spanning decreases typicality. We note, however, that the effect size is rather small – this is partly due to the inclusion of cuisine- and ZIP code fixed effects. This nevertheless casts shadows on the general practice of assuming that category spanners are clearly less typical to their categories than single-category organizations.
The effect of category spanning on typicality (linear regression).
p < 0.01, **p < 0.05, *p < 0.1.
Robust standard errors are given in parentheses; N = 59,605.
In Table 6, we test the second step of the mediation: the effect of typicality on ratings. In Model 1, we only include typicality and the main control variables. In this model, typicality does not have a significant effect on ratings. However, as we demonstrated above, the quality scores strongly affect ratings. Thus, to model the effects of typicality, one needs to control for the quality effects. In Models 2–4, we include the food quality scores. To investigate whether the effect of typicality varies depending on the characteristics of the restaurants, we include interaction variables with niche width, food quality, and price. As the comparison of the log-likelihoods shows, the full models provide significantly better fit. In Models 3 and 4, the estimate of typicality becomes highly significant, and its effect size increases substantially. As Models 3 and 4 are the best fitting models, we conclude that typicality indeed increases ratings. This finding confirms Hypothesis 3b.
The effect of typicality, quality, and category spanning on Yelp ratings (ordered logit regressions).
p < 0.01, **p < 0.05, *p < 0.1.
Robust standard errors are given in parentheses; N = 59,605.
Hypothesis 4 posits that high-quality restaurants can benefit from atypicality. The results in Models 3, 4, and 5 of Table 6 confirm Hypothesis 4: the main effect of typicality is positive, but the interaction with food quality is significant and negative. The Zagat quality scores range from 1 to 30, and Model 4 suggests that atypicality is beneficial for restaurants that score 23 or higher (calculated as 4.157/.181). Thus, typicality is beneficial for lower quality restaurants, but as the food quality of the restaurant increases, the benefits of typicality wane, and for high-quality restaurants, it turns negative. Figure 7 visualizes effect of typicality for restaurants at different levels of niche width and food quality. Finally, note that typicality increases ratings especially for middle-priced restaurants.

Predicted effect of typicality (based on the estimates of Model 4 of Table 6).
The negative interaction effect between quality and typicality also shows up in the negative pairwise correlation between quality and typicality. The restaurant-level −.127 correlation between Zagat food score and restaurant typicality (significant at p < .01) indicates that high-quality restaurants are more likely to be atypical.
The strong negative interaction effect between niche width and typicality in Table 6 is worth further discussion. This finding, together with the positive main effect of niche width, shows that category spanning is detrimental to the extent that the restaurant actually engages in the multiple cuisines it claims in its labels. In other words, if a restaurant claims multiple labels and it tries to be typical to all the labels it claims, then it is worse off than restaurants that claim multiple labels but concentrate their efforts to one cuisine and are atypical to the other labels. This finding is consistent with current understanding of category spanning. Kovács and Hannan (2010), for example, assert that organizations that span multiple high-contrast categories suffer most from spanning because being member of multiple strong identity categories sends conflicting signals to audience members.
Hypothesis 5 posits that “Fusion” type of category spanning is more beneficial for high-quality restaurants. Model 5 in Table 6 investigates this hypothesis by adding to Model 4 “fusion” and its interaction with food quality. Recall that to assess whether a restaurant engages in “fusion” kind of spanning, for each dish on the menu, we calculate the Herfindahl index of typicality values and then we take the average of these Herfindahl indices for each restaurant. The reason behind this measure is that fusion restaurants tend to have low average Herfindahl values as they combine ingredients and techniques within dishes. As these average values are not normally distributed, we use a binary measure here: a restaurant is tagged as fusion if its average within-dish Herfindahl index is below the population average. Model 5 shows that on average it is detrimental to be a fusion restaurant. However, high-quality restaurants can benefit from being a fusion. This result confirms Hypothesis 5.
Finally, in an additional set of analysis not shown here, we investigated whether reviewer activism moderates the effect of category spanning and typicality on ratings. In line with Kovács and Hannan (2010), we found that category spanning and typicality have a weaker influence on ratings by activist reviewers. We share the explanation by Kovács and Hannan (2010), who argue that activist reviewers and reviewers with domain expertise are less likely to use category cues to navigate the organizational space, so the possible confusion arising from category spanning is less likely to affect them.
Robustness checks
To investigate the robustness of the above findings, we conducted several sensitivity analyses. As none of the robustness checks lead to estimates that are substantially different from the results presented above, we do not show them here, just list the alternative specifications we tried. First, as an alternative specification for category spanning, we used a binary version of the niche width variable (0 if single-category organization, 1 if multiple-category organization), and also tried models in which niche width is measured with the number of categories the restaurant populates. Second, we reran all the above analyses on the organization level, aggregating ratings and organizational and reviewer characteristics by organizations. Although some coefficients have changed slightly and some estimated coefficients lost significance in these alternative specifications, all results are consistent with the general empirical patterns described above.
Discussion
In this article, we set out to study the mechanisms that drive how category spanning affects audiences’ value ratings of organizations. We focused on the two mechanisms that are most often evoked in current organizational literature: the quality-based explanation and the typicality-based explanation. Our empirical setting was restaurants and restaurant reviews in San Francisco. To assess the typicality of restaurants, we collected the menus of restaurants. To assess the quality of the restaurants, we collected the food quality scores from the Zagat Guide.
The first contribution of this article is that we demonstrate that quality mediates the negative average effect of category spanning on value ratings: lower quality leads to lower ratings, and restaurants in multiple categories on average confer lower quality than single-category restaurant. These findings provide a direct corroboration of the assumption behind the principle of allocation (Hannan and Freeman, 1989). We also demonstrate the positive relationship between typicality and audience value ratings. While these two relationships have been widely held true (Hannan et al., 2007; Hsu et al., 2009; Kovács and Hannan, 2010, 2011), researchers have only provided indirect evidence by demonstrating the negative effect of category spanning on audience value (Hsu, 2006; Kovács and Hannan, 2010, 2011). We argued that a more direct test of typicality was needed. On the one hand, no previous research has shown that category spanning indeed diminishes typicality in the categories spanned. On the other hand, the multiple-category approach cannot say anything about the typicality of single-category organizations. By collecting and analyzing restaurant menus, we were able to assess the offering of the restaurants and directly assess the fit to the schema of the categories claimed by the restaurants. We believe that such a test is unique in the organizational literature. Combining the effects of category spanning, typicality, and quality in a single model (Table 6, Model 4), we find that these mechanisms provide related yet distinct effects on value. Quality does not explain away the effects of typicality, nor does typicality account for the importance of quality.
Besides demonstrating that category spanning has a negative effect on value on average, we also argued that certain types of organizations could benefit from atypicality and category spanning. We found that being typical is especially important for low- and mid-quality restaurants. High-quality restaurants, on the contrary, may benefit from category spanning and atypicality. This pattern is consistent with multiple streams of previous research. Phillips and Zuckerman (2001) demonstrate that high-status actors bear lower costs of illegitimacy than middle-status or lower status actors. Waguespack and Sorenson (2010) show that high status of film producers increases their chance of getting favorable classification. Another explanation consistent with our findings could concern audience heterogeneity and the tendency toward novelty seeking. Ward and Loken (1988) demonstrate that the positive effects of product typicality are more pronounced if customers do not look for novelty or exclusiveness. Babin et al. (2004) although confirming an overall positive relationship between typicality and value, also note that some subjects are novelty seekers, and slight deviance from the prototypes could often result in higher value. High-quality restaurants can thus benefit from spanning if customers who visit high-quality restaurants are more likely to be novelty seekers than customers of lower quality restaurants.
Finally, we differentiate between two types of category spanning, “fusion” and “food court” (Baron, 2004). We argued that high-quality restaurants could benefit from fusion type of spanning as they attract novelty-seeking patrons, and the high quality or prestige of the restaurant might make these audiences less suspicious of innovative dishes. Restaurants that are of low- or mid-quality should not try to innovate, however, as atypicality will hurt them. Thus, if they are to engage in any kind of spanning, it should be the less innovative “food court” kind of spanning.
Here we would like to note that while we contrasted the two types of spanning at the restaurant menu level, this distinction has interesting parallels at other levels of analysis as well. For example, at the organization brand level, it relates to the issue whether it is more beneficial for a firm to structure its products into separate but specialized brands or to combine the products under the umbrella of a unified brand. Or at the organization structure level, the “fusion” versus “food court” differentiation relates to whether different lines of businesses should be integrated or kept separate. While further elaborating on these parallels is beyond the scope of the current article, we note that, for example, the typicality-based argument would imply that keeping brands and lines of businesses separate is beneficial. This would be an argument favoring conglomerates that keep their brands focused and separate, thus not confusing their audiences. Combining the brands would be beneficial for brands with high quality (or status).
The corroboration of the positive average effect of typicality has implication for current research in organizational authenticity. Carroll and Wheaton (2009) argue that restaurants that are viewed as authentic are viewed more favorable by patrons. One of their constructs of authenticity is type-authenticity: an organization is type-authentic if its attributes and practices are consistent with its type. This construct is similar to what we call typicality. Our findings indicate that type-authenticity should be especially important for low- and mid-quality level restaurant. Future research could address how authenticity relates to quality and category spanning (see also Kovács et al., 2013) and how the four types of authenticity proposed by Carroll and Wheaton (2009) help explain the consequences of category spanning.
Limitations and future research
This study is not without limitations. First, we focused on the restaurant domain. A main advantage of the restaurant domain is that while it contains multiple categories, these categories and their schemas are highly commensurable as the offerings (menus) have the same structure. The question of generalizability to other contexts naturally arises. While we believe that the effects of category spanning, typicality, and quality are generalizable to other settings, we acknowledge that their relative importance might vary across settings. Take, for example, Hypothesis 4, which argued that high-quality restaurants could benefit from being atypical but mid- and low-quality institutions benefit from being typical. This finding might be specific to domains that value creativity, innovation, and novelty seeking but at the same time put emphasis on category fit. In domains where novelty seeking or artistry is not important (e.g. traditional financial institutions) or very important (such as design or creative art), this pattern might not hold. We encourage future research in such domains.
Second, we focused on the effects of category spanning on value ratings by audiences, and one could ask whether our findings generalize to other outcome variables such as profitability or survival. Investigating the effect of category spanning, typicality, and quality on these outcome variables is not possible with our data, and given the intricate relationship between specialism, generalism, profit, and survival, we refrain from making firm predictions. One the one hand, as restaurant ratings positively influence revenues (Luca, 2011), and value ratings negatively influence mortality (Hannan and Freeman, 1989), one could argue that our findings would generalize to profitability and survival. On the other hand, however, as population ecologists have shown, the trade-off between specialism and generalism often depends on environmental factors such as the stability audience taste and demand (Hannan and Freeman, 1989). When the environment is in flux, category spanning could be a more beneficial strategy in order to hedge against changes in taste. Similarly, atypicality might be beneficial if the organization can predict changes in taste and create their own market niches. We leave the disentangling of these effects for future work.
Third, for the categorization of restaurants, we use the categories the restaurants self-declare on MenuPages.com. This is possibly problematic on two accounts. First, a large proportion of consumers probably do not choose restaurants based on MenuPages.com; so this categorization might not be for the consumers. Consumers might decide based on alternative classifications for the restaurants or might not use category information at all. In this article, we assumed that the category claims the restaurant makes on MenuPages.com coincide with the category claims it makes in other places such as on other consumer websites, on their menus, or during advertising. Second, the self-declared categorization of a restaurant might differ from the “real” categorization for strategic reasons: restaurants might intentionally try to misrepresent themselves by either claiming labels they should not or by not claiming labels that they should. Although we do not think that this behavior is pervasive, with our data, we cannot rule it out. We leave it for future research to explore and test the possible consequences of the above two self-reporting biases.
Admittedly, our tests suffer from survival bias: our sample consists of restaurants that operated in October 2011, at the time of data collection. Unfortunately, data limitations prohibit us from addressing this selection problem. Future research could address this issue by prospectively downloading menus and coding the Zagat Guide in future years. In addition, we focused on a single city, San Francisco, and mapped the restaurant schemas in that city. This is adequate for our empirical study as we use local schemas to predict local value, but it would be interesting to gather and analyze menus and reviews from other settings and understand how restaurant schemas differ across cities or countries.
We also assumed in this article that the category system is stable. This is an especially sensitive assumption in the case of category spanning because in certain cases category spanners may create a new category by combining two previously distinct categories (e.g. this is how the “minivan” category was created; see Rosa et al., 1999). We believe, however, that such cases are rare. In the domain of restaurant, one could mention “minivan-like” new categories such as “Tex-Mex” or “Californian.” But these are rare cases and are very atypical. We (the authors) can personally recount numerous novel combinations from the San Francisco dining scene (Indian pizza, sushi taco and curry frozen yoghurt, just to mention a few), but very few of these caught on. Most category combinations never become successful and do not establish a new category. Second, not only are these new combinations unlikely to stick, but they also often take a longer time period to really become an established new category. We checked on MenuPages.com if any new category has been added during our observation period. No new categories have appeared on the San Francisco MenuPages.com website since 1 January 2010. Slightly before our observation period started, the “Gastropub” category was added. As this example demonstrates, new categories do occasionally emerge. Given that only seven restaurants in our data set are classified as gastropub, these cases do not threaten the validity our findings (our results hold after dropping these seven restaurants). One would need to take into account the changes in the category structure, however, if the study were to encompass a significantly longer time period, say 20–30 years.
A related limitation is that we analyze the data as cross-sectional, and thus, we cannot claim to show causality of the effect. Ideally, we would need panel data on both the Zagat scores and on the changes in the restaurant menus. Future research on category spanning is advised to collect panel data on the relevant variables in order to explore the causal relationships among category spanning, typicality, and quality.
Another limitation of the article is that we could not account for the effect that labels (beyond the effect through typicality) might directly influence the perception of the restaurants and the food served. Research shows that customers change their perception of the products as a function of the labels applied to them. For example, Wansink and Park (2002) demonstrate that subjects rate the same (non-soy containing) snack less tasteful but more healthy if the label indicates that the snack contains soy. Wansink et al. (2005) show that subjects who ate foods with evocative menu names (such as “Succulent Italian Seafood Filet”) generated a larger number of positive comments about the food and rated it as significantly more tasty than those eating regularly named counterparts (such as “Seafood Filet”). Future research should address how these effects interact with typicality effects.
Some might also find it problematic that our measure for restaurant quality is perceptual. We would argue that quality is often ambiguous and in most domains of life quality is relative to viewpoints, tastes, preferences, or customs. Not only is this true for food and most cultural products but for most consumer goods as well (Shepard, 1987). If so, a researcher cannot do otherwise but measure quality as perceived by relevant audience members. Having said this, we encourage future studies of the effect of category spanning in domains where quality can be measured more objectively.
A final avenue for future research could be to use more advanced computational linguistic approaches to map category schemas, such as topic modeling (Griffiths and Steyvers, 2004). On a related point, future research could further scrutinize how category schemas are stored, specifically, whether categories should be represented with the prototype, the exemplar, or some other frameworks (Murphy, 2002). In this article, we followed a combination of the prototype and exemplar views, but future research could investigate which representation describes audience members’ behavior better.
Footnotes
Appendix 1
This section provides an illustration how the typicality values are calculated for single- and multiple-category restaurants. In this hypothetical example, there are four restaurants, each with only one or two items (Table 7, see left panel). First, we transform this table to a category-word occurrence table (Table 7, see right panel). In case of multiple-category restaurants, we divide the occurrence of the menu items with the number of categories the restaurant belongs to. For example, in Restaurant C, the “mushroom” item will get 0.5 value in the “French” and 0.5 value in the “Spanish” category.
Then we calculate the Jaccard similarity index for all word-category pairs (Table 8). Each cell is calculated according to the following formula
where #(wordi&categoryJ) denotes the number of times the word i appears on menus in category J, #(wordi) denotes the total number of times the word i appears on the menus, and #(categoryJ) denotes the total number of items in category J. For example, the word “mushroom” appears 2.5 times in “French” menus, it appears three times in total, and there are 5.5 words in total in French menus
From Table 8, we calculate the typicality of each restaurant in each category by taking the weighted average of the Jaccard similarities of the menu words in that category. For example, the typicality of “Restaurant A” in the “French” category is calculated as (2 × 0.42 + 0.15 + 0.25)/4 = 0.31, whereas the typicality of “Restaurant C” in the “French” category is calculated as (0.5 × 0.42 + 0.5 × 0.25 + 0.5 × 0.084) = 0.13. Finally, to assess typicality of the restaurant in the labels it claims, we sum the typicality values for each organization in all categories it is in (these values are 0.31, 0.31, 0.13, 0.09, and 1; Table 9).
Acknowledgements
We appreciate the help and detailed comments of Nathan Betancourt, Gianluca Carnabuci, Jerker Denrell, Michael Hannan, Ming Leung, and Serden Özcan. The article benefitted from seminar discussions at Carnegie Mellon University, UC Berkeley, University of Lugano, and Yale University.
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
