Abstract
A long-standing concern in the study of culture is understanding how culture is distributed, often discussed in terms of cultural “coherence.” Cultural diversity, defined as the degree to which people share beliefs or meanings, is one dimension of cultural coherence that has been associated with many social outcomes. This article contributes to this area of research by considering how to measure cultural diversity in text and introducing a simple approach that uses word counts and sets of diversity indices (called “diversity profiles”). Text is useful for social-psychological analysis because as an artifact of individual thought, it provides a way to measure how beliefs and meanings are distributed and made salient across groups. The measurement approach outlined here contrasts to many contemporary computational approaches to measuring culture in text, which employ a relational logic of meaning based on word co-occurrences. While these more sophisticated approaches are well suited to measuring diversity in many instances, I show that there are some cases for which simpler measures based on word counts are ideal. After discussing the measurement of cultural diversity using word counts, I present a computational analysis of interview transcripts of American religious parents discussing the ages at which it is appropriate for children to participate in different practices often considered inappropriate for young children. The analysis points to the homogenizing influence of institutions on discussions of age appropriateness. I conclude by discussing implications for cultural analysis more generally.
Keywords
A long-standing concern in the study of culture is understanding how culture is distributed. The question is often described in terms of “coherence”: Do individual cultural elements cohere together as a unified whole, or is culture a collection of things with no necessary relation to each other? Early anthropological and sociological theories of culture tended to assume that culture cohered into shared “systems,” while contemporary theories argue that cultural elements are unevenly distributed across people rather than comprising a single entity shared by an entire group (Adams and Markus 2003; D’Andrade 2001; Dimaggio and Markus 2010; Hannerz 1992:7).
Contemporary cultural theories emphasizing the “collection” metaphor of culture over the “unified system” metaphor have transformed questions about cultural coherence, moving the question from whether culture is or is not coherent to the question of when and why coherence happens. While few today claim belief in necessarily coherent cultural systems, scholars recognize that collections of cultural elements vary widely, with some collections exhibiting more or less coherence than others. Measuring and explaining these patterns of cultural distribution across individuals or groups has become an important part of contemporary cultural analysis.
Cultural coherence has two dimensions: the degree to which people share a given cultural element (defined here as cultural diversity and the focus of this article) and the degree to which having one cultural element is associated with having or not having another. These dimensions have been given different names, including “consensus” and “concentration” and “tightness” and “consistency” (Ghaziani and Baldassarri 2011:3; Martin 1999, 2002). I opt for the term “cultural diversity” here to describe the first dimension because it matches the common interdisciplinary usages of “diversity.”
This article considers the measurement of cultural diversity in text. Textual data play an increasingly important role in cultural analysis, as digitization and the internet make textual data more voluminous and accessible and as computational tools open new possibilities for analyzing texts of all kinds, including ethnographic field notes, interview transcripts, archival records, books, and social media posts (Bail 2014; Rawlings and Childress 2021). As artifacts of individual thought (Ignatow 2016), texts provide an opportunity to observe how people think and communicate.
Applied to text, “cultural diversity” is a measure of the distribution of meanings and/or beliefs expressed via written words. A set of documents (a corpus) produced by a group of individuals may be said to manifest a high degree of cultural diversity within the group if the texts contain a wide range of meanings or beliefs across different documents. Alternatively, a group may be said to manifest a low degree of cultural diversity if different people tend to communicate the same meanings or beliefs across different documents. In this way, measuring cultural diversity in text provides insight into the cultural diversity of the creators of the texts.
Measures of cultural diversity in text may be of interest to cultural analysts both as an outcome variable and as an explanatory variable. As an outcome, text-based measures of diversity may be useful for studying whether institutions make certain ideas more salient (Bail 2012; Swidler 2013) or, when longitudinal data are available, for measuring the degree to which people grow more or less homogenous in their thinking about important topics like life goals (Alcaraz, Hayford, and Glick 2022; Frye and Trinitapoli 2015), causes of poverty (Homan, Valentino, and Weed 2017), or meanings of sexual harassment (Chawla et al. 2021). As an explanatory variable, cultural diversity has been associated with outcomes such as organizational effectiveness (Corritore, Goldberg, and Srivastava 2020; Ghaziani and Baldassarri 2011), sexual behavior (Harding 2007), religious belief and practice, (Berger 2011; Chaves and Gorski 2001; Hick 1997), and social network composition (Barnes-Mauthe et al. 2013; Grabowski and Kosiński 2006).
Textual data present tremendous opportunities for the study of cultural diversity, but the range of available measures and methods makes it difficult to know where to start. Different approaches have different strengths and weaknesses, so choosing the right tool for the job requires familiarity with different approaches and knowing what they offer for measuring cultural diversity. This article aims to facilitate and improve the analysis of cultural diversity in text by discussing the potential and limitations of contemporary methods that employ a relational logic based on word co-occurrences and then introducing a new approach using a simpler, alternative logic based on word counts, which has its own strengths and weaknesses that make it ideal for certain cases.
I begin with a discussion of three contemporary methods for measuring culture in text that have been or could be used to study cultural diversity. Many methods could be discussed here, but I chose three methods that represent different applications of a shared, dominant framework for thinking about the measurement of meaning that applies to most new and popular text measures. Although these methods differ in important ways, each employs a relational approach to measuring cultural meanings based on word co-occurrences. After discussing the strengths and possibilities of the relational approach to measuring meaning for studying cultural diversity, I discuss cases where relational measures of meaning may struggle and where simpler, “discrete” measures of meaning based on word counts (Stoltz and Taylor 2021:5), which ignore word co-occurrences, are optimal.
The main body of the article introduces a measure following this latter approach and illustrates it with a brief study of the way a set of religious American parents discuss the age appropriateness of a set of common practices often considered inappropriate for young children (e.g., trying alcohol, starting to date, and using social media). The analysis focuses on the diversity of ages mentioned by parents across the different practices, revealing that some practices are associated with a wide range of salient ages and that others are concentrated on a limited few. I argue that this variation can partially be explained by the influence of institutions that focus attention on the timing of certain activities while ignoring others. After presenting these findings, I conclude with a discussion about the implications for cultural analysis more generally.
Measuring Cultural Diversity In Text
Relational Measurement
Many methods exist to measure culture in text that have been or could be used to measure cultural diversity. Although considerably varied, most contemporary measures of culture in text use relational logic to identify meanings in a way that does not assume that different words have wholly different meanings. There are good reasons for this. Imagine, for example, that we wanted to measure how similar or diverse the following three statements are: “the shop is great,”“the store is fantastic,” and “the emporium is to die for.” Although a human would likely evaluate these three as semantically similar, a simple algorithm that treats unique words as semantically discrete would evaluate them as very different and thus a more diverse set because the three statements use different nouns and adjectives. The simple algorithm would err by incorrectly assuming that “store” and “shop” are categorically different, no more different than say, “store” and “bologna.” Contemporary computational methods avoid assuming discrete words carry discrete meanings by observing the contexts in which words are situated. As Firth (1957:11) famously quipped, “You shall know a word by the company it keeps!” I discuss three of these methods and their usefulness for measuring cultural diversity. The goal here is not to create an exhaustive list of possible methods for measuring cultural diversity in text but to identify a family of related methods based on this relational logic and discuss their strengths and weaknesses before introducing an alternative approach.
Latent Dirichlet allocation (LDA) topic modeling is a common relational approach to measuring culture in text, and it has been used to measure cultural diversity (Corritore et al. 2020; Hall, Jurafsky, and Manning 2008; Hashimoto et al. 2021). LDA topic modeling is an inductive method that divides texts into different topics based on word co-occurrences with the goal of uncovering the latent structure of the text. If a set of words tends to appear in close proximity within a given corpus (a set of texts), LDA topic models will likely group them together. Once these topics have been created, analysts can measure the diversity in the distribution of topics—within documents or across the entire corpora—measuring, for instance, whether authors give equal space to different topics or whether certain topics receive more or less attention than others.
Word embedding models offer another way of measuring cultural diversity (Stoltz and Taylor 2021). Word embedding models assign to each word a vector summarizing the contexts in which the word appears. Analysts can either use “pretrained” models trained on millions of documents or create their own using their corpus. There is considerable enthusiasm about word embedding models because, as Stoltz and Taylor (2021: 3) recently pointed out, “[i]nferring a word’s meaning by summarizing its (linguistic) context aligns with relational theories of meaning” that have been central in cultural theory (Kirchner and Mohr 2010; Mische 2011; Pachucki and Breiger 2010; Zelizer 2012). Word embedding models offer a useful method for measuring cultural diversity at the level of corpora. For example, say we have a set of documents and we wish to know the degree to which they are similar or different. Word embedding models can be used to create document similarity measures by way of a method called “word mover’s distance” (WMD; Kusner et al. 2015; Pomeroy, Dasandi, and Mikhaylov 2019; Stoltz and Taylor 2021). The result is a weighted document-document similarity matrix based on the overall semantic “distance” between documents. 1 The distribution of these document similarities can be used as a relative measure of cultural diversity when compared to distributions of document similarities from other corpora.
Concept mover’s distance (CMD) is a new computational text method that could also be used to measure cultural diversity (Stoltz and Taylor 2019). CMD uses word embeddings, but unlike WMD, which measures general semantic similarity between documents, CMD measures the degree to which a given text engages with a focal concept. Unlike topic modeling, CMD is a deductive method, requiring the analyst to provide one or more concept words. When applied to a corpus, CMD generates a measure of engagement for each text in relation to the focal concept. The distribution of engagement scores can be used as the basis for a relative measure of cultural diversity focused on a specific concept, for example, by measuring whether documents in a corpus have a more homogenous or diverse engagement with a focal concept than another corpus.
Discrete Measurement
The three methods previously discussed are examples of using the relational logic of word co-occurrences to measure meanings in a way that goes beyond the limitations of measuring meaning with discrete, individual words. They are justifiably popular and promising approaches to measuring culture in text and can be used profitably to measure cultural diversity in many cases. However, somewhat counterintuitively, there are times when measuring meanings using the relational logic of word co-occurrences and not coterminous with discrete words is not only unnecessary but also suboptimal. For these kinds of cases, I suggest considering an alternative, simpler approach based on word counts.
Despite the impressive usefulness of relational measures of meaning, cultural analysts may want to measure cultural diversity using frequencies of discrete, individual words when cases meet three specific conditions: (1) The goal is to measure diversity in engagement with a known set of concepts, (2) these concepts are closely associated with specific words or phrases, and (3) the concepts within the set are similar to each other yet significantly different. I will discuss each of these conditions in the following, but first, to get a better sense of the kinds of cases I am describing, imagine the following questions:
Do children from working-class families express a more or less diverse set of ideal occupations than children from middle-class families?
Has the United States’s international news coverage featured a more or less diverse set of countries over time?
People have different ideas about the appropriate time to try different activities. If we look at the ages that are associated with different activities, is the distribution of these ages more diverse for certain topics?
These hypothetical research questions exemplify the conditions described here. They all focus on measuring diversity in references to a limited set of knowable concepts (occupations, countries, and ages) that are closely associated with specific words or phrases (e.g., “baker,”“Azerbaijan,” and “18”). Additionally, each of these cases likely include concepts that are similar to each other but are significantly different. For example, nations like El Salvador and Honduras may appear similar (perhaps especially to a language model) because they are both Spanish-speaking countries in Central America, but they are clearly different.
Why might cases like this warrant a simpler method based on simple word frequencies? First, because the goal in these cases is to measure the distribution of a known set of concepts, LDA topic modeling (which aims to uncover the latent topics distributed throughout texts) is less fitting. A case could be made for using supervised approaches to LDA topic modeling that take seed words or labeled features (Andrzejewski, Zhu, and Craven 2009; Druck, Mann, and McCallum 2008; Li et al. 2016), but if the set of topics is already known and closely associated with a limited set of words, these methods will be unnecessary and likely introduce undesirable noise.
Second, measuring meaning by counting key words and phrases may be desirable in some cases because existing approaches that rely on word co-occurrences to measure similarity can wash over socially significant differences. 2 To understand why similar-yet-different concepts warrant special consideration, consider two phrases: “elementary school” and “high school.” In a word embedding model, these two bigrams would likely be quite similar because they are both related to children and primary education. For many applications, this similarity makes sense and would be a desirable outcome. However, imagine these two phrases were spoken by parents who were asked about the appropriate age to start dating. Although “elementary school” and “high school” may be semantically similar, in the context of age appropriateness, the difference between them is large and significant. In sum, there are cases in which seemingly small differences are culturally significant, and in these cases, cultural analysis may be more hindered than helped by washing out these differences with relational measures of meaning based on word co-occurrences. In such cases, the cultural diversity of meaning may be more effectively measured by counting keywords.
There are certain types of concepts that are more likely to meet the three conditions described earlier. These include (1) proper nouns, including names of people, places, and things; (2) discrete numbers (with some exceptions, discussed more later); and (3) categories identified by specific labels, such as genres (e.g., “folk” or “rap”) or time periods (e.g., “yesterday” or “high school”). These types of concepts are good candidates because they have relatively stable meanings that deriving inductively would likely confuse. Given current available methods, when we want to measure cultural diversity within a set of concepts like these, the best approach is likely to count keywords. 3 In the next section, I discuss how to measure cultural diversity using keyword counts and identify potential challenges and obstacles that need to be addressed. I then show how this approach can be used with a brief analysis of interview transcripts of American religious parents discussing the ages at which it is appropriate for children to participate in different practices often considered inappropriate for young children.
Measuring Cultural Diversity With Keyword Counts
Any kind of text data can theoretically be used to study cultural diversity in meaning using keywords, from interview transcripts to tweets to newspaper articles. Once text data are obtained, measuring diversity in meaning using keywords consists of three main tasks: (1) creating a keyword dictionary, (2) preparing the text, and (3) calculating diversity scores. There are special considerations associated with each of these, which I discuss in turn.
Creating a Keyword Dictionary
The validity of a cultural diversity measure based on keywords is limited by the quality of the keyword dictionary used to create it. A keyword dictionary is a list of keywords (individual words or phrases) and associated concepts. To understand why a good keyword dictionary matters, consider an ecologist measuring species diversity in a given area. If the ecologist’s diversity measurement is not based on accurate taxonomy, such that they do not know the full range of possible species and/or they do not know their distinguishing traits (e.g., they lump wasps and bees together), the measure will be inaccurate. When measuring cultural diversity in text using keywords, we need an accurate “taxonomy” of relevant concepts and an adequate list of associated keywords (Leinster and Cobbold 2012:487).
The “taxonomy” of relevant concepts will vary depending on the specific research question. In some cases, the domain we are studying has a clearly delimited set of concepts. For example, if we are studying diversity in mentions of foreign countries in international news coverage, the set of relevant concepts includes all the foreign countries of the world. In other cases, the full taxonomy of relevant concepts is more ambiguous. For example, what are the relevant age categories parents use when discussing age appropriateness? In cases like this, where there is not a clear census of concepts, some qualitative investigation may be useful to develop the requisite background knowledge. For example, my own reading of the interview data discussed later revealed that when discussing age appropriateness, parents sometimes referred to age categories by referring to institutions, such as “middle school,”“college,” or “marriage,” instead of referring to a numeric age.
Once we have the set of relevant concepts, we need a set of keywords to identify them. The ideal cases for keyword analysis are those in which there is one unique keyword for each concept. When a concept is evoked by multiple keywords, we must add each keyword as an instance of that concept. Failure to do this produces false negatives—skipping cases that should be included as manifestations of a concept. For example, Alexandria Ocasio-Cortez is commonly referred to as “AOC” or “Ocasio-Cortez.” An accurate keyword dictionary of congressional representatives would include known variants of their names. Knowing the relevant keywords may require a degree of qualitative analysis or the use of semiautomated methods to “fill in” gaps in manually created keyword dictionaries (Guermazi, Hammami, and Hamadou 2007). In some cases, block modeling may be a useful automated tool for facilitating the creation of keyword dictionaries. Block modeling has been applied to study similarities and differences in the meanings of different identities (e.g., “woman” and “wife”) by comparing the network structure of these different “nodes.” In other words, by comparing the contexts in which the different identity words appear, one can determine whether a set of identities is synonymous in the discursive space in which the texts are created (Mohr 1994; Mohr and Neely 2009; Mohr and Rawlings 2012). This approach may thus be useful for developing a parsimonious taxonomy of concepts and keywords.
In some cases, we know our concepts and our keywords but a given keyword is associated with multiple concepts. For example, “black” may refer to the color or to a person’s race. In these cases, we risk generating false positives in our data if we do not filter out the irrelevant instances of the keyword. Although this can be a problem, it is often not an issue because the text sample is focused in a way that makes these kinds of homophones highly unlikely. For example, if the sample includes texts about race, then “black” is unlikely to refer to the color. Analysts can test whether this is an issue in their data by identifying the passages in which the polysemous keyword occurs and counting how many are false positives (i.e., instances where the word does not refer to the desired concept). If it proves to be a serious problem, the analyst can attempt to differentiate the desired keyword by attending to co-occurring keywords or in extreme cases, contextualized word embeddings such as BERT (Devlin et al. 2018).
To summarize, a keyword dictionary requires knowledge of relevant concepts and knowledge of relevant keywords. The less these are known, the more difficult the task becomes. The further our research case gets from the ideal (i.e., a bounded set of concepts with unique associated keywords), the more likely it is that the research task is best approached with different methods. 4 However, for cases that meet the aforementioned conditions, keyword analysis is not only suitable but also optimal.
Preparing the Text
The next step to calculating cultural diversity of meaning using keywords is preparing the text. This entails three main tasks: formatting the text, using the keyword dictionary to tag the data, and preparing the resulting data set for measurement.
To increase the performance of concept identification, the text format should match the format of the keywords in the dictionary. For example, if the keywords in the dictionary are lowercase and without punctuation, the text should be made lowercase, and punctuation should be removed. Once the text is properly formatted, the keyword dictionary is used to identify all instances of the concept based on those keywords. The result is a document by keyword matrix, with each cell indicating the number of times a given keyword appears in a given document. 5
Once we have created the document-keyword matrix, the next step is to ensure it is properly structured as a “document-concept” matrix (see Table 1 for an example). Because the goal is to measure diversity in meaning with respect to a set of concepts, we need to take care of any instances in which a single concept is associated with multiple keywords. If our matrix has multiple columns referring to the same concept, such as the “AOC”/“Ocasio-Cortez,” these columns need to be merged. If these instances are not merged into instances of the same concept, the diversity measure will appear higher than it really is.
A Hypothetical “Document-Concept” Matrix, Recording the Occurrences of Five Target Concepts across Four Documents
Next, we may want to take steps to minimize the ability of individual documents to bias our measure. Single documents with many references to a given concept may bias the diversity measure if the cells are continuous counts. The diversity measures discussed later use the sums of all instances of each concept across documents, so a single document with a very high count could have an outsized impact. To remedy this, I suggest binarizing the values in the matrix, thereby measuring presence or absence of a concept within a text rather than frequency of lexicalization (Pang and Lee 2008).
Finally, rows in the document-concept matrix (representing separate documents) need to be aggregated into the desired unit of analysis. Returning to the example of species diversity, an ecologist must determine the boundaries in which to count species. In social research, we determine our boundaries by choosing which kinds of texts will be aggregated for the basis of our diversity measure. Although we can calculate diversity within single texts like computational linguists, cultural analysts will typically be interested in diversity in collections of text generated by individuals belonging to some social category, such as members of an organization or a demographic group. If the desired boundary has not already been established by the chosen sample of documents, the analyst should now divide the document-concept matrix according to the traits that are apropos to the research question. For example, if we were comparing news outlets, we would create subsets for each institution. This allows the diversity of any created subset to be calculated separately. Subsets should include the same or at least very similar numbers of documents because any disparity here will bias the comparison, most likely by giving the larger subset a larger diversity score.
Calculating Diversity Scores
Once the text is prepared and the data are properly structured, we can calculate the diversity scores. There are multiple ways to go about this, depending on the research question. The approach I advocate is a multipronged approach, following the standard practices of ecological sciences (Chao and Jost 2015; Leinster and Cobbold 2012; Tóthmérész 1995).
There are three primary ways to measure diversity. The simplest measure is richness (or “species richness” in ecology), which in the current context is the number of concepts that appear in our group of texts. For example, if we were measuring the richness of mentions of other countries in newspapers’ international coverage, the richness would equal the total number of foreign countries mentioned at least once. Richness is sometimes insightful, but it is limited because it does not take into account unevenness in how frequently different concepts are mentioned (referred to as “abundance” in ecology). For example, a newspaper may mention every country at least once, but some countries are mentioned much more abundantly than others; measuring diversity solely in terms of richness would miss the bigger picture.
Two common indices are used to calculate diversity that account for both richness and abundance: the Shannon index (Shannon 1948) and Gini-Simpson index (Simpson 1949). The Shannon index is rooted in information theory and is a measure of uncertainty in predicting the identity of an unknown entity in a given set. If the richness in a set is high (i.e., there are many different kinds of things) and the things in the set are evenly distributed, it is much more difficult to predict which entity is selected than in a set where there are only a few kinds of things and one occurs far more often than the rest (Morris et al. 2014:3515). The Simpson index, by contrast, is a measure of the probability that two randomly selected things in a set are the same type (e.g., the same species). The Simpson index is the same as the Herfindahl-Hirschman index (Herfindahl 1950; Hirschman 1980). 6 If you subtract the Simpson index from one, you get the Gini-Simpson index, which is a measure of the probability that two randomly selected things in a set are of different types. For this reason, the Gini-Simpson is more common than the Simpson index when measuring diversity. The Gini-Simpson index is the same as the Gibbs-Martin index (Gibbs and Martin 1962), the Blau index (Blau 1977), and Hurlbert’s probability of interspecies encounter (Hurlbert 1971).
The Shannon and Gini-Simpson indices are useful measures of diversity but are easy to misinterpret because they are nonlinear. To facilitate interpretation, ecologists recommend converting them to “effective numbers,” or “true diversity” (Chao, Chiu, and Jost 2010; Jost 2007, 2006). True diversity is calculated by figuring out how many unique equally abundant things are needed to generate the diversity index of the observed sample. True diversity based on richness, Shannon diversity, and Gini-Simpson diversity are mathematically similar. In fact, they can all be calculated by changing a single parameter (q) in the following formula (Hill 1973):
R refers to richness (the number of unique types), and Pi refers to the proportional abundance of the ith type. When q is set to 0, the effective number equals richness; when q = 1, the effective number is calculated according to Shannon diversity; and when q = 2, it is calculated according to Gini-Simpson diversity. The differing values of q affect how sensitive the measure is to rare versus abundant things. Richness (q = 0) is biased toward rare things because every new type of thing (e.g., every new species or concept) increases richness the same amount, regardless of how abundant it is. On the other hand, Gini-Simpson diversity (q = 2) is biased toward abundant things such that changes in abundance have more impact on the measure than changes in richness. For this reason, Gini-Simpson is often used in studies seeking to measure unequal concentration in things like market share or wealth. Shannon diversity (q = 1) is considered the most neutral because it weights each type by its frequency, giving richness and abundance similar weight.
One takeaway of this discussion is that distributions of things can be diverse in different ways. Although analysts can and often do use only one measure, this gives only a partial understanding of diversity. Because each diversity measure gives slightly different information, ecologists often create “diversity profiles” by calculating all three, yielding a more complete understanding of how things are distributed (Chao and Jost 2015; Leinster and Cobbold 2012; Tóthmérész 1995). In the next section, I show how using multiple measures of diversity can enhance understanding of cultural distribution.
Measuring Semantic Diversity In Parents’ Discussions Of Age Appropriateness Using Keyword Counts
This section contains an empirical demonstration of the method previously outlined. The goal of the analysis is to measure diversity in a set of religious American parents’ discussions about the appropriate age to begin different practices often considered inappropriate for young children. The focus is on diversity in ages associated with these different practices and ascertaining whether certain practices are associated with a more diverse set of ages than others. More generally, measuring variation in diversity in temporal associations may shed light on mechanisms that induce more cultural coherence.
The data for the analysis come from interviews conducted from 2014 to 2015 as part of the Intergenerational Religious Transmission Project. Researchers interviewed parents of various religious backgrounds from different regions of the United States on topics relating to parenting and the transmission of religion. Researchers chose participants using stratified quota sampling to include a diverse sample of religious adults from different religious traditions, family structures, ethnic and racial categories, and social classes and with varying levels of religious commitment. For this analysis, I used the transcripts from the 45 interviews (22 men and 23 women) that responded to a module on agedappropriateness (only a subset of the full sample was given the module). Interviewers told participants they would be asked at what age it was appropriate to begin certain practices and then read a list of practices, one at a time, allowing the parent to respond to each one. Table 2 lists the practices. 7
List of Practices
Following the steps outlined earlier, I began by creating a keyword dictionary. I began by reading interview transcripts and recording unique references to age categories and continued this until further reading consistently yielded no new findings. In the reading, I identified several patterns. In their discussions of age appropriateness, parents commonly referred to specific ages (e.g., 12, 18, 21), general age ranges (e.g., “teenage,”“adult,”“twenties”), types of educational institution (e.g., middle school, high school, college), grades in school (e.g., 10th, 11th, 12th; or sophomore, junior, senior ), institutional markers (e.g., “marriage” or “legal age”), and several ways of indicating “never” (e.g., “never,”“not appropriate,”“absolutely not”). Once I identified common phrases, I extended the observed ranges to include keywords that would logically fit that I may have missed (e.g., including “junior high”).
The list of keywords revealed a new challenge to be addressed: Different parents referred to age periods in different ways. For example, some parents referred to specific numerical ages, such as “13,” and others referred to institutional markers, such as “eighth grade.” To avoid artificially inflating the diversity scores, I replaced references to K–12 grades with their average accompanied ages. “Kindergarten” became “5,”“first grade” became “6,” and so on.
After assigning references to K–12 grades to the corresponding ages, another problem remained: Some parents referred to a range of ages such as “teenage,”“high school,” or “college” instead of a specific age. This reveals an additional challenge in measuring diversity in thinking about time, which is that people can think at different timescales. When it comes to age periods, for example, people can think in terms of years or groups of years. For calculating the diversity scores, we can either leave these as discrete concepts or combine the more fine-grained age concepts into groups according to similarity. There is no clear right or wrong answer, but the decision may significantly affect the result. Leaving the concepts as they are will result in a higher diversity score because concepts that could be grouped together are counted separately, increasing richness. Combining concepts into higher-level categories will yield a lower diversity score by decreasing richness. For this project, I decided to measure diversity without grouping together categories because increased temporal specificity on the question of age appropriateness suggests a significantly different orientation to the issue—meaning that is worth preserving.
For measuring semantic diversity with keywords, only minimal text preprocessing is required. I made everything lowercase, made sure that all numbers were the same format by converting numbers to text, and removed punctuation. I then created a data set that tagged instances of each of the age period concepts using the keyword dictionary. Before moving on, I checked the accuracy of the classification. I was particularly concerned with low numbers like “two,” which could refer to the age periods I intended to measure or irrelevant phrases, like “the two of them.” I considered dropping these keywords altogether because parents rarely used them, if at all, but given the small size of the data set, I manually checked each case, keeping only appropriately tagged data. In the end, however, keeping these rare numbers did not significantly change the diversity scores. It is worth noting that this appears to only be an issue with keyword analysis that includes lower numbers as keywords (e.g., 1–10), which according to Benford’s law, occur more frequently in human communication (Berger and Hill 2015).
The final analysis included 45 discrete age categories. I calculated the effective number of age periods using the Shannon and Gini-Simpson indices, meaning the number of equally distributed age periods that would be necessary to generate the observed value of the respective diversity indices. In what follows, I give examples of how to present and interpret diversity measures using different values of q. My primary goal is to highlight the kind of information that can be gleaned about cultural diversity using the keyword-based measure and provide a template for interpreting the measures. In so doing, I answer three questions: First, which topics yield more/less diverse distributions of culture (measured here as age categories)? Second, which topics exhibit higher concentration on one or more age categories? Third, what might explain the observed variation in cultural diversity?
Figures 1 and 2 provide detailed representations of the diversity of salient age categories for each topic. Figure 1 shows the effective number of the age categories for each topic based on the Shannon and Gini-Simpson indices. The topics are listed in order of increasing diversity (according to the Shannon index), with the most diverse topics at the bottom. Figure 1 shows that among this sample of religious parents, there is indeed variation in the degree of diversity across different topics in terms of the distribution of salient age categories. The distribution of salient age categories for discussing starting to date, owning a smartphone, and wearing makeup are much more diverse than the distributions for discussing swearing and trying drugs. However, Figure 1 also reveals that there are significant differences between the Gini-Simpson and Shannon measures that warrant extra attention. Figure 2 plots the diversity measures in a way that facilitates this comparison.

Effective Number of Age Periods by Topic

Diversity Profiles for Each Topic
Figure 2 shows diversity profiles for each topic. Following a common approach in ecology, I created the diversity profiles by calculating diversity scores at many values of q between 0 and 3 for each topic and plotting these as lines. The points on the lines are set to the most used values of q (0, 1, and 2) that correspond with richness, Shannon diversity, and Gini-Simpson diversity. The topics are organized from top left to bottom right according to the steepness of the slope (and by the gradient of the line to facilitate comparison). Steeper slopes and darker lines mean more unevenness in the distribution of age categories (i.e., more concentration).
Recall that the different diversity measures are sensitive to different features of distributions. Richness is sensitive to rare things and ignorant of abundance, the Gini-Simpson index is more sensitive to abundant things and less sensitive to rare things, and the Shannon index weighs rare and abundant things evenly. Because the Gini-Simpson index is more sensitive to concentration (abundant things) than the Shannon index, topics with greater concentration in some age periods have large differences between these two values, represented in Figure 2 by the steepness of the slope of the diversity profile curve and the gradient of the line. This provides additional information about the relative cultural diversity associated with these topics.
Consider the difference between discussions of social media and R-rated movies. Figure 1 shows that discussions about using social media and watching R-rated movies yielded almost identical distributions of salient age categories measured by the Shannon index. However, the diversity profiles of these two topics are quite distinct. The plots in Figure 2 suggest that discussions about watching R-rated movies were more concentrated (there was at least one abundant age category) than discussions about social media use. Indeed, if we compare the raw distributions of age categories for these topics, shown in Figure 3, we find that among the salient age categories mentioned in discussions about watching R-rated movies, there was a peak at age 17, with a quarter of parents mentioning it. At the same time, discussions of R-rated movies yielded more salient age categories than discussions of social media (shown in Figure 2 as the effective numbers at q = 0). Thus, we can say that discussions of R-rated movies and discussions of social media were similarly diverse (measured by the Shannon index) but in different ways. Discussions of R-rated movies were both less diverse and more diverse than discussions of social media insofar as they yielded more concentration and more age categories overall. By contrast, discussions of social media yielded fewer age categories overall but less concentration on any single age category. Thus, we see that comparing values of the three diversity indices yields detailed information about cultural distribution that single measures cannot capture alone.

Distribution of Age Categories by Topic
What might explain the observed variation in diversity scores and diversity profiles across the different topics? One possibility, following Swidler (2013:136), is that an “encompassing institutional matrix . . . accounts for the shared elements of a common culture.” Institutions may act directly on those who are subject to them or indirectly by instantiating widespread patterns that individuals perceive as salient (Lynn, Walker, and Peterson 2016; Salganik and Watts 2008). In the case of age appropriateness, an initial clue of the importance of institutions lies in the fact that many parents responded to the interview prompt specifically asking about an “age” with institutional markers (e.g., “high school” or “marriage”). This suggests that institutions offer a heuristic that “clumps” various possible ages together. 8 Additionally, institutions induce homogeneity in parental reasoning more directly by setting age standards for certain activities. For example, in the United States, watching an R-rated movie without an accompanying adult is legal at age 17, and drinking is legal at age 21, which results in the shared salience of these age categories among parents. Similarly, churches sometimes dictate that certain activities are permissible only at certain times. Many churches prohibit sexual relations until marriage, which renders “marriage” a widely salient category among religious parents. Similarly, when discussing dating, Latter-day Saint (Mormon) parents frequently mention that their church has said that 16 is the appropriate age for children to start dating.
Institutions also often induce homogeneity in parents’ reasoning about age appropriateness by prohibiting certain activities outright. This occurs via legal means but perhaps more frequently via religious proscription (Mollborn and Sennott 2015:1287). It is no coincidence that many of the topics with the smallest diversity scores are activities that are commonly prohibited by religious authorities; the prohibitions against these activities are widely salient among religious parents, resulting in high concentrations of “never” mentions. For example, of all the salient age categories mentioned when discussing swearing, half of them were “never.” This is because even parents who allowed swearing often discussed conditions wherein it was never appropriate. Similarly, “never” accounted for 44.2 percent of the salient age categories mentioned in discussions about viewing porn. Even when parents disagree with these prohibitions, their salience is often manifest in phrases like “the church teaches never, but . . . .” In other words, institutions influence salience, but this does not preclude parents’ agency (Guhin 2016; McPherson and Sauder 2013).
If much of the observed variation in diversity scores is attributable to institutional proscription, what would happen if we bracketed these responses from our calculations? In other words, how diverse are parents’ discussions of these topics if we pay attention only to the positive age categories they mention? To answer these questions, I removed “never” from the list of categories and recalculated the diversity scores. The results, shown in Figure 4, present a different picture. Several things stand out. First, the scores for some topics, including swearing, viewing porn, and trying drugs, which had the lowest diversity scores in the original measure, became much more diverse. This suggests that absent the widespread salience induced by prohibition, there is little to align parents’ reasoning about these topics. Second, some topics, including having sex and watching R-rated movies, still show modest amounts of concentration (manifest by the gaps between Shannon and Gini-Simpson scores) because they are associated with some shared salience via standard, positive age categories (i.e., marriage for sex and 17 for R-rated movies). Third, the diversity scores for some topics, including owning a smartphone, using social media, and wearing makeup, hardly change at all because these are activities that are rarely prohibited outright. Taken together, the results suggest that institutions do play an important role in aligning parents’ reasoning about age appropriateness by banning activities and by setting standards of age appropriateness.

Effective Number of Age Categories by Topic (Excluding “Never”)
The “encompassing institutional matrix” account is promising, but there are cases in which we find a high degree of concentration without a direct, unambiguous institutional connection. For example, 29.1 percent of parents in the sample mentioned “never” when discussing having a television in the bedroom. “Thy children shall not have TVs in their bedrooms” is unlikely a common sermon, but it is clear that the prohibition is salient to many religious parents. Whither, then, does this salience arise? One possibility is that shaping institutions are responsible but in less direct ways. For example, direct religious proscriptions and warnings of pornography may push parents to think critically about television access. Similarly, if a religious institution explicitly discusses the importance of family togetherness, this may make parents more aware and concerned about the potentially atomizing effects of bedroom televisions.
In summary, calculating diversity using simple keywords can provide a depth of knowledge about collective salience. Additionally, calculating and comparing multiple diversity indices reveals subtle insights, which would be less visible when using only a single measure. Plotting diversity profile curves alongside diversity scores can provide detailed information to guide inference about why the topics or groups we are studying have the diversity scores they do.
Discussion and Conclusion
In this article, I aimed to facilitate and improve the measurement of cultural diversity in text. I did this by discussing two different frameworks for measuring meaning and by introducing and demonstrating a simple method based on word counts that is ideal for certain cases. The article has several implications for cultural analysis.
First, in terms of method, one of the challenges of contemporary cultural analysis is knowing which tool is right for the job given one’s specific problem (Mohr et al. 2020:155). When it comes to measuring cultural diversity, there are many available measures to consider and evaluate. I argued that popular measures of meaning based on the relational logic of word co-occurrences are useful for measuring cultural diversity in many cases but not all. When studying a set of known categories that are closely associated with key words or phrases, a simple word count approach is likely preferable. I then presented a simple method for measuring cultural diversity in these kinds of cases that can be used with both small- and large-scale data sets. More generally, analysts measuring cultural diversity will benefit from carefully considering their case and asking whether a simpler approach is warranted even though more sophisticated methods exist.
Second, the word count approach to measuring and studying cultural diversity gives an additional way of thinking about computational grounded theory (Nelson 2017), which combines interpretation and formal measurement to emphasize an iterative process alternating between induction and deduction. Nelson’s (2017) original model begins with inductive computational exploration, followed by deep reading and then pattern confirmation using more computational tools. The approach outlined here entails a similar back-and-forth between induction and deduction but in a slightly different order. The approach begins with qualitative induction to generate a list of keywords and uses some computation tools to evaluate the fit of the keyword dictionary. The resulting dictionary is deductively applied, and then the resulting profiles are inductively interpreted. The analysis also included some simple pattern confirmation by recalculating diversity indices after removing “never” from the keyword dictionary, although more sophisticated tools could be used here. In this way, the method presented here should not be read as an isolated measure but one tool in a set of possible methods that facilitate computational grounded theory.
Third, when it comes to formal diversity measures, I argued that using a single diversity index is limited and that analysts may benefit from following the example of ecologists and creating “diversity profiles” using multiple diversity indices (Chao and Jost 2015; Leinster and Cobbold 2012; Tóthmérész 1995). Diversity profiles are unnecessary if analysts are interested in only a single dimension of diversity (e.g., concentration), but attending to multiple indices may reveal additional insights that deepen understanding and facilitate explanation. For example, as we saw earlier, distributions could be described as exhibiting both higher concentration and more richness. Substantively, contrasting diversity indices like this may indicate a powerful but limited homogenizing agent (e.g., an institution with limited reach or power).
Fourth, in terms of substantive issues of cultural sociology, measuring cultural diversity in text with the added insight of multiple diversity indices is particularly well suited to testing Swidler’s (2013) claims about the role of institutions in inducing cultural homogeneity. Following Swidler’s hypotheses, the results of the analysis suggested that religious and legal standards and prohibition (or the lack thereof) shape the cultural diversity in parents’ reasoning about age appropriateness.
Finally, it should be noted that these methods are not intended to replace interpretation but create simplifications of the text that reveal difficult-to-see patterns and enhance interpretation (Chakrabarti and Frye 2017; Lee and Martin 2015; Stoltz and Taylor 2021). Relatedly, the word count–based approach to measuring cultural diversity presented here is not designed to replace more sophisticated methods but is an additional tool for measuring culture in text that is optimal for a subset of research cases. In sum, research on cultural diversity will benefit from specificity in conceptualization and intentionality in method, keeping in mind that the best method may very well be one of the simplest.
Research Data
sj-csv-6-spq-10.1177_01902725231194356 – Measuring Cultural Diversity in Text with Word Counts
sj-csv-6-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Research Data
sj-csv-7-spq-10.1177_01902725231194356 – Measuring Cultural Diversity in Text with Word Counts
sj-csv-7-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Research Data
sj-csv-8-spq-10.1177_01902725231194356 – Measuring Cultural Diversity in Text with Word Counts
sj-csv-8-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Supplemental Material
sj-docx-1-spq-10.1177_01902725231194356 – Supplemental material for Measuring Cultural Diversity in Text with Word Counts
Supplemental material, sj-docx-1-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Research Data
sj-r-1-spq-10.1177_01902725231194356 – Measuring Cultural Diversity in Text with Word Counts
sj-r-1-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Research Data
sj-r-2-spq-10.1177_01902725231194356 – Measuring Cultural Diversity in Text with Word Counts
sj-r-2-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Research Data
sj-r-3-spq-10.1177_01902725231194356 – Measuring Cultural Diversity in Text with Word Counts
sj-r-3-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Research Data
sj-r-4-spq-10.1177_01902725231194356 – Measuring Cultural Diversity in Text with Word Counts
sj-r-4-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Research Data
sj-r-5-spq-10.1177_01902725231194356 – Measuring Cultural Diversity in Text with Word Counts
sj-r-5-spq-10.1177_01902725231194356 for Measuring Cultural Diversity in Text with Word Counts by Michael Lee Wood in Social Psychology Quarterly
Footnotes
Acknowledgements
The author would like to acknowledge the following people for their constructive feedback on one or more versions of this article: Dustin S. Stoltz, Omar Lizardo, Marshall A. Taylor, Terry McDonnell, Christian Smith, David Gibson, Clayton Childress, Craig Rawlings, members of the Culture Workshop at Notre Dame, and anonymous reviewers.
Supplemental Material
Supplemental material for this article is available online.
1
Document similarity measures can be created using topics generated by LDA topic modeling, but they are less fine-grained than those created using WMD.
2
From here on, I use the term “keywords” to refer to both individual key words and key phrases.
3
It is possible that future methods could address these limitations. For example, perhaps someone will develop dynamic word embedding models that take into account a speaker’s current practical considerations when estimating similarities between words. However, until these methods arrive, humble keywords often get the job done.
4
If target concepts are not clearly associated with a limited set of keywords, then this increases the amount of required interpretive work and raises the problems discussed by Biernacki (2014) and
.
5
6
Interestingly, Herfindahl independently rediscovered the index after Hirschman but included it only in his unpublished dissertation. Herfindahl’s work probably would have been forgotten had his dissertation not been read by Rosenbluth, a fellow Columbia economics PhD student, who gave Herfindahl the credit for the index in an influential article (Rosenbluth 1955). After that, the index became known as the “Herfindahl” index in economics even though Herfindahl never published the index himself. Hirschman tried to set the record straight in an article published in the American Economic Review titled “The Paternity of an Index” (Hirschman 1964), and eventually his name was added to the index. At the time of the article, Hirschman did not seem to know that Simpson’s identical index, published in 1949, was making waves in ecological research. For more on the history, see Adajar, Berndt, and Conti (2019).
7
The online appendix includes demographic information about the sample.
8
I thank an anonymous reviewer for pointing this out.
Bio
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
