Abstract
Textual social media data have become indispensable to researchers’ understanding of message strategies and other marketing practices. In a new departure for the field of brand communication, this study adopts and extends a semi-supervised machine-learning approach, guided latent Dirichlet allocation (LDA), which incorporates human insights into the discovery and classification of topics. We used it to analyze tweets from businesses involved with an emerging food technology, cultured meat, and delineated four key message strategies used by these brands: providing functional, educational, corporate social responsibility, and relational content. We further ascertained the relationships between brands and the key topics embedded in their Twitter data. A comparison of model performance suggests that guided LDA can be an advantageous alternative to traditional LDA, which is characterized by high efficiency and immense popularity among researchers, but—because of its unsupervised nature—yields findings that can be difficult to interpret. The present study therefore has critical theoretical and methodological implications for communication and marketing scholars.
Keywords
The emergence and proliferation of social media have provided brands with new means of connecting with consumers. There is growing scholarly and practical interest in deciphering emerging communication and marketing practices, and discerning their influences, based on the vast array of textual data embedded in social media messages from and about brands (e.g., Ashley & Tuten, 2015; Berger et al., 2020; de las Heras-Pedrosa et al., 2020; Liu et al., 2021). Understanding branded messages can be particularly valuable for new and evolving industries, given that brand communication on social media can play a key role in shaping public perceptions in the early stage of the attitude-formation process, and in fostering stakeholder engagement (e.g., Chen et al., 2017).
However, if textual social media data are to provide useful insights, researchers must be able to reliably and efficiently extract and analyze them. Traditional approaches that work with textual data commonly rely on manual coding (e.g., Ashley & Tuten, 2015), which in the context of big data suffers from both low reliability and low efficiency (Su et al., 2017). To quickly and reliably quantify information embedded in a body of large-scale unstructured or semi-structured social media data, computational text analysis, also known as text mining (Netzer et al., 2012), and automated text analysis (Berger et al., 2020) are especially useful. Topic modeling, a subtype of computational text analysis, enables researchers to extract prominent topics discussed in textual messages such as social media posts, which can provide references to public opinion, consumer perceptions, and message strategies (Berger et al., 2020; Yun et al., 2020). Within the sphere of topic modeling, traditional latent Dirichlet allocation (LDA) has been popular among researchers for topic discovery, but has often been criticized for its unsupervised nature: that is, it derives topics strictly from data, and determining underlying topics from its results can therefore be challenging (Berger et al., 2020; Chandrasekaran et al., 2020; Ramage et al., 2009; Toubia et al., 2019; Watanabe & Zhou, 2020).
To address this drawback, the present study aims to make improvements to guided LDA, a text-analysis procedure for semi-supervised topic modeling (Jagarlamudi et al., 2012), that will make it more suitable for analyzing brands’ message strategies on social media such as Twitter. Guided LDA, sometimes known as seeded LDA, has the potential to overcome traditional LDA’s main drawbacks while remaining automated and scalable (Chandrasekaran et al., 2020; Toubia et al., 2019). It differs from traditional LDA in that, instead of completely lacking human input or theoretical guidance during its model construction, it allows researchers to use their prior knowledge and an existing theoretical framework to guide the classification of topics, via the provision of a set of seed words representative of underlying important topics (Chandrasekaran et al., 2020; Jagarlamudi et al., 2012).
The present study provides details of the refined procedures we developed for building seed-word lists and conducting guided LDA models, as well as of those models’ application to the analysis of Twitter posts from startups and small businesses that produce cultured meat and related products, which are expected to capture 35% of the growing market in meat by 2040 (Digital Food Lab, 2020). Most cultured-meat companies are currently small businesses, usually with limited budgets for marketing and public relations (Grimmer et al., 2017), and many rely on social media to raise awareness of their brands and products, reach potential customers, and build relationships with other stakeholders (e.g., Cirlugea et al., 2020). Despite the increasing prevalence of social media over the past decade having dramatically changed how businesses manage their stakeholder relationships (e.g., Taecharungroj, 2017; Tafesse & Wien, 2017), most studies of social media message strategies have focused on established, global industries (e.g., Liu et al., 2021), and seldom targeted smaller enterprises and startups.
Our study’s contributions are both theoretical and practical. With regard to the former, it provides a theory-driven examination of the message strategies employed by small businesses and startups on social media. Via effective leverage of guided LDA, these brands’ posts were classified into four main message-strategy themes—that is, functional, educational, corporate social responsibility (CSR), and relational—and various subcategories of those themes, following prior brand-communication literature (e.g., Coursaris et al., 2013; Dolan et al., 2019; Sinclaire & Vogus, 2011; Taecharungroj, 2017; Tafesse & Wien, 2017; Yang et al., 2020). In combination with brand-topic analysis informed by Rosen-Zvi et al.’s (2004) author-topic model, which allows us to analyze the salient message strategies adopted by each brand, our findings provide further insights into how a variety of cultured-meat businesses and startups use Twitter to strategically communicate about their products, promote their brand images, and engage with consumers. Our findings will help deepen understanding of social media marketing in the understudied area of new brands, and provide practical guidance for other aspiring brands.
Methodologically, our new procedure for efficiently and effectively building seed-word lists for guided LDA models integrates some of the best practices developed through previous studies. Additionally, the present study compares the performance of guided LDA against that of traditional LDA in terms of model interpretability. Although traditional LDA has been used extensively by communication and marketing scholars (e.g., Feng et al., 2020; Liu et al., 2017; Lou et al., 2019; Yun et al., 2020), our results suggest that guided LDA can be an advantageous alternative for scholars and practitioners in these fields who require automatic analysis, yet seek a systematic understanding—guided by human intelligence and conceptual frameworks—of the dominant message strategies and foci employed by their focal social media users.
Brands’ Message Strategies on Social Media
Thanks to the digitization of information, textual data are readily available. In particular, the emergence of social media has made textual messages an important data source for scholars and practitioners seeking insights into brands’ image-development, equity-creation, and stakeholder-communication strategies (Ashley & Tuten, 2015; de las Heras-Pedrosa et al., 2020; Gómez et al., 2019). Social media have also changed the landscape of businesses’ marketing and public-relations activities, by allowing companies to strategically “position the product in the mind of the prospect” (Ries & Trout, 2001, p. 2) through increasing the salience of particular brand attributes (Ragas & Roberts, 2009).
One growing line of scholarship focuses on delineating the message strategies adopted by brands on social media. An early classification of such strategies distinguishes between informational and transformational types, with the former emphasizing the factual attributes of products (e.g., comparative and generic), and the latter, the psychological aspects of experiencing them (e.g., user image and brand image; Laskey et al., 1989). Similarly, Taylor (1999) proposed a broad dichotomy between transmission-based message strategies, which focus on consumers’ need for functional, rational product information; and ritual-based ones, which center on their affective needs, such as for emotional fulfillment and resonance. Similarly, Taecharungroj (2017) classified Starbucks’ marketing-communication strategies on Twitter into three types: information-sharing, emotion-evoking, and action-inducing. Tafesse and Wien (2017) went further, using Facebook posts from selected global brands operating in a diversity of market contexts to identify 12 message typologies, for example, functional, educational, customer-relationship, and cause-related.
Challenges of Using Traditional LDA Models to Decipher Messages
To date, studies of brand messages’ textual attributes have relied predominantly either on manual content analysis (e.g., Cvijikj & Michahelles, 2013; Taecharungroj, 2017), which is prone to human coder fatigue and low reliability (Su et al., 2017), or on unsupervised computational approaches (e.g., Swaminathan et al., 2022). Among the latter, topic modeling has been widely adopted. As briefly discussed above, it can be used to extract prominent topics of discussion from documents and datasets, thus allowing researchers to discover central themes by identifying word groups rather than individual words (Berger et al., 2020; Jelodar et al., 2019). Tools commonly used to perform topic modeling include LDA, Poisson factorization (PF), and latent semantic analysis (LSA). LDA has emerged as “the most common form of topic modeling” (Yun et al., 2020, p. 51; see also Jelodar et al., 2019) due largely to two important advantages: (1) its greater accuracy than LSA when applied to Twitter data (Tijare & Rani, 2020) and (2) its scalability (Puschmann & Scheffler, 2016).
Topic modeling, a subtype of natural-language processing, is conducted using algorithms capable of identifying latent themes in massive accumulations of text (Jelodar et al., 2019). LDA, for instance, is a Bayesian learning algorithm used to extract thematic topics by grouping semantically related words based on their co-occurrence patterns (Blei et al., 2003). As an unsupervised machine-learning technique that pairs an inductive approach with quantitative measurements, LDA requires no prior annotations of documents, as topics emerge directly from the analysis (Blei et al., 2003; Maier et al., 2018). In other words, topics can be viewed as latent content-related categories that represent the text collection as a whole (Maier et al., 2018). Each document within such a collection is modeled as a multinomial distribution of topics as part of a probabilistic model, whereas a topic is represented as a multinomial distribution over words. That is, a topic is chosen from the respective topic distributions of each document in the collection. Then, a word is sampled from each distribution, over words for the topic chosen in the previous step (e.g., Hong & Davison, 2010). For a detailed introduction to the technical aspects of LDA, see Blei et al. (2003).
While LDA has been praised for the ease and speed with which it automatically detects themes in large numbers of documents, it is not without its limitations and challenges. As briefly noted above, a major critique of LDA relates to the interpretability of its models. Because of its unsupervised nature, topics simply arise from the data (Toubia et al., 2019) and thus are not always easy to delineate; and some may not be meaningful (Berger et al., 2020; Chandrasekaran et al., 2020). Indeed, Jagarlamudi et al. (2012) argued that LDA models tend “to explain only the most obvious and superficial aspects” of a collection of text documents, and to perform relatively poorly when topics are “rare” (p. 204). Providing an appropriately descriptive name for a given topic can also be challenging (Ramage et al., 2009; Toubia et al., 2019), and—because they are based solely on co-occurrences of words—topics produced by LDA models are often inconsistent with theoretical frameworks (Watanabe & Zhou, 2020). Therefore, it has become necessary to develop alternative approaches, notably ones incorporating human supervision, that allow more direct interpretability of topics and their associated labels (e.g., Toubia et al., 2019).
Guided LDA as an Alternative Topic Modeling Technique for Analyzing Brand Messages
One way to increase the interpretability of topic models’ classification results is supervised/semi-supervised machine learning that incorporates prior lexical knowledge. Guided LDA allows users to specify “sets of seed words” that are representative of the focal text collection (Jagarlamudi et al., 2012, p. 204). As compared to its traditional counterpart, guided LDA has significantly better accuracy (Jagarlamudi et al., 2012; Watanabe & Zhou, 2020) and quality (Yu et al., 2017). Moreover, because they are semi-supervised, guided LDA models allow for more efficient theory-driven text analysis (Toubia et al., 2019; Watanabe & Zhou, 2020) and can reveal rarer topics within a given dataset (Shanthakumar et al., 2020) than their non-supervised counterparts.
Another area in which guided LDA models outperform traditional ones is the creation of topic-word and document-topic probability distributions (Jagarlamudi et al., 2012). In a topic-word distribution model (e.g., Model 1 as specified by Jagarlamudi et al., 2012), each topic consists of a combination of a regular topic distribution and its associated seed-topic distribution. Regular topic distribution can generate any word, whereas seed-topic distribution can only generate words from a corresponding seed-word set. This differs from traditional LDA, in which each topic is defined by only one multinomial distribution over words. After it has been created, a topic-word distribution can be improved via the use of seed-topic distribution to classify relevant words into their corresponding regular topics. To improve document-topic probability distribution (e.g., Model 2 in Jagarlamudi et al., 2012), seed information can be transformed from words into documents that include those words. That is, every seed-word set is defined by a distribution over regular topics (i.e., group-topic distribution). Then, the document-topic distribution is informed by a sampled seed set and its group-topic distribution.
Seed-Word Selection
In any semi-supervised topic-classification approach such as guided LDA, the selection of seed words is of vital importance. Such words must be unambiguous as well as accurate reflections of the topics (Watanabe & Zhou, 2020). Conversely, seed-word sets that match irrelevant text can confound actual associations between topics and words during the classification process, and decrease algorithms’ performance (Jagarlamudi et al., 2012; Watanabe & Zhou, 2020).
There are two main methods whereby scholars select seed words. In the first, such words are identified based on one’s domain knowledge and relevant prior literature (e.g., Chandrasekaran et al., 2020; Watanabe & Zhou, 2020). More specifically, the sources used to identify relevant topics and their associated seed words range from scholarly literature and theoretical frameworks (Chandrasekaran et al., 2020; Toubia et al., 2019) and subject-matter experts’ recommendations (Toubia et al., 2019) to glossaries, books’ indexes (Watanabe & Zhou, 2020), and crowdsourced feedback from platforms such as Amazon Mechanical Turk (Toubia et al., 2019). The second main method of seed-word selection involves researchers conducting preliminary explorations of the same collections of text documents they plan to analyze. They may manually classify a subset of a dataset, known as the training set, into topics and associated seed-word terms (Yu et al., 2017); or find the most frequently occurring words in the corpus and manually classify them into relevant topics (Shanthakumar et al., 2020; Watanabe & Zhou, 2020). Also, researchers have identified topics and seed words based on traditional LDA models’ results (Chandrasekaran et al., 2020).
Application
Researchers have shown a growing interest in using guided LDA to examine large corpora of text ranging from speech transcripts to user-generated online content. For example, Watanabe and Zhou (2020) recently utilized it to classify United Nations (UN) General Assembly speeches into six main topics: greetings, the UN, security, human rights, democracy, and development. Guided LDA has also been used to analyze complaints filed by restaurants to supply-chain orchestrators as a means of identifying the main quality issues associated with supply chains, which were found to be freshness, packaging, and delivery (Yu et al., 2017). Among those who have applied guided LDA to the analysis of social media content, Shanthakumar et al. (2020) looked at tweets with COVID-related hashtags that were created during the early days of the pandemic, and identified five thematic topics: general COVID-19 information, school closures, panic buying, lockdowns, and quarantine. Despite this growing use of guided LDA, however, it has seldom been applied in the contexts of brand-communication and marketing research.
The Present Study
Employing computational text analysis to examine masses of digital text is widely regarded as reliable (e.g., Berger et al., 2020; Yun et al., 2020). Accordingly, the current study extends existing guided LDA research to the analysis of brands’ message strategies on social media. Building on previous studies (Chandrasekaran et al., 2020; Shanthakumar et al., 2020; Toubia et al., 2019; Watanabe & Zhou, 2020; Yu et al., 2017), we adopted a novel mixed approach to identifying appropriate seed words for classification of salient themes and topics embedded in brand tweets, thus incorporating human intelligence into our guided LDA model.
To better delineate the relationships between brands and their main message strategies on Twitter, we also extended our guided LDA approach to include authorship information (e.g., Rosen-Zvi et al., 2004). While a considerable quantity of existing scholarly work on brands’ social media communication focuses on the collective, industry level (e.g., Liu et al., 2021), Zhang and Su (2022) encouraged the exploration of brand-level differences on the grounds that “every brand has a unique culture and identity” (p. 16). Information about such differences may help us answer a variety of important queries, including how to distinguish among brands’ dominant product framings and which marketing strategies are being pursued by different brands. To the best of our knowledge, this has rarely been attempted.
Additionally, the current paper is intended to demonstrate the potential of guided LDA for studying emerging-business contexts—in particular, brand communication by small and/or startup businesses—in which established topic categories may not be evident, and the non-interpretability of topics derived from traditional LDA models may therefore be more of an issue. Moreover, little if any research has empirically compared the performance of guided LDA against traditional LDA. Previous LDA studies of social media strategies have focused mostly on established firms or markets occupying extensively researched categories, so that the topics that emerged from traditional LDA models could be easily interpreted (e.g., Ayele & Juell-Skielse, 2018). Specifically, this study seeks to answer the following questions and test the following hypothesis.
Methods
This study involved five steps: data acquisition, text preprocessing, building the seed-word list, running the guided LDA model, and brand-topic analysis. Figure 1 summarizes the steps, and the details of each one are discussed in the following five sections. The steps involved in data collection, text preprocessing, and analysis.
Step 1: Data Acquisition
Sampled Cultured-Meat Company Details
aCompany names and Twitter handles are those in use when the data were collected.
bNow UPSIDE Foods, with the new Twitter account @upsidefoods. Both @upsidefoods and @MemphisMeats represent the same account.
cIt is now Believer Meats, with the new Twitter handle @believermeats, created in August 2022. Data were only collected from the original account.
dThe Twitter handle is now @NewAgeEats.
eThe Twitter handle is now @futurefieldsHQ.
fThe Twitter handle is now @itsjustvow.
gThe number of collected tweets refers to the tweets collected for the data analysis.
hUnlike most other cultured-meat companies, Biftek.co, Future Fields, and Future Meat Technologies did not produce and sell meat directly to consumers at the time of data collection. Rather, the former two companies produced growth-medium supplements to grow muscle stem cells for meat cultivation, whereas the latter sold cell lines and bioreactors to manufacturers to help them scale up cultured-meat production.
Step 2: Text Preprocessing
Among the 7,942 tweets we collected, 1,441 were removed for being replies. The remaining 6,501 were preprocessed before data analysis. The first step in our preprocessing was tokenization, in which all the tweets were broken into basic units called unigrams or tokens. Then, we removed Twitter handles (i.e., @username), symbols (e.g., # and &), URLs, non-alphabetical items other than symbols (e.g., punctuation marks and numbers), and words with fewer than three characters. In addition, all words were transformed into lowercase. Next, words on the English stop-word list from the Natural Language Toolkit (NLTK) package in Python (see https://www.nltk.org/) were removed. After these data-cleaning procedures, we performed stemming, whereby all words were transformed into their stemmed forms using the NLTK Porter stemmer (Porter, 1980). Lastly, we identified a group of non-English words that commonly appeared in our dataset and removed tweets that contained any of them. Tweets that contained no words after these data-cleaning steps were also excluded from the dataset. This left 6,248 tweets for our final analysis.
Step 3: Building the Seed-Word List
Themes/Topics, Seed Words, and Top Words Captured by the Guided LDA Model, with Proportions of Themes/Topics
Note.
aVarious definite numbers of seed words ranging from single digits to more than 20 have been proposed as appropriate for textual analysis (Chandrasekaran et al., 2020; Shanthakumar et al., 2020; Watanabe & Zhou, 2020).
bThe top 20 topic words associated with the highest coefficients are listed. Consistent with previous studies (e.g., Nagai et al., 2019; Ramesh et al., 2014; Toubia et al., 2019; Watanabe & Zhou, 2020), many words derived from guided LDA were the same as seed words, but relevant topic words were also identified. The percentage of each topic under product attributes was rounded, and the sum is based on the numbers before rounding up.
cThe “Other” category refers to unseeded topics in our model, to account for tweets that did not fall into any classified topic. Two such topics were used, following Shanthakumar et al. (2020), and accounted for 9.8% and 3.9% of all tweets, respectively.
Step 4: Running the Guided LDA Model
We used the GuidedLDA package in Python to run guided topic modeling. Following Nagai et al. (2019), we set the parameters as .01 for α and .01 for η, which are respectively Dirichlet priors on the per-document-topic distribution and per-topic word distribution (Gangadharan & Gupta, 2020). Meanwhile, seed confidence, that is, the probability of biasing the selection of seed-word distribution, was set at .7 (Nagai et al., 2019). Researchers running guided LDA models have included topics in addition to the number of identified seeded topics (i.e., 12, in this case) to cover documents that did not fall under any of the latter (e.g., Li et al., 2019; Ramesh et al., 2014; Shanthakumar et al., 2020), and we followed this practice as well. After manually evaluating the interpretability of the models with varying numbers of unseeded topics, the one with two unseeded topics yielded the best categorization, which was consistent with prior findings by Shanthakumar et al. (2020). Therefore, our guided LDA model was run with 14 topics (i.e., 12 seeded topics and two unseeded ones), and trained with 1,000 iterations (Jagarlamudi et al., 2012).
To evaluate the performance of our guided LDA model, we calculated two performance metrics—a coherence score and Hellinger distance (Rosalind & Suguna, 2022; Yu et al., 2017)—and compared them against the same metrics for the traditional LDA model. Coherence measures are used to evaluate the interpretability of topics based on the co-occurrence of keywords, with higher scores indicating better interpretability by human judges (for details, see Mimno et al., 2011). We used C V , a topic-coherence measure that has been shown to outperform others in terms of correlation with human topic-ranking data (e.g., Röder et al., 2015). Hellinger distance, on the other hand, measures how far apart topics are from each other, on average, in the topic-keyword distribution. A greater Hellinger distance indicates more distinctness of and less overlap between topics, and thus that they are more interpretable (Blei & Lafferty, 2009). We set the same topic number, 14, and α and η values of .01 for both the guided and the traditional LDA models, to facilitate systematic comparison between them. The best-performing number of topics (K) in the traditional LDA model, based on C V score, was also 14.
Step 5: Brand-Topic Analysis
Based on the generated document-topic distributions, each extracted tweet was assigned to its most salient topic (Hosseini et al., 2020; Nugroho et al., 2015). Slightly more than four-fifths (81.6%) of the tweets had a prominent factor 3 of 1.4 or higher, indicating that each was dominated by one topic, as Nugroho et al. specified. Analyzing tweets’ salient topics can tell us how each topic is represented among all tweets, and allows us to characterize the relationship between brands and their dominant tweet topics. Lastly, informed by Rosen-Zvi et al.’s (2004) work, we conducted a brand-topic analysis that took account of the content of tweets and thematic interests of their authors, with the aim of understanding which companies were more likely than others to discuss certain topics. We first used the pandas package in Python (McKinney, 2010) to calculate the proportion of each topic among all tweets from each brand, and then plotted a heatmap to showcase the relationships between topics and brands using the Python package Matplotlib (Hunter, 2007).
Results
Identified Themes and Topics
In answer to RQ1, regarding the prominent message strategies used by cultured-meat brands on Twitter, our guided LDA model yielded 12 subtopics that we framed into four broad themes (see Table 2). Of the 6,248 tweets examined, 30.1% consisted of relational content, for example, (1) attempts to initiate conversations and other interactions with consumers, for example, via announcements about fundraising (8.4%); (2) dissemination of hiring information (9.2%); and (3) promotion of events such as conferences and summits (12.5%).
The second most prominent theme was functional content (22.6%): that is, factual claims, product attributes, and product benefits, including for (1) seafood products (4.5%); (2) meat products (7.8%); (3) the complementarity of cultured meat with plant-based diets as an alternative protein (6.5%); and (4) the health and nutritional aspects of cultured meat (3.8%).
This was followed by the theme of CSR marketing (20.4%), which highlighted how the brands achieved social and environmental imperatives, with the goals of building positive brand images and stimulating affection and other positive emotions among consumers looking to boost sustainability via their purchases. This primarily involved articulating how cultured-meat products (1) help achieve environmental and food-system sustainability (9.3%) and (2) enhance animal welfare (11.1%).
Lastly, 13.4% of the final sample of tweets delivered information aimed at educating and informing consumers. Its three main subtopics were (1) products’ manufacturing process (2.8%), (2) regulatory frameworks (3.8%), and (3) industry trends and market developments (6.8%).
Brands and Their Dominant Topics
The heatmap shown in Figure 2 was created to answer RQ2, regarding the relationship between brands and their respective predominant tweet foci. Each cell in the heatmap represents the proportion of a given brand’s tweets that reflected a particular message strategy. The darker a cell in the heatmap is, the more likely a focal company was to post about the corresponding topic. The heatmap suggests that most of the 19 focal brands placed a relatively equal emphasis on the identified strategies. Nevertheless, tweets posted by SuperMeat primarily focused on the animal-welfare aspect of cultured meat, emphasizing how this emerging technology can decrease the number of animals slaughtered. Brands that produced lab-grown seafood, including Wildtype and Shiok Meats, tweeted predominantly about seafood; and similarly, a high portion of tweets from Aleph Farms, which specialized in the development of cultured beef, were found to contain meat-related information. Tweets from several brands, notably New Age Meats, Mosa Meat, Future Fields, and Vow, reflected a focus on relational content.
4
Heatmap representing the relationships between brands and their dominant tweet topics, as measured by that topic as a proportion of all their tweets, with darker colors indicating higher proportions
Evaluation of the Guided LDA Model
Our C V and Hellinger-distance results suggest that, in terms of interpretability, our guided LDA model (14 topics, C V = .7881, Hellinger distance = .8141) was preferable to the traditional LDA model (14 topics, C V = .3497, Hellinger distance = .7085). The results from the traditional LDA model further suggest its low interpretability, as the focus of most of its topics cannot be clearly discerned (see Appendix 1). Therefore, H1 was supported.
Discussion
Theoretical Contributions
To the authors’ knowledge, this is the first empirical attempt to apply guided LDA to the examination of brands’ social media message strategies in the context of an emerging industry. By using guided LDA to analyze tweets from the cultured-meat industry, it has provided one of the first views of the landscape of social media message strategies employed by small businesses and startups. These strategies reflect a mixture of informational, emotional, social, and relational values (e.g., Dolan et al., 2019; Sinclaire & Vogus, 2011; Taecharungroj, 2017; Tafesse & Wien, 2017), which may be driven by consumer needs for product- and service-information acquisition, identity formation/projection, and social interaction when participating in brand-related social media use, in line with uses and gratification theory (e.g., Muntinga et al., 2011).
In line with prior scholarship showing that informative content such as product attributes and technical details can be a key factor in consumer engagement in online brand communities (e.g., Cvijikj & Michahelles, 2013), around 40% of the tweets we sampled focused on functional and educational content; and within that set of tweets, almost two-thirds highlighted products’ functional attributes. This strong emphasis on product information may also reflect cultured meat’s status as an emerging industry with little consumer familiarity and low product availability (e.g., de Oliveira Padilha et al., 2022; Tan, 2021).
It is important to note, however, that the sampled cultured-meat brands also actively incorporated CSR information into their social media strategies, presumably to depict themselves “in a positive light” (Tafesse & Wien, 2017, p. 18). Tweets reflecting this approach proactively and strategically emphasize companies’ relevance to issues of societal concern (Sinclaire & Vogus, 2011), such as environmental and food-system sustainability and animal welfare, echoing prominent frames of news coverage and public discourse about cultured meat (Laestadius & Caldwell, 2015; Painter et al., 2020). CSR campaigns on social media have been linked with more positive brand impressions, such as higher levels of brand equity (Yang et al., 2020), and this may be related to social media users’ motivation for supporting social causes (Tafesse & Wien, 2017) and self-fulfillment (Muntinga et al., 2011).
On the other hand, our findings that around a third of the sampled tweets primarily focused on sharing corporate news and brand-organized events might reflect an intent by brands to establish relationships with stakeholders and members of the public. One such embodied practice is based on facilitating interaction with consumers via encouraging participation in brand events (Taecharungroj, 2017). Probably because of the sampled cultured-meat brands’ shared status as players in a new market, our analysis also identified the sharing of fundraising updates and successful outcomes as a strategy for building relationships and creating a sense of community. Building consumer trust has been found to be directly related to enhanced consumer commitment (Kwan Soo Shin et al., 2019), so such relational content may be critically important to new businesses’ creation of positive emotions among existing and potential customers.
Methodological Contributions
This study has provided detailed steps for utilizing guided LDA models to computationally identify prominent branding strategies used on Twitter by small businesses and startups, building upon existing theoretical frameworks and integrating human insights. In particular, it illustrates key considerations and practices for creating seed-word lists, which are essential in guiding LDA models and can profoundly affect the quality of data analysis. There have hitherto been no uniform guidelines on how to build a seed-word list for topic classification in guided LDA. This study therefore summarized and illustrated how to use domain knowledge and traditional bigram models to find prominent seed words, in a mixed approach that combines the established scholarly methods of relying on subject-matter expertise (Toubia et al., 2019) with exploration of text collections or their subsets (Shanthakumar et al., 2020; Yu et al., 2017). One unique contribution of this study is that it combined a traditional bigram model run on a small subset of data (i.e., a training set) with domain knowledge to identify themes and seed words. Using a bigram model was advantageous because it helped us identify the dominant themes in the training set via providing a list of frequent terms with better interpretability than a unigram model would have. Adopting bigram models therefore enables researchers to build seed-word lists both objectively and relatively quickly. Our findings also show that the interpretability of guided LDA’s findings was much better than that yielded by traditional LDA, despite the small number of additional steps the former required. This confirmed the efficiency and efficacy of our approach.
This study has also showcased steps for performing brand-topic analysis that will help researchers delve into inter-brand differences in branding strategies. Hardly any previous research on brand communication on social media has distinguished among the brands within its focal industry. While we found that brands in the cultured-meat industry largely employed similar Twitter branding strategies to one another, there were also discernible differences among them in terms of topic emphasis. Such findings suggest the importance of not always treating brands in the same industry monolithically, as they may have distinct stakeholder-communication agendas and strategies, and because consumers may express a wide variety of attitudes toward them (Liu et al., 2017).
Practical Implications
Guided LDA has also clear advantages for scholars and practitioners in the fields of communication and marketing. First, its semi-supervised nature allows researchers to perform topic modeling with the guidance of human insights by developing externally valid seed-word lists, and thus to define classification categories that are consistent with pre-existing theoretical frameworks and/or practical considerations. Moreover, guided LDA’s flexibility allows other relevant thematic dimensions to appear, even if the researchers do not initially suspect their presence (Toubia et al., 2019). For example, by changing seed confidence, researchers can tune their guided LDA models to classify tweets based not only on the selected seed words, but also on words’ co-occurrence patterns. Following Nagai et al.’s (2019) recommendations, we set our seed confidence at .7, but future studies could usefully explore the potential impact of other seed-confidence levels on model performance.
Second, guided LDA maintains the efficiency and speed advantages of traditional LDA while overcoming its well-attested limitation of low topic interpretability. As the present study has demonstrated, better interpretability is particularly helpful when analyzing thematic patterns of brand communication, particularly in the form of textual messages from emerging product categories or startup companies. Specifically, we found that guided LDA models had better performance than traditional LDA ones in terms of both topic coherence and interpretability, even when the latter were provided with the best-performing number of topics.
Third, we have demonstrated that guided LDA can be automated, and is thus scalable. That is, a verified set of seed words may be used repeatedly in guided LDA analysis of text corpora from varied social media platforms or other sources of information. A possible collateral benefit of such an approach is that it would allow longitudinal or real-time tracking of brand strategies or consumer attitudes, as well as systematic cross-platform comparisons.
Limitations and Future Directions
While we endeavored to maximize this study’s quality by adopting appropriate procedures and validation techniques in line with prior scholars’ recommendations, it is not without its limitations and areas ripe for further exploration. First, although our developmental process for a seed-word list was consistent with existing approaches (Chandrasekaran et al., 2020; Shanthakumar et al., 2020; Toubia et al., 2019; Watanabe & Zhou, 2020), and our model’s superior performance relative to traditional LDA was confirmed, the procedure we demonstrated is only one possible way of developing a seed-word list. Therefore, future researchers could usefully explore alternative approaches and compare their model results across two or more of them.
Second, in this research, we examined branding strategies from an emerging industry, cultured meat, to illustrate how to use the proposed guided LDA procedure to analyze branded messages on Twitter. We believe that the proposed analysis procedure, and particularly its approach to seed-word list development, would also be applicable in wider research settings; yet, the degree of model improvement of guided LDA over traditional LDA may vary across brands, industries, and technological foci of interest. Future studies should therefore apply our approach to compare brand communication in both established and other emerging industries. To fully understand and extend this approach’s applicability and utility, larger datasets with a higher volume of textual content or multiple case studies should also be analyzed.
Conclusion
Using a dataset comprising thousands of Twitter posts created by cultured-meat brands in multiple countries, this article illustrated an in-depth guided LDA procedure for examining and delineating prominent message strategies employed by an emerging industry and by the prominent brands within it. Our primary categorization of the sampled tweets’ message strategies represents a comprehensive framework that complements the existing brand-communication literature and provides additional insights into the social media practices of small businesses and startups. Our findings are also likely to have important managerial implications for the development and adaptation of marketing practices on social media.
Methodologically, the semi-supervised nature of guided LDA was found to yield better model performance than traditional LDA via the integration of human intelligence and pre-existing theoretical and practical knowledge. While semi-supervised models require knowledge of both methodology and subject matter, this paper has demonstrated their clear methodological advantages; and we hope that its set of recommended procedures will encourage communication and marketing researchers to apply guided LDA in their analyses of social media messages.
Footnotes
Acknowledgments
Leona Yi-Fan Su, Yee Man Margaret Ng, and Yi-Cheng Wang would like to acknowledge funding support from the Center for Digital Agriculture at the University of Illinois Urbana-Champaign. Dr. Su and Dr. Wang would also like to acknowledge funding support from the United States Department of Agriculture (2022-67024-36149).
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Notes
Author Biographies
Appendix
Results of Traditional Unigram Topic Modeling With 14 Topics aTopics were labeled as “Not applicable” if no underlying theme could be identified. bThe stemmed form of “cms,” which refers to “cultured meat symposium.” Because we removed words with fewer than three characters before stemming, “cms” was kept, and then transformed to “cm” as part of our data-preprocessing step.
Number
Topic
Top 20 Words
1
Not applicable
a
cleanmeat, steak, foodtech, cellag, futureoffood, cultivatedmeat, world, culturedmeat, first, farm, cellbasedmeat, cm,
b
biotech, food, via, meat, cellularagricultur, make, cell, grown
2
Not applicable
goodfoodconfer, meat, food, plantbas, billion, get, burger, new, sustain, protein, industri, sept, talk, great, impact, anim, see, acceler, futur, action
3
Not applicable
meat, product, sale, lab, grown, like, launch, interview, wrap, cell, space, plant, replac, star, cleanmeat, cultiv, protein, base, learn, super
4
Fundraising
meat, food, invest, base, new, startup, global, product, cell, plant, isra, compani, via, rais, clean, fund, sustain, protein, million, world
5
Not applicable
food, meat, futur, base, industri, technolog, scienc, excit, look, biolog, investor, plant, innov, cleanmeat, tech, clean, program, cultur, seafood, foodtech
6
Not applicable
meat, aleph, food, year, age, new, farm, job, anim, set, use, happi, autom, per, impact, approach, product, take, like, global
7
Not applicable
sustain, mission, team, excit, via, meat, world, say, achiev, wild, synthet, industri, challeng, scale, manufactur, differ, parti, consum, vast, caus
8
Not applicable
meat, cleanmeat, futurefood, anim, use, cultivatedmeat, food, less, plantbas, cell, protein, muscl, product, grow, opportun, free, public, berkeley, alt, talk
9
Not applicable
thank, congratul, scienc, award, meat, agtech, focu, standard, includ, come, gfi, open, end, million, coronaviru, professor, industri, univers, progress, fight
10
Not applicable
meat, food, futur, cell, protein, join, new, base, anim, cultur, dairi, ceo, industri, cleanmeat, farm, sustain, shrimp, week, milk, plant
11
Not applicable
meat, chicken, cell, base, cultur, year, egg, new, market, cleanmeat, food, produc, consum, engin, demand, world, predict, move, restaur, term
12
Sustainability
meat, anim, eat, fish, beef, climat, livestock, cow, goodfood, chang, tast, percent, antibiot, environment, world, cleanmeat, would, human, water, tackl
13
Not applicable
meat, cell, anim, base, everi, sign, year, seafood, speci, clean, engin, today, earth, effort, huge, grate, cleanmeat, agricultur, product, world
14
Not applicable
meat, futureoffood, year, cellag, cleanmeat, plant, base, cultur, culturedmeat, ago, compani, beyond, march, seed, talk, think, cell, first, soon, agricultur
