Abstract

Today, both marketing practitioners and researchers have access to a wide variety of multimedia data, such as combinations of text from social media posts and online reviews, images and videos posted on Instagram or Snapchat, and audio data from customer service interactions. Despite the potential of multimedia data to provide rich insight for marketing research and marketing practice, the field has only just begun to tackle its formidable theoretical and methodological challenges and benefit from the substantive insights it allows us to uncover. This special issue intends to bring together cutting-edge research using one or more types of multimedia data to address both methodological and substantive topics related to the broad discipline of marketing.
In the marketing literature, the term “multimedia” has been used to describe both marketing messages from multiple media sources, such as television, radio, and newspaper (e.g., Danaher et al. 2020; Naik and Raman 2003), and information processed using multiple modalities, such as auditory and visual processing (Tavassoli 1998; Tavassoli and Lee 2003). Although in some cases media are equivalent to modalities (e.g., radio relies only on auditory processing), in other cases the same medium uses multiple modalities (e.g., television advertisements rely on both auditory and visual processing).
For analyzing multimedia data, both medium and modality are critical distinctions. If images and audio come from a single media source, such as a YouTube video, it is important to acknowledge their connection in the analysis. The visual and auditory information in the video derive their meaning, in part, by the context each modality creates for the other, making their shared source an important part of the analysis. At the same time, when analyzing multimodal data, it is important to recognize that visual and auditory data may have distinct properties, such as signal-to-noise ratio, and that data from each modality may not contribute equally to the perceived meaning of the whole (Mehta 2018). Further, because visual processing and auditory processing are distinct, the use of multiple modalities rather than a single modality can enhance learning (Mayer 2001; Tavassoli 1998), though both synergies and interference may be observed (Tavassoli and Lee 2003).
In this introduction to the special issue, we first explore, from a theoretical perspective, the implications of multiple media and multiple modalities for marketing messages and data. In Table 1, we highlight the articles in this special issue and selected other works in the literature that aggregate data across multiple media, those that leverage multimodal data, and those that leverage multimodal data from multiple media. Next, we highlight opportunities offered by multimedia data for the field of marketing, including expanding the field’s methodological capabilities and addressing important substantive challenges.
Examples of Articles Using Data from Multiple Media Sources and Multiple Modalities
Indicates articles in this Multimedia Special Issue.
Theoretical Development
Both the medium from which data originate and the modality are critical to analyzing and interpreting multimedia messages and data. In this section, we discuss the implications of both medium and modality.
Why Does the Medium Matter?
Research suggests that the medium through which a message is communicated can influence its effectiveness. For example, Venkatraman et al. (2021) illustrate that consumers use more relational processing when encoding ads presented in print format than in digital format, potentially enhancing their memory retrieval when exposed to ad cues during purchasing decisions. Danaher et al. (2020) tracked the number of brand ad exposures across multiple media, including email, catalog, and paid search, and show that while all three media were effective in increasing sales, digital media (email and paid search) were more effective for increasing the brand’s online sales, while catalogs were more effective for increasing in-store sales.
Why might the effectiveness of marketing communications vary across media? Media richness theory suggests that media sources vary in their ability to transmit cues, with “rich media,” such as television ads, transmitting more cues simultaneously (e.g., both verbal cues, such as words, and nonverbal cues, such as facial expressions) than “poor media,” such as a text-only email (Kahai and Cooper 2003). The term “media richness” combines several dimensions, including the number of cues, the modality of the cues, immediacy of transmission, and interactivity of the message (i.e., the potential for feedback and personalization; see Kahai and Cooper 2003). A rich media source, such as face-to-face communication, enables feedback and includes nonverbal cues, such as tone and facial expressions, whereas email does not; thus, face-to-face communication should be more effective than email for messages that have the potential to be unclear or confusing (Daft and Lengel 1986). In today’s digital advertising environment, the term rich media refers to digital ads with “features like video, audio, or other elements that encourage viewers to interact and engage with the content.” 1
Notably, although early work on media richness suggests that the persuasiveness of messages decreases progressively as media richness decreases (e.g., from television to radio to text-only formats such as newspaper), because less rich media transmits less information to consumers (Keating and Latane 1976), later work suggests there are also costs of media richness. Advertising presented in audio or video formats may increase consumers’ attention to communicator cues compared with text-only ads, but focusing on the communicator can reduce attention to information about the product being advertised (Chaiken and Eagly 1983). As we discuss in the next section, combining some modalities may be more detrimental to cognitive processing than combining other modalities (Tavassoli and Lee 2003).
Beyond its richness, the medium also influences the structure of the data, the source of the data, and the norms observed when using the medium and interpreting data from the medium. Structural factors may include whether the data are organized into variables, whether a medium is multidirectional (allowing feedback rather than only one-way transmission), the constraints of the medium (e.g., Twitter allows only a certain number of characters), and the number of modalities that are used (e.g., some media allow for multimodal data while others do not). For example, when a consumer creates a review on Amazon, they are identified as the source of the data and prompted to rate the product on a 1–5 scale, perhaps rate specific product features, create a headline, add a photo or video, and provide a written review.
The medium of the message often restricts the source of the data, meaning that the medium and the source are nonindependent. The source of multimedia data may be a consumer, a firm, or a third party. For example, Amazon identifies some data on its website as being produced by consumers, such as consumer-generated reviews; other data, such as product prices and descriptions, are provided by Amazon or by third-party sellers. However, the same medium may be used by multiple sources: either a firm or a consumer can post a video on YouTube, firms may respond to online customer reviews (Wang and Chaudhry 2018), or firms may respond on Twitter when a consumer tweets about a service experience with the firm (Golmohammadi et al. 2021). Notably, although Amazon chooses to post consumer-generated reviews on its website, consumers may interpret this consumer-generated content differently because it appears on the seller’s website rather than on a third-party platform. The relationship between the medium and the source is important to consider because the source of a message is perceived to be informative by those receiving the message, and it guides their processing of the message. Research on persuasion suggests that source credibility is a critical factor in the persuasiveness of a message (Wilson and Sherrell 1993). The same data (e.g., text stating, “This product is great”), if provided by a firm (e.g., via text on the firm’s website) versus by a consumer (e.g., in the text of an online review hosted on a third party’s platform), may be interpreted quite differently on the basis of the reader’s understanding of the inferred motives of the firm and the consumer (Lantzy et al. 2021).
Finally, the medium is important because norms that guide usage of the medium should be considered when interpreting data from the medium. For example, Twitter is often used to lodge complaints, especially about service experiences (Golmohammadi et al. 2021), whereas Facebook and Instagram tend to highlight more positive experiences (Dholakia 2018). Given these divergent norms, the same message describing a customer experience may be interpreted differently by other consumers depending on the platform on which it is posted (the medium). Analogously, Melumad, Meyer, and Kim (2021) show that retelling a story leads to systematic differences between the initial news story and the retold versions of the story; anticipating these differences will allow those reading the stories to draw more accurate conclusions.
Why Does the Modality of the Data Matter?
Consumers experience the environment using several sensory systems, including visual, olfactory, auditory, gustatory, and haptic systems, and each system processes data in a different modality. Consumers tend to encode modality-specific content when they process information (Unnava, Burnkrant, and Erevelles 1994). For example, when they process auditory data, consumers may encode the order of the data (e.g., which lines of a song come first, second, and third), whereas when they process visual data, they may encode visual elements such as the font style (Unnava, Burnkrant, and Erevelles 1994).
Although each of these sensory systems is distinct, the processing of sensory inputs across modalities tends to be interconnected and interdependent (for a review, see Biswas and Szocs [2019]). For example, participants were better able to distinguish a visual stimulus from distractors (an object segregation task) when an auditory tone was presented simultaneously (Vroomen and De Gelder 2000). A great deal of research has examined the relationship between visual and auditory processing, especially in the domains of learning and memory. When consumers process information in multiple modalities, prior work has demonstrated greater integration in memory between pieces of information that are encoded using similar processes (Tavassoli 1998). For example, experiments conducted with Korean consumers who are fluent in both Hangul (a phonological script relying on auditory processing) and Hancha (a logographic script relying on visual processing) show that auditory brand identifiers (sonic logos) are more strongly integrated with words written in Hangul, whereas visual brand identifiers (visual logos) are more strongly integrated with words written in Hancha (Tavassoli and Han 2001).
When information is presented using more than one modality, consumers may simultaneously use multiple cognitive processing channels (e.g., visual and auditory; Tavassoli and Lee 2003), but information processed using distinct channels must be reassembled to arrive at an integrated solution (Penney 1989), meaning that there are trade-offs to consider. Multiple input modalities can impair information processing when used to convey descriptive information (e.g., a list of facts) but facilitate processing when used to convey explanative information (i.e., information concerning relationships between facts; Lim and Benbasat 2002). Further, matching the modality of retrieval over time leads to greater consistency in the attitudes expressed by consumers, suggesting that response modality may influence the manner in which attitudes are represented (Tavassoli and Fitzsimons 2006). Building on this literature, Zhou et al. (2021) examine learning outcomes in response to video courseware that combines auditory and visual modalities.
When analyzing multimodal data, it is important to recognize that data in different modalities (e.g., text data, auditory data, visual data, spatial data) may have distinct properties, such as signal-to-noise ratio, and may be given different weights in interpreting the whole (Mehta 2018). Thus, there may be dissimilar steps required to prepare data from different modalities for analysis or assemble data across modalities. For example, before data from electroencephalography can be meaningfully analyzed, ocular artifacts from eye movements and blinks must be removed (Singh and Wagatsuma 2017). Less extensive preprocessing is required when analyzing consumer-generated online reviews, but researchers may want to distinguish reviews posted by verified customers from those that are not verified, or reviews suspected to be “fake” versus those that seem to have been created by real customers (Anderson and Simester 2014). Recent research suggests that nonverbal cues, such as the source’s review-posting history and other social interaction activity on the platform, can be more useful than the text of the review (verbal cues) for detecting fake reviews (Zhang et al. 2016).
A second issue is that data from different modalities may contribute unequally to the perceived meaning of the whole (Mehta 2018). A consumer who has processing resources available may be influenced more by visual information than information in other modalities because humans tend to prioritize visual information processing (Jia, Shiv, and Rao 2014). Consistent with prioritization of visual information, photographs presented as part of a Facebook profile had more impact on judgments of extraversion than textual self-disclosures, holding constant other characteristics of the profile (Van der Heide, D’Angelo, and Schumaker 2012). In contrast, a consumer who is unable to pay close attention to visual content may be influenced more by information in other modalities. Research shows that even short-term visual deprivation can improve sensory processing from other modalities, such as increasing accuracy in identifying the location of sounds (Lewald 2007). Indeed, when textual information was presented alone as part of a Facebook profile, this verbal information more strongly influenced judgments of extraversion than visual information alone (photographs; Van der Heide, D’Angelo, and Schumaker 2012).
Modalities also differ in the degree to which there is an agreed-upon structure for data in that modality. For example, there are several clearly defined features in auditory data (e.g., pitch, rhythm, volume) and spatial data (e.g., latitude, longitude, altitude), whereas researchers are still defining the features of images (e.g., pixels, contrast, brightness), and these features have less ability to communicate emergent properties such as “what the image is” than the features of auditory or spatial data. Because much of the data generated via social media platforms is unstructured, we see extensive use of techniques designed to structure the data, such as natural language processing techniques to structure textual data. For example, the articles in this special issue by Humphreys, Isaac, and Wang (2021), Melumad, Meyer, and Kim (2021), Toubia (2021), and Lee (2021) rely heavily on natural language processing techniques to derive insights from unstructured textual data. Notably, although machines may be able to detect subtle patterns in large data sets (e.g., interpreting medical scans) and identify emergent properties more effectively than human processors, whether consumers can detect the features of the data (e.g., pitch, brightness) may bias the identification of features.
Differences in the underlying structure of the data make it challenging to integrate data across modalities. As we discussed previously, the researcher may use the medium and source of the data (e.g., a common platform, the same consumer-generated online review) to link data from different modalities in a meaningful way. In other cases, time (e.g., a visual image presented at the same time as an auditory signal in a video), shared spatial location (e.g., geotags on apps) or a shared unit of analysis can be used to link data across modalities. For example, Chen et al. (2021) link location data with eye-tracking data to better understand where consumers are looking during a shopping trip. Boughanmi and Ansari (2021) develop a machine learning framework that allows them to combine multiple types of data (e.g., ranking data, metadata, acoustics, user-generated textual data) for a specific musical album or playlist.
Opportunities for Multimedia Data
By integrating data across media sources and across modalities, researchers will be able to derive nuanced and reliable inferences about consumers and the effectiveness of marketing strategies. Data from different modalities often provide complementary insights, and, to the extent that results converge across multiple modalities and media sources, confidence in conclusions increases. For example, Boughanmi and Ansari (2021) combine metadata, acoustic features of songs, and user-generated textual data to predict the success of musical albums and playlists. Lee (2021) combines data from Twitter, Instagram, and lab studies to examine status branding. Hartmann et al. (2021) combine data from multiple media (Twitter and Instagram) and multiple modalities (visual imagery and text) to better understand consumers’ relationships with brands.
One of the opportunities multimedia data present to the field of marketing is the incentive to engage in methodological innovation. The immense volume of data being generated by consumers, devices, and firms across modalities and media sources suggests the need for methodological innovations that enhance insights into substantive marketing problems. In the next section, we discuss methodological challenges of multimedia data, and we organize our discussion by stages of the research process: data gathering, analysis, and interpretation.
A second opportunity multimedia data present to the field of marketing is the ability to shed new light on substantive marketing problems. We highlight several of these substantive challenges in this editorial: sentiment-based targeting, location-sensitive targeting, computers as consumers, and privacy. The articles in this special issue highlight others, such as understanding the consumer decision journey (Humphreys, Isaac, and Wang 2021), consumer–brand relationships (Hartmann et al. 2021; Lee 2021), online learning (Zhou et al. 2021), and word of mouth (Melumad, Meyer, and Kim 2021).
Methodological Implications
Research based on multimedia data is different not only due to the nature of the data but also due to the methods that analysts utilize as well as the opportunities and challenges these methods present. We organize this discussion in terms of the key stages of the research process—namely, data gathering, analysis, and interpretation.
Data Gathering
Data has long been the fuel that has powered academic marketing research. The domain of multimedia research has benefited greatly from the public availability of vast amounts of data, as well as the development of modern technologies to scrape websites and use application programming interfaces (APIs) to access the data. In this special issue, we find several examples of creative approaches to obtain and combine such public data. Boughanmi and Ansari (2021) obtain data from four different sources: scraped rankings of best-performing albums from the Billboard magazine website, acoustic features from the Spotify API, textual tags from the Last.fm API, and music genres from the Discogs API. Hartmann et al. (2021) obtain data from a vendor that has Twitter-firehose access to a random sample of 10% of all tweets. Lee (2021) obtains over 160,000 tweets of luxury brands of shoulder bags, which include over 91,000 images, and scrapes the official websites of the brands to obtain prices. Toubia (2021) downloads scripts and synopses of 858 movies from the Internet Movie Database as well as full text and abstracts of articles published in several top marketing journals.
Other authors in this special issue obtained large proprietary multimedia data sets. Zhou et al. (2021) obtain online course videos and anonymized individual-level viewing records from MasterClass, a large online education platform based in the US. Toubia (2021) collaborates with a media company to obtain closed captions for a large set of TV show episodes. Humphreys, Isaac, and Wang (2021) obtain data on search queries that drove traffic to websites of leading CPG brands.
In contrast, a few articles in the special issue rely largely on primary data. Chen et al. (2021) use ambulatory eye-tracking technology to obtain detailed information about both where a shopper is located in a grocery store and visual fixations during the shopping trip. Melumad, Meyer, and Kim (2021) capture the full text of sequential retelling of stories by almost 11,000 participants in ten experiments.
We hope and expect data availability will continue to grow and nourish further research productivity. Yet, two burgeoning issues are notable. First, in recent years, concerns about consumer privacy have grown dramatically; subsequently, we discuss in depth the special privacy sensitivities of multimedia data. Second, several legal and ethical concerns have emerged around the procurement and use of web-scraped data for research. Academic journals are starting to develop policies to regulate the use of web-scraped data in submitted manuscripts, and we know that the American Marketing Association is actively working on this issue.
Data Analysis
Multimedia data are often voluminous, unstructured, and high dimensional, calling for the use of analytical methods that are well-suited to these characteristics. These methods draw from multiple disciplines, including computer vision, economics, machine learning, and natural language processing. Common goals of the analysis are extraction of features (e.g., Zhou et al. [2021] obtain features of video content); summarization or dimensionality reduction (e.g., Toubia [2021] develops a Poisson factorization topic model to summarize text); characterization of the content (e.g., Melumad, Meyer, and Kim 2021), such as concreteness (e.g., Humphreys, Isaac, and Wang 2021) and emotionality (e.g., Lee 2021); and classification and prediction (e.g., Boughanmi and Ansari 2021).
The unique characteristics of multimedia data sometimes imply that analytical techniques need to be creatively combined to accomplish more than one goal. For instance, Boughanmi and Ansari (2021, p. 1034) face data that are “structured, unstructured, discrete, continuous and textual … and high dimensional.” To summarize the semantic content of a voluminous set of textual tags, they first use a supervised hierarchical Dirichlet process to infer latent album themes. The output of the summarization process—the themes—then enter a predictive model as covariates. Hartmann et al. (2021) use deep learning to fuse image mining and text mining. Zhou et al. (2021) combine features extracted from video data with the “speaking rate” from audio data and sentiment analysis based on subtitles. Thus, as noted previously, the challenges posed by multimedia data are spurring innovation in modeling techniques.
The novelty of multimedia data in marketing has another important implication. While some articles that employ multimedia data are intended to test theory (e.g., Hartmann et al. 2021; Lee 2021), others aim to build and validate predictive models (e.g., Boughanmi and Ansari 2021; Zhou et al. 2021). In the latter case, there may be neither well-developed theory nor prior empirical literature about the expected effects. As a result, researchers prefer to specify models that allow flexibility in the effects (e.g., the use of penalized splines in Boughanmi and Ansari [2021]) to “let the data speak.” Similarly, they may favor nonparametric and semiparametric models (e.g., the use of gradient boosting machines in Zhou et al. [2021]) relative to parametric models. Many machine learning models have these desirable features.
Increasingly, we find that researchers rely on proprietary algorithms for analyses of multimedia data. For instance, Lee (2021) uses Microsoft Azure’s cloud-based Face API to analyze the emotionality of faces in over 91,000 images and to quantify variables such as happiness, neutralness, and sadness. This algorithm is one among several commercial black-box options such as Google Cloud Vision and Amazon Rekognition. Similarly, Hartmann et al. (2021) use a proprietary machine learning solution provided by the data vendor that identifies images that contain brand logos. The availability of these algorithmic tools has put powerful analytic capabilities in the hands of researchers whose goal is to address substantive questions without being encumbered by the need to develop the tools themselves. In our view, this is a positive development for marketing academia because it expands the methodological capabilities of many researchers dramatically. However, because researchers, reviewers, and readers are unable to directly examine the inner workings of the black boxes, the onus is on the researchers to provide adequate evidence to validate the outputs of these algorithms. For example, Lee validates the results of the Face API by comparison with human-coded emotionality ratings in a sample of Instagram brand posts. Similarly, Hartmann et al. employ human coders to inspect a random sample of images to identify false positives. As the use of proprietary algorithms becomes increasingly common, journals may need to develop policies about forms of validation with which authors will need to comply.
Interpretation
As we have discussed, many machine learning models have features that make them suitable for multimedia data. However, these models are often severely handicapped in terms of interpretability, a characteristic that researchers in marketing have traditionally valued heavily. Some articles in the special issue tackle this problem. For instance, Zhou et al. (2021) want to understand the importance of different video features in predicting consumption (how much of a video is watched). To do so, they examine both classic approaches (feature permutations) as well as newly developed frameworks (Shapley additive explanations). Similarly, Hartmann et al. (2021) employ gradient-weighted class activation maps to provide post hoc interpretation of the aspects of an image that played an important role in the classification task.
Interpretability of model outputs is important to assess their face validity, which generates user trust, as well as to obtain insights into how to improve a model’s predictive accuracy. We believe this is an area in which further research will be especially important to allow the models to gain greater traction among marketing practitioners.
Substantive Challenges
An additional opportunity presented by multimedia data is the ability to shed new light on substantive marketing problems. We highlight several of these substantive opportunities and challenges here: sentiment-based targeting, location-sensitive targeting, computers as consumers, and privacy.
Sentiment-Based Targeting
As the popularity of social network platforms such as Facebook, LinkedIn, WeChat, Twitter, YouTube, Twitch, and TikTok increases, collectively these platforms spew voluminous multimedia (e.g., consumer presence on multiple social media sites) and multimodal (e.g., text, images, video) data in the form of user-generated content and firm-generated content. These data are fodder for sentiment analysis with a focus on analyzing opinions, evaluations, attitudes, and emotions.
Sentiment analysis methods have grown well beyond using textual information to include all forms of nontextual information such as audio, images, and video. For example, the analysis of emotions based on images containing facial expressions posted on social media platforms is a rich area (as previously noted, Zhou et al. [2021] uses FACE to do this). Affectiva (www.affectiva.com) has built an extensive database of emotional reactions to 53,000 ads over 90 countries and eight years using artificial intelligence classification of facial images. Dupré et al. (2020) provide a performance comparison of eight commercially available automatic classifiers for facial affect recognition using both images and videos.
The revenue model for many live-streaming social media platforms relies on viewer engagement (Lin, Yao, and Chen 2021; Lu et al. 2021). Consequently, the analysis of user engagement with live video that appears in the form of emojis, texts (e.g., chat rooms), and tips becomes critical. Insights that emerge from such analysis are likely to be instrumental in the success of these new platforms.
Social media platforms allow users to connect with each other and thus develop social networks, which aid in propagating sentiments. Analysis of multimedia data can help identify the features that lead to sentiment flow. For example, Berger and Milkman (2012) identify features of online content, such as news stories, that make them more likely to “go viral.”
Location-Sensitive Targeting
Advances in geographic positioning systems and mobile technology are leading to the emergence of location-aware multimedia data. For example, a consumer who carries a mobile communication device with enabled location services on multiple apps may generate data from several media sources (the multiple apps) and in several modalities (e.g., graphics, images, videos posted on social media) with location and time stamps. Geosocial data not only incorporates the exact coordinates (latitude–longitude) of the user but also includes the geographic surroundings of those coordinates.
It is important to explicitly recognize two aspects of these location-aware multimedia data. First, geotagging of posts by users enables the creation of social networks around physical geographic spaces or based on geographic interests (e.g., hiking, birdwatching), referred to as “geosocial networking.” 2 Thus, location awareness provides an impetus to grow geosocial networks, which in turn benefits the social media firms supporting such networks. As a result, there is increasing interest among marketing practitioners in drawing location inferences from social media posts (e.g., Boyd and Ellison 2007). We expect location-sensitive targeting to become important in the growth of both social media and consumption behaviors that such geosocial networks perpetuate.
Second, location-aware multimedia data enable firms to implement real-time marketing strategies based on the geographic location of the consumer. For instance, marketers can target products and services to consumers based on their current location and the geographic surroundings. In an early academic study, Dubé et al. (2017) assess the returns to “geoconquesting” by examining the targeting of coupons delivered by competing firms via mobile phones. Implications for firms go beyond marketing existing products and services to developing new products and services that cater to geosocial networks and the associated consumption behaviors. Thus, these geosocial networks touch every aspect of marketing—value creation, value communication and delivery, and value appropriation.
Computers as Consumers
Another area of substantial innovation is the emergence of devices that communicate with consumers as well as with other devices over the internet (Internet of Things). Examples include smart home devices (e.g., Amazon Echo, Nest Thermostat, Nanit Pro, Furbo Dog Camera, iRobot Roomba, Peloton Bike, and Sleep Number 360 Smart Bed), connected cars (Lehner 2019), and wearables (e.g., Fitbit, Oculus Rift). These devices generate vast amounts of data on consumer behaviors, presenting unique research opportunities for developing and adapting theories (e.g., Novak and Hoffman 2019) and methods (e.g., Fortino et al. 2021).
The data that these connected devices generate are often multimodal. For example, smart home assistants, such as Alexa and Apple’s Home Pod, generate audio and video data along with location and time stamps. Similarly, wearables with embedded sensors generate (or would potentially generate) 3 data on multiple health indicators coupled with location-time information.
When we consider data generated by a single connected device, the data-analytic needs are similar to those of other multimodal data, but the data-analytic challenges multiply when we consider device–consumer interactions (perhaps with multiple consumers) and device–device interactions. Such data possess network characteristics. From the firm’s perspective, smart devices provide opportunities to design data collection systems that capture rich data on consumers, devices, consumer–device interactions, and device–device interactions. The design of such data collection systems determines the structure and richness of these data, which present opportunities for firms and researchers but also privacy concerns.
Privacy
Recent years have brought about heightened consciousness about consumer privacy in several domains. In part, this has resulted in new and stricter regulations such as the General Data Protection Regulation (https://gdpr-info.eu/) and the California Consumer Privacy Act (https://oag.ca.gov/privacy/ccpa).
What are the implications of the growth in multimedia data for consumer privacy? Over 20 years ago, Sweeney (2000) used data from the 1990 U.S. Census to show that 87% of people born in the United States could be identified through a combination of just their zip code, gender, and date of birth. The risk of identification increases sharply when a consumer can be observed through multiple modalities. Unbeknownst to many consumers, images uploaded to social media and social networking sites include metadata, some of which may be innocuous (e.g., camera information) and some of which may pose a risk to privacy (e.g., geolocation, time stamps). Moreover, background details in images may disclose unintended information such as context or other individuals. Audio data of spoken voices can reveal accents, ethnicities, gender, age, and emotions. User reviews can disclose locations and shopping and travel patterns. Thus, combining data across modalities rapidly increases the risk of disclosure of personally identifiable information and potentially sensitive information.
Second, as a direct consequence of the increased disclosure risk, companies and other entities that gather consumer information face a much higher burden of data protection. When organizations share data either internally or externally, they need to not only meet regulatory requirements, but also manage the enormous risk that security lapses such as data breaches pose to their brand reputations. Traditionally, companies have focused on managing these risks by controlling who has access to the data and how the data are accessed, but the rapid growth in the reported number of data leaks is testimony to the weakness of this approach. Other conventional protection methods have proved to be vulnerable as well, as demonstrated dramatically when an anonymized data set of movie ratings of 500,000 subscribers released publicly by Netflix was successfully deanonymized by researchers (Narayanan and Shmatikov 2008). More recently, techniques for statistical disclosure control have gained traction (Schneider et al. 2017, 2018; Zhou, Lu, and Ding 2020). These approaches rely on aggregation, or adding controlled statistical noise to data, to limit disclosure of private information to an intruder, with the goal of preserving the information utility contained in the original data. However, these techniques have focused primarily on structured data.
Similar techniques for preserving privacy in unstructured data, such as text, images, and videos, are still in their infancy. For instance, current approaches to transform textual data by adding noise often reduce the utility of the data for natural language processing tasks to impracticably low levels. Because textual data contained in user-generated reviews (Rocklage and Fazio 2020; Schoenmueller, Netzer, and Stahl 2020) or Google searches (Li and Ma 2020) have become a mainstay of marketing research, improvements in statistical disclosure control techniques for such data are crucial.
The increased privacy risk from multimedia data and the attendant costs of managing these risks is likely to inhibit sharing of such data by companies. In December 2020, Apple began requiring every app on the App Store to provide privacy labels that disclose what data it collects and how it uses these data (Chen 2021). Further, in June 2021, Apple asked apps to seek opt-in permission by users. We believe this strategy is not only consistent with Apple’s brand positioning, which emphasizes privacy, but also an early attempt by a major technology company to be forward-looking in addressing consumer privacy concerns. Previous literature (Miller and Tucker 2009) has found that state privacy regulations hindered the adoption of electronic medical records, because hospitals could not easily avail of the benefits of exchanging patient information electronically. An unfortunate consequence of the increased barriers to data sharing could be that academic researchers face reduced access to high-quality data from collaborating organizations. Another possible implication is a shift to reliance on more self-collected data through practices such as web scraping, which, as discussed, is an area in which academic journals may need to develop policies. We hope that these challenges will spur more research into methodologies for protection of multimedia data that enable companies to better manage the trade-off between privacy risk and loss of information utility.
Conclusion
The rapidly growing volume of data from multiple media sources and in multiple modalities offers several exciting opportunities for marketing. The goal of this special issue was to bring together cutting-edge research that relies on such data to address substantive and methodological issues related to the broad discipline of marketing. The articles that appear in the special issue and the emerging research we cite here reveal that the use of data from several media sources and/or multiple modalities helps researchers triangulate analyses and conduct more nuanced theory tests, increasing confidence in the results, and developing richer insight into consumer and firm behaviors. At the same time, these data pose continuing challenges for marketing executives and researchers that stem from the immense volume and complexity of the data (e.g., combining structured data with several forms of unstructured data), opening up a wealth of opportunities for methodological innovation. We believe that combining data across media sources and multiple modalities will enable researchers to take important strides in addressing ongoing and new substantive problems in the field of marketing, and we hope to see continued research leveraging these data in the pages of the Journal of Marketing Research.
Footnotes
Acknowledgments
The special issue titled “Marketing Insights from Multimedia Data” was launched by the authors during their term as Editor in Chief (RG) and Coeditors (SG and RH) of the Journal of Marketing Research. The authors are grateful to Peter Danaher, Robert Meyer, and Nader Tavassoli for providing thoughtful comments on a previous version of this editorial.
