Abstract
Literature suggests that when the image and text of news are congruent, it enhances news salience. However, when incongruent, people prioritize the viewpoint of the picture over the words. But what is congruency? This study experimentally investigates perceived image–text congruency in climate change online news, examining how newsreaders perceive and are influenced by congruency. It also explores the impacts of news format on perceptions of congruity, concerns about the climate crisis, and trust in the news outlet. Findings suggest congruency is partially subjective, highlighting that responsible visual discourse in the news may depend less on congruency than previously thought.
Keywords
Prior research suggests that when the image and text of a news story are congruent, this congruency can increase news learning and salience (Brosius et al., 1996; Paivio, 1991; Reese, 1984). In contrast, “when words and pictures are not congruent, people remember the information or perceive the viewpoint of the picture over the words” (Coleman, 2010, p. 242). Similarly, Gibson and Zillmann (2000) found that attention-grabbing images may dominate textual information in non-matching visual-verbal arrays. But what is congruency anyway? The literature has often treated congruency in news as an objective concept based either on an informational or emotional match between image and text, without considering newsreaders’ perceptions (e.g., Feldman & Hart, 2018; Huang & Fahmy, 2013; O’Neill et al., 2023). But what if newsreaders disagree on what is a “match” between an image and text in a news story?
To answer this question, we investigate perceived image–text congruency in a climate change news context. This is an especially relevant context to study audience perceptions of image–text congruency. Not only is climate change one of the most pressing issues of our time, but scholars also have systematically identified a mismatch between visual and textual information in climate change news stories (DiFrancesco & Young, 2011; O’Neill et al., 2023). For example, O’Neill et al. (2023) found that images in news stories about climate change-related heat waves often do not convey the seriousness of the crisis and instead depict summer “fun-in-the-sun.” This incongruency could have significant consequences for how news consumers perceive both the threat of climate change and the news itself, yet these effects have not been tested. We also do not know whether news readers will perceive conceptually mismatched visual and textual information in climate change news as incongruent in the first place. The current study addresses this gap.
Moreover, communication research has increasingly focused on journalism’s role in engaging the public with climate change and promoting their understanding of the issue (e.g., Appelgren & Jönsson, 2021; Feldman & Hart, 2018; Hart & Feldman, 2016; Schäfer, 2012; Stecula & Merkley, 2019; Wang et al., 2018; Wong-Parodi & Feygina, 2021). This investigation thus aligns with a broader discourse within communication research and contributes to a deeper understanding of how image editing in particular and journalistic practices in general can effectively communicate complex societal challenges like climate change.
This study experimentally examines how newsreaders perceive and are influenced by the (in)congruency between images and text in online news articles about climate change. Specifically, the study tests the main and interactive effects of variations in news text and images on audiences’ interpretation of congruency as well as the impacts of (in)congruent image–text pairings and of the perceived congruency of those pairings on individuals’ concerns for climate change impacts and trust in the news outlet that publishes climate change news. The study also investigates whether presenting news as a full article or as a social media-style headline affects perceptions of congruency, concerns about climate change impacts, and trust in the news outlet given that the social media-style format may draw more attention to the image. Thus, the key objectives of this study are to determine whether news readers perceive congruency between conceptually matched images and text, under what conditions, and with what effects.
To our knowledge, this is the first study to explore how image–text congruency is perceived by climate change news readers. This is especially crucial for clarifying the specific message editors intend to convey to their readers. Given the polysemic nature of media (e.g., Fiske, 1986) and of visuals in particular (Krause & Bucy, 2018), image–text congruency is likely to be a subjective phenomenon. Thus, if we are to fully theorize visual-verbal (in)congruency in news stories and its effects on news audiences, it is critical to understand the contextual factors that may affect newsreaders’ perceptions of congruency.
Literature Review
Effects of Multimodal Content
Prior research has investigated the impacts of multimodal content on learning, issue perception, and information recall. Cue summation theory (Severin, 1967), for example, suggests that multimodal content enhances learning by providing additional relevant cues via each channel compared with single-channel communications. This effect is also supported by the dual-coding theory (Paivio, 1971), which proposes two distinct cognitive subsystems specialized for two different types of stimuli. It implies that text and images are processed separately by independent subsystems, and their combination facilitates information recall more effectively than either mode alone (Paivio, 1986). On the same note, Graber (1990) argues that visuals are a crucial component of television news and can significantly enhance viewers’ learning, retention, and recall of information. This is echoed in other studies suggesting that successfully combining images and texts significantly improves learning (e.g., Eitel et al., 2013; Schnotz, 2002). Research from the advertising field has also focused on the interplay between visual and verbal message content, finding, for example, that in advertisements containing visual metaphors, verbal information can play an important role in recipients’ understanding and appreciation of the ad (e.g., Phillips, 2000; Ryoo et al., 2021).
Several theories have been developed to explain how message characteristics shape individuals’ information processing and subsequent effects on learning and persuasion. For example, the limited capacity model of motivated mediated message processing (LC4MP; Lang, 2006), which has been applied to online news media (Wise et al., 2009), posits that individuals possess limited cognitive resources to devote to continuous encoding, storage, and retrieval processes. Whether controlled or automatic, resource allocation is influenced by personal goals, individual differences, and message content and structure. Message features, such as the (in)congruency of multimodal content, can affect the availability of processing resources, potentially influencing cognitive load and resource allocation by activating people’s motivations or creating a distraction.
Another helpful framework for understanding information processing is the Elaboration Likelihood Model (ELM; Petty & Cacioppo, 1981). While the ELM focuses more on persuasion, it introduces the concept of cognitive elaboration—or close processing of message content—as a crucial determinant of a message’s effects on lasting attitude change. According to the ELM, variables such as individuals’ motivation and ability, along with message features, determine whether an individual will engage in elaborative processing (Petty et al., 2009).
Along these lines, the picture superiority effect (Paivio, 1991) holds that images are recalled more easily than words because images contain more informational content (Graber, 1996). Also, images, more than texts, are attention-grabbing (Garcia & Stark, 1991). They evoke heightened emotions compared with texts (Iyer & Oldmeadow, 2006), which, subsequently, may result in augmented persuasion in the audience (Powell et al., 2015).
Moreover, when the image and text are incongruent (i.e., provide different meanings), attention-grabbing images may overshadow contrasting text (Gibson & Zillmann, 2000) and neutralize or even oppose the message embedded in the text. Thus, it can be argued that the congruency of image–text news compounds is another variable that determines the resources individuals will allocate to processing a given message or the extent to which they engage in cognitive elaboration.
Image–Text Congruency in Climate Change News
Although news media play a critical role in shaping how we understand and act on climate change, climate change represents a challenge for news producers as climate change’s most severe consequences will occur in the future, which makes it difficult to predict (Van der Linden et al., 2015) and, therefore, to visualize. Given the growing risks posed by climate change, the issue is increasingly being covered by news outlets compared with the past, with the United States, for example, reaching “an all-time high” in its coverage of climate change in October and November 2021 (Simpkins, 2021). Newsrooms have devised various strategies to raise the effectiveness of climate change news. Given the impacts of multimodal content discussed earlier, one strategy that news organizations have used to increase audiences’ interest and promote their understanding of the issue has been adopting multimodal news content, particularly by adding visual elements such as photos to the text of news articles.
Indeed, research on climate change news confirms that visual representations of climate change in news media play an essential role in how people construct meaning about climate change (e.g., O’Neill et al., 2013). Some of these studies have employed content analysis and analyzed the mismatch between textual and visual elements of climate change news. DiFrancesco and Young (2011), for instance, in analyzing Canadian print media coverage of climate change, find what they describe as “a profound disjuncture between images and text” (p. 517). They argue that “image and article frequently refer to completely different dimensions of the climate change issue, thus presenting multiple and sometimes even competing narratives to readers” (p. 532). Other research has pointed out that images used in climate change stories often do not effectively convey the seriousness of the crisis. For instance, in examining the visual news coverage of the 2019 heat waves in Europe, O’Neill et al. (2023) found that many of the visuals were incongruent with article texts, focusing on “fun in the sun” (i.e., positively valenced leisure activities in or near the water) rather than the seriousness of the heat waves. This result is consistent with prior research that sheds light on the occurrence of “positive” pictures (e.g., children enjoying unusual weather conditions) in “negative” stories in the context of climate change news (N. W. Smith & Joffe, 2009). Similarly, Wozniak et al. (2015), in talking about an article on COP17 in Durban, mention that “although suggesting a close connection to the written text through layout choices, the photograph tells the reader a separate story” (p. 471).
Conceptualizing Congruency
Congruency is a taken-for-granted concept in existing theoretical and empirical studies, often defined for the purposes of content analyses or experimental studies without considering readers’ and participants’ interpretations of congruency. For example, Gibson and Zillmann (2000) define incongruency as instances when the information is contained only in photographs, not text. Alternatively, Grimes and Drechsel (1996) refer to incongruency as the juxtaposition of non-matching voice and image where images unintentionally link those they portray with defamatory references contained in accompanying text. In the context of climate change news, O’Neill et al. (2023) refer to incongruency as “a dissonance between images and text for media heatwave coverage” where most images “were positively valenced, yet hardly any texts were” (p. 98). In another example, DiFrancesco and Young (2011) define “narrative disjuncture” in newspaper content as instances “where image and language are telling different stories about climate change” (p. 528). In the same study, congruency is defined as the condition where both the text and the imagery of news articles “are packaged together with the intention (presumably) of presenting cogent narratives to audiences” (p. 531).
Outside the news and journalism research domain, the literature on consumer behavior offers additional perspectives on defining congruency. However, first, it is necessary to distinguish congruency from the related but slightly different concept of (in)congruity. In the marketing literature, incongruity is usually defined as the deliberate combination of deviant elements in an advertising message to attract audience attention and promote attitude change (e.g., Lee & Schumann, 2004). The humor literature also uses the concept of incongruity to refer to the contrasting or unexpected elements in a joke that evoke humor (e.g., Warren & McGraw, 2016). Notwithstanding, the marketing literature offers useful insights for defining congruency in multimodal news content (i.e., the consistency of images and text), which is our concern in the current study.
Heckler and Childers (1992) hold that, similar to the literature on news and journalism, the conceptual bases of congruency are not clearly identified in consumer behavior research. To address this gap, they rely on social cognition research and the associative model of person perception, which introduces three forms of behavior: “congruent, incongruent, and irrelevant or neutral to the prior expectancy formed through a personality impression” (p. 477).
Borrowing the concept of theme—“the general focus of a story to which the plot adheres” (Heckler & Childers, 1992, p. 477)—and theme-based incongruencies, the authors develop a theoretical framework that contains two essential dimensions of incongruency: relevancy and expectancy. Relevancy refers to “material pertaining directly to the meaning of the theme and reflects how the information contained in the stimulus contributes to or detracts from the clear identification of the theme or primary message being communicated.” Expectancy is defined as “the degree to which an item or piece of information falls into some predetermined pattern or structure evoked by the theme” (p. 477). Accordingly, congruency is identified as both relevant and expected, incongruency is both relevant and unexpected, and irrelevancy is uninformative. Notably, Heckler and Childers’s (1992) study is on TV advertising, a topic that does not necessarily share the characteristics of a news item.
Although relevancy still carries a significant vagueness considering the polysemic nature of images, the notion of expectancy offers a useful clue for understanding congruency. Expectancy shares similarities with schemata as described in the schema theory. Schema theory posits that “individuals’ perceptions are guided, in part, by cognitive structures—called schemata—that help individuals construct meaning out of the otherwise overwhelming number of external stimuli to which they are exposed” (Grimes & Drechsel, 1996, p. 170). Notably, schemata are primarily based on one’s background and previous experiences and expectations. Hence, when media content activates topic-specific schemata, it leads to a particular perception of the depicted event, aligning with an individual’s expectations.
Returning to news content, it is possible that congruency is different in the minds of news audiences given their varying schemata and expectations. Due to the “interpretive agency of media audiences” (Boxman-Shabtai, 2023, p. 1089), media possess a polysemic quality, and individuals exhibit diverse news consumption patterns (e.g., Fiske, 1986; Liebes & Katz, 1990). For example, in a study of how individuals interpret images about hydraulic fracturing, or fracking, Krause and Bucy (2018) found that individuals do not always interpret images in line with the images’ intended meaning. Consequently, the congruency between images and text may be more subjective than objective.
This study investigates this possibility by examining whether audiences’ perceptions of congruency follow the interpretation of congruency used in previous studies. Drawing from prior research, our operationalization of image–text congruency includes two dimensions: informational consistency and emotional consistency. Informational consistency refers to when both text and image focus on the same aspects of an issue, such as climate change impacts or solutions, as observed in Feldman and Hart’s (2018) and Hart and Feldman’s (2016) research. Emotional consistency, on the contrary, denotes a match between the emotional valence of the image and text, whether negative, neutral, or positive, as in O’Neill et al.’s (2023) research whereby a positively valenced “fun in the sun” image was considered incongruent with a negatively valenced news story about extreme heat. Despite the absence of a singular, universally accepted definition and a consensus in this field, combining these two aspects of congruency offers a relatively stable foundation to establish the design for this study.
To examine the potential subjectivity of image–text congruency among news audiences, we test the following hypothesis:
Image–text congruency matters in journalism because of its potential effects on public knowledge and perceptions of the issues covered in the news. As discussed previously, news stories that contain congruent image–text pairs may help promote learning and attitudes consistent with the news story because the meanings of the image and text reinforce one another (Brosius et al., 1996). Thus, in the context of the present study, which examines the effects of image–text congruency in news stories about climate change and extreme summer weather events (e.g., heatwaves, wildfires, drought, etc.), a congruent image that exemplifies the risks of climate change may encourage news audiences to take climate change and its consequences more seriously. In contrast, if the image is incongruent with the text, the image may predominate in news audiences’ information processing, as predicted by the picture superiority effect (Geise & Baden, 2015), or the inconsistency between the image and text may confuse audiences and thus hinder understanding of the news story, as speculated by McIntyre et al. (2018, p. 985), thereby lowering concern about climate change. This was implied by O’Neill et al. (2023), who argued that the prevalence of “fun in the sun” images in news stories about the dangers of climate change-related heat waves could “displace and marginalize vulnerability” (p. 100).
Yet research on the effects of congruency has yielded mixed findings. One reason that congruency may increase news learning and issue salience may be due to its effects on elaboration, or the extent to which people engage in message-relevant thinking, which may increase when the image and text provide reinforcing meanings (S. M. Smith & Shaffer, 2000). However, some research has found that incongruency actually increases elaboration and improves information processing, perhaps because the unexpectedness or ambiguity of the image–text combination draws attention (Lagerwerf et al., 2016; Russell, 2002). Other studies have found no effects of congruency (Powell et al., 2015; Tran, 2015).
Moreover, the influence of visuals in climate change news is still a nascent research area. Prior studies have found that news images that depict the negative consequences of climate change can increase individuals’ risk perceptions and concern about the impacts of climate change (e.g., Bolsen et al., 2018; Chapman et al., 2016; Hart et al., 2023), although these studies did not examine congruency. Some previous research has analyzed how the match between text and imagery in climate news stories affects issue perceptions and emotional reactions, finding that the effects of text and imagery were independent of one another (i.e., that congruency did not make a difference; (Feldman & Hart, 2018; Hart & Feldman, 2016).
Despite mixed evidence in the literature, our baseline expectation is that congruency between imagery and text in news will enhance learning and salience of the issue discussed in the news. Thus, we hypothesize:
Trust in institutions is a fundamental pillar of democracy and social order, particularly in modern societies (Kalogeropoulos et al., 2019; Kohring & Matthes, 2007). Specifically, media trust becomes a critical measure for news production, circulation, and understanding of media effects and audiences’ perception of media content in communication and media research (Moran & Nechushtai, 2023). News outlets earn their audience’s trust when they rely on facts and provide evidence for their claims (Elizabeth et al., 2017). Thus, imagery that aligns with the text and provides factual illustration for the argument in the news text may increase trust, although this has not been explicitly studied in prior research. Therefore, we hypothesize:
Variations in News Consumption Behavior
With the advent of digital platforms, news consumption, distribution, and production have changed fundamentally. As reported by Möller et al. (2020), much news use has turned out to be more incidental in the context of general information searches and social media experiences. For example, politically uninterested users rarely use social media intentionally for news. Consequently, these users are unlikely to read beyond quoted headlines (Möller et al., 2020). Meanwhile, as emphasized in recent studies of news consumption behaviors (Molyneux, 2018), younger news readers “nibble away at the news, whenever and wherever they feel like it. They prefer frequent news snacks to regular full meals” (Sauvageau, 2012, p. 32).
Furthermore, eye-tracking studies have indicated that visual elements are central to how users interact with multimodal news content. Bode et al. (2017), for example, in exposing users to a simulated Facebook news feed, found that images received more attention than text. Also, interviewees in Vergara et al.’s (2021) study recognized the importance of visual stimuli in their use of Facebook. In contrast to news presented on social media, images in online newspapers have usually been found not to elicit significant visual attention compared with the text for various reasons, from the designs of online newspapers to the goal orientation of online news readers (Leckner, 2012).
These studies combined with the news consumption behaviors outlined earlier could potentially make visuals accompanying headlines of news stories in a social media context more influential and alter users’ attention to and understanding of news messages compared with users who read the entire text of an article on a news website. Given these potential differences, it is important to examine how users’ attention to imagery may vary based on the format of news to which they are exposed. This leads us to our fourth hypothesis in this study:
Furthermore, it is possible that because, as hypothesized, people pay more attention to the image in the social media format, they will be more sensitive to image–text alignment. Thus, the last hypothesis is:
Method
An online experiment was fielded in December 2022. The study utilized a 2 (format: full article vs. headline only) × 3 (image: congruent image, incongruent image, no image) between-subjects factorial design. The experimental news stimulus was embedded within an online survey designed and hosted on the Qualtrics platform.
Sample
A sample of 660 adults was recruited online through the Lucid Marketplace. We excluded participants from the final sample who failed attention checks. Specifically, we eliminated individuals who either failed to accurately report the topic of the stimulus or neglected to correctly identify a specific answer in a question (i.e., “select ‘strongly agree’ for this question”). Sampling quotas were implemented to ensure the final sample matched U.S. census demographic distributions. The sample was 47.3% female, 77.3% white, and 10.7% Hispanic, with a mean age of 46.9 (SD = 17.04). Median education was “some college, no degree,” and median income was $25,000 to $49,999.
Procedure and Stimuli
After consenting to participate and answering demographic questions, participants were randomly assigned to see one of six versions of a news article about climate change (i.e., two format conditions crossed with three image conditions; n = 111 per cell). The text of the stimuli (available in Appendix) referred to the same real-world news article adapted from The Washington Post about how climate change is causing extreme summer weather events in the United States, such as heatwaves, wildfires, and drought. The headline in all conditions was “Summers in America are becoming hotter, longer, and more dangerous due to climate change.”
The news format condition was manipulated by altering the format of the news stimulus (i.e., headline-only or full article). Participants in the headline-only + image condition saw a mock Twitter post with the headline and image; participants in the headline-only with no-image condition saw a mock iPhone lock screen with a popped-up news headline. We chose to examine the headline-only condition in this manner to uphold ecological validity, considering that users rarely come across a tweeted news story without an accompanying image. Finally, participants in the full article condition saw a news story from a mock news organization webpage, either with or without an image.
The image condition was manipulated by altering the image that participants saw in the news piece. We operationalized congruency based on the emotional and informational alignment of the image with the text. For the congruent condition, participants saw one of two stimulus-sampled versions of a color image depicting a wildfire in the United States (Figures 1 and 2), which is also mentioned in the text of the news article, reflecting informational alignment between the image and text. For the incongruent condition, participants saw one of two stimulus-sampled versions of a color image of unconcerned beachgoers (Figures 3 and 4), which is not something to which the text refers. The choice of incongruent images follows from O’Neill et al. (2023) who found that climate change news stories often misrepresent heat wave risks by including images that depict people enjoying the beach. Beachgoing is generally considered a pleasurable activity, whereas heat and wildfire risk has negative connotations; thus, there also was incongruency in the emotional valence of the text versus imagery. Moreover, the explicit reference to “danger” in the article headline and text is informationally and emotionally consistent with the image of wildfires, whereas it is inconsistent with the image of beachgoing.

Congruent Image 1.

Congruent Image 2.

Incongruent Image 1.

Incongruent Image 2.
The content of the stimuli, attributed to Reuters, was factually accurate and adapted from actual news published by the legacy press (the text was taken from The Washington Post, and the images were obtained from climate crisis news that appeared in Reuters, Associated Press, The Wall Street Journal and The New York Times) to be reflective of social media and website posts that users would encounter online in their everyday lives.
We chose Reuters as it is considered a relatively neutral news source (Meylan, 2022). The original content of the article was not altered other than to shorten the length (the full-text condition was 326 words) and to vary the news image. Participants were instructed to read the stimulus carefully and told that questions about the news piece would follow. After seeing the stimulus, participants completed a survey that measured their perceived congruency of the text and image, concern about the summer heat, and trust toward Reuters, among other variables.
Measures
Perceived Congruency
Perceived congruency, or the extent to which participants saw a fit between the text and the image of the news piece, was measured by asking participants how much they agreed or disagreed with the following statements on a scale from 1, strongly disagree to 7, strongly agree: (1) “The image matched, or conveyed the same message as, the headline/text of the news article,” (2) “The image is an appropriate choice to reflect the content of the article,” (3) “The image helped me better understand the content of the article/headline.” Answers to the three statements were then averaged together (α = .90) into a single scale that ranged from 1 to 7 (M = 4.69, SD = 1.53). Only participants assigned to one of the image conditions (n = 456) answered these questions.
Image Preference Among Participants in the No-Image Condition
To investigate perceived congruency among the participants in the no-image condition (n = 204), the following question was asked: “If you were a news editor and were asked to pair this headline/text with an image, which of the following images would you select?” Participants were then asked to choose one of two images: one incongruent image (the beach image in Figure 3) or one congruent image (the wildfire image in Figure 2).
Concern for Extreme Summer Heat
Concern for extreme summer heat was measured by asking participants two questions. The first question asked “Overall, how much of a threat do you think extreme summer heat poses?” on a scale from 1, No threat at all to 7, A very large threat. The second question asked participants “How concerned are you about extreme summer heat?” on a scale from 1, Not at all concerned to 7, Extremely concerned. Responses to the two questions were then averaged together (α = .93) into a single scale that ranged from 1 to 7 (M = 5.01, SD = 1.73).
Trust in Reuters
Trust in Reuters, the news organization to which the news stimulus was attributed, was measured by asking participants how much they agreed or disagreed with the following statements on a scale from 1, strongly disagree to 7, strongly agree: (a) “I trust Reuters as a news source,” (b) “I believe that Reuters is a credible news source,” and (c) “I am likely to seek out news from Reuters in the future.” Answers to the three statements were then averaged together (α = .93) into a single scale that ranged from 1 to 7 (M = 4.73, SD = 1.60).
Attention to the Image
Attention to the image was measured by asking participants how much they paid attention to the image accompanying the tweeted/posted news article on a scale from 1, None to 7, A great deal (M = 4.82, SD = 1.78). Only participants who were assigned to one of the image conditions (n = 456) answered this question.
Results
Before conducting the main analyses, t-tests were conducted to determine whether the two stimulus-sampled congruent images and two stimulus-sampled incongruent images, respectively, had similar effects on the dependent variables. The two incongruent images did not vary significantly on perceived congruency, t(227) = −1.333, p = .18; Cohen’s d = .18, heat concern, t(226) = .998, p = .31; Cohen’s d = .13, and Reuters trust, t(225) = −1.424, p = .15; Cohen’s d = .19, and thus were collapsed into a single incongruent image condition for the main analyses. However, the two congruent images varied significantly in perceived congruency, t(225) = 1.991, p < .048; Cohen’s d = .26, such that Congruent Image 1 (Figure 1) was perceived as more congruent than Congruent Image 2 (Figure 2). We thus treated these as separate congruent image conditions (“Congruent Image 1” and “Congruent Image 2”) in all analyses.
To test our hypotheses, a two-way, full factorial analysis of variance (ANOVA) was used to examine whether there were any main or interactive effects of news format and image condition on perceived congruency, heat concern, and Reuters trust (see Table 1 for means and standard deviations for all dependent variables across experimental conditions). As a test of H1, the main effect of the image condition on perceived congruency was not significant, F (2, 450) = 2.902, p = .056; partial η2 = .01. Looking at the mean differences between each of the congruent images and the incongruent image condition (Table 1), although the pattern of means showed that Congruent Image 1 (M = 5.00, SD =.14) was perceived as more congruent than the incongruent images (M = 4.60, SD = .11), this difference was not significant after using the Bonferroni correction to adjust for two pairwise comparisons (p = .056). 1 The perceived congruency means for Congruent Image 2 (M = 4.58, SD = .14) and the incongruent images were nearly identical (p = 1.00). Thus, H1 was not supported.
Means and Standard Errors of Dependent Variables Across the Experimental Conditions.
As an additional test of H1, we examined the image preference among participants in the no-image condition. Here, we found that 67.2% preferred the congruent wildfire image compared with 32.8% who preferred the incongruent beach image; a one proportion z-test confirmed that these are significantly different from one another (z = 4.90, p < .001; Cohen’s h = .70).
Turning to H2, the main effect of the image condition on heat concern was not significant, F (3, 651) = 1.997, p = .113; partial η2 = .009. However, post hoc tests comparing each of the congruent image conditions to the incongruent condition, using the Bonferroni correction to adjust for two pairwise tests, showed that the incongruent images (M = 5.18, SD = .11) resulted in significantly higher heat concern than Congruent Image 2 (M = 4.70, SD = .16, p = .036), which is opposite of what H2 predicted. This difference should be interpreted cautiously given that the omnibus effect of image condition was not significant. There was no difference between Congruent Image 1 (M = 4.95, SD = .16) and the incongruent condition (p = .55).
Similarly, a test of H3 found that the effect of image condition on Reuters trust, F (3, 650) = 1.809, p = .144; partial η2 = .008, was insignificant, although again the post hoc comparisons between the incongruent and congruent images using the Bonferroni adjustment demonstrated significantly higher trust in response to the incongruent images (M = 4.88, SD = .11) compared with Congruent Image 2 (M = 4.46, SD = .14, p = .036). This is the opposite of H3’s prediction, although the results should be interpreted cautiously given the lack of a significant omnibus effect. There was no difference between Congruent Image 1 (M = 4.69, SD = .15) and the incongruent condition (p = .58).
In support of H4, we found a significant relationship between news format and attention to the image, with the mean attention being higher in the headline condition (M = 5.20, SD = 1.670) than in the full-text condition, M = 4.39, SD = 1.818; t(454) = −4.916, p < .001; Cohen’s d = .46, lending support to H4. We also note that there was no interaction between news format and image congruency, F (2, 450) = .400, p = .671; partial η2 = .002; thus, the effect of news format on image attention occurred regardless of manipulated congruency.
H5, which predicted that news format would moderate the effects of image congruency, was not supported, as evidenced by the lack of a significant interaction between the news format and the image conditions for perceived congruency, F (2, 450) = .464, p = .629; partial η2 = .002, heat concern, F (3, 651) = .410, p = .746; partial η2 = .002, and Reuters trust, F (3, 650) = .516, p = .671; partial η2 = .002. We also note that news format did not have a significant main effect on perceived congruency, F (1, 450) = .0004, p = .983; partial η2 < .0001, heat concern, F (1, 651) = .339, p = .561; partial η2 = .001, and Reuters trust, F (1, 650) = .038, p = .845; partial η2 < .0001. However, when examining image preference among participants in the no-image condition, a significant chi-square test, χ2 (1, N = 204) = 5.967, p < .015; φ = .17, indicated that participants’ preference for the congruent wildfire image over the incongruent beachgoing image was significantly smaller in the headline-only condition (59.4% vs. 40.6%) than in the full-text condition (75.5% vs. 24.5%). This provides some evidence that perceptions of congruency depend on news format.
Post hoc Analysis
Given the differences in perceived congruency between the two congruent images, and the limited effects of the congruency manipulation on heat concern and trust, we explored the possibility that perceived congruency, rather than manipulated congruency, is what influences these outcomes. Using ordinary least square regression, controlling for demographics, political party identification (measured on a five-point scale ranging from 1 strong Democrat to 5 strong Republican; M = 3.09, SD = 1.40), climate change news consumption (based on the percentage of respondents who selected climate change from a list of news topics that they follow regularly, 42.6%), and the experimental conditions (see Table 2), we found that the relationship between perceived congruency and concern about extreme summer heat was significant and positive (B = 0.40, p < .001), as was the relationship between perceived congruency and trust in Reuters (B = 0.52, p < .001). Notably, once the positive effect of perceived congruency is accounted for, the effects of both Congruent Image 1 and Congruent Image 2 (relative to the incongruent condition) on concern and trust, respectively, are negative and significant (trust: Congruent Image 1 B = −.40, p = .007; Congruent Image 2 B = −.38, p = .01; heat concern: Congruent Image 1 B = −.39, p = .01; Congruent Image 2 B = −.39, p = .01). The negative effects are consistent with the mean differences reported in Table 1; however, the negative effects of both congruent images are now significant in the regression analyses, indicating a suppressor effect of perceived congruency (Conger, 1974). Still, when examining the standardized coefficients (Betas) in Table 2, the effects of perceived congruency are larger in magnitude than the effects of manipulated congruency.
OLS Regression Analysis Predicting Reuters Trust and Heat Concern.
The incongruent condition is the reference category. bThe headline-only news format is the reference category. cFemale is the reference category. dNon-White is the reference category.
p <.05, **p <.01, ***p <.001.
To further understand this suppressor effect, we tested a mediation model using SPSS PROCESS 4.1 (Model 4), with Congruent Image 1 and Congruent Image 2 as multicategorical predictors, perceived congruency as the mediator, heat concern and Reuters trust as the respective dependent variables, and with the same covariates as in Table 2. This model allows us to estimate the total effect of the congruent conditions on each dependent variable (i.e., without perceived congruency in the model; similar to the ANOVA results), the direct effect of the congruent conditions on each dependent variable after accounting for perceived congruency (i.e., identical to the results in Table 2), and the indirect effects of the congruent conditions on each dependent variable via perceived congruency. These estimates are presented in Table 3; mediation models for heat concern and trust are depicted visually in Figures 5 and 6, respectively. Briefly, for both dependent variables, Congruent Image 1 has nonsignificant negative total effects, significant negative direct effects, and significant positive indirect effects via perceived congruency. (The positive indirect effects result from Congruent Image 1’s higher levels of perceived congruency relative to the incongruent condition, and the positive associations between perceived congruency and both dependent variables). This pattern indicates that perceived congruency suppresses the negative effect of Congruent Image 1 on heat concern and trust (see MacKinnon et al., 2000). This is due to the opposite signs of the direct effects (negative) and indirect effects (positive) of Congruent Image 1, which cancel each other out, resulting in a nonsignificant total effect. For Congruent Image 2, the total and direct effects on both dependent variables are significant, negative, and of similar magnitude, whereas the indirect effect via perceived congruency is not significant. Thus, the effects of Congruent Image 2 on concern and trust are negative regardless of perceived congruency. Overall, these results suggest that manipulated congruency has negative direct effects on heat concern and trust; yet, when manipulated congruency increases perceived congruency, as it did in the case of Congruent Image 1, this results in a positive indirect effect on concern and trust via perceived congruency.

Mediation Model Depicting the Effect of Manipulated Image Congruency on Heat Concern via Perceived Congruency.

Mediation Model Depicting the Effect of Manipulated Image Congruency on Reuters Trust via Perceived Congruency.
Total, Direct, and Indirect Effects of Manipulated Congruency (Compared to Incongruency) on Heat Concern and Reuters Trust, With Perceived Congruency as Mediator.
Note. Estimates are based on mediation models conducted using SPSS PROCESS 4.1 (Model 4); see Figures 5 and 6. Indirect effects were computed with 10,000 bootstrapped samples. Age, gender, race and ethnicity, political party identification, climate change news consumption, and news format (i.e., full text vs. headline) were included as covariates. Standard errors are in parentheses. Significant effects are bolded. BootLLCI/BootULCI = Bootstrapped lower and upper confidence intervals.
Discussion
The purpose of this study was to test the effects of image–text congruency in news stories about climate change on individuals’ perceptions of congruency, their concern about climate change impacts, and their trust in news, while also examining how the effects of congruency may vary depending on the format of news (i.e., full online article vs. social media post). This is the first study we are aware of to experimentally test perceived congruency among online climate change news audiences. In this study, congruency was operationalized using imagery of wildfires, which echoed the textual news stimulus discussing the effects of climate change on extreme summer weather, including wildfires, intense heat, droughts, and flooding. The imagery of beachgoers was used to represent incongruency, as beachgoing was not discussed in the news story and is not illustrative of the climate change dangers emphasized in the headline and body of the news article, yet commonly appears in news stories about the impacts of climate change on extreme heat (O’Neill et al., 2023). Overall, the findings reveal small and inconsistent effects of image–text congruency and suggest that perceptions of congruency may lie in the eye of the beholder rather than strictly follow scholars’ conceptualization of congruency. Moreover, while manipulated congruency had limited effects, perceived congruency was robustly related to both concern about extreme summer heat and news trust. Moreover, perceived congruency appeared to suppress a small, negative effect of manipulated congruency on both concern and trust. We elaborate on and discuss the implications of these findings below.
Regarding the effects of image–text congruency on perceived congruency (H1), the results were mixed. On one hand, Congruent Image 1 (Figure 1) was perceived as more congruent than the incongruent images, although it did not reach the conventional threshold of statistical significance. Moreover, when participants in the no-image condition were asked to select an image that echoes the text, they exhibited a much higher preference for the congruent wildfire image over the incongruent beach image, helping to validate scholars’ conceptualization of congruency. On the other hand, the two congruent images varied in their perceived congruency despite both illustrating wildfires, with Congruent Image 2 perceived as less congruent than Congruent Image 1 and no different from the incongruent images, which provides evidence for congruency as a subjective phenomenon. One plausible explanation for the latter finding is rooted in the ideational resources of visuals, that is, “the capacity of semiotic systems to represent objects and their relations in a world outside the representational system” (Kress & Van Leeuwen, 2020, p. 47). Beyond formality, visual composition holds significant semantic and ideological dimensions, and changing elements within an arrangement alters meaning. In our study, the presence of an individual in Congruent Image 2 (Figure 2), seemingly unaffected by the surrounding wildfire, may disrupt the expected narrative coherence, influencing participants’ perceptions.
Regarding H2 and H3, the congruency manipulation, overall, did not significantly affect heat concern or trust in Reuters. However, opposite of our hypotheses, post hoc comparisons showed that the incongruent images resulted in significantly higher heat concern and trust than Congruent Image 2. Moreover, after controlling for perceived congruency in the post hoc regression analyses, both congruent images, compared to the incongruent condition, were significantly and negatively related to concern and trust, although the magnitude of the effect was small. While these trends are only suggestive, it is possible that people paid more attention to the text when the image was incongruent or unexpected, and thus were more likely to internalize the text’s message (i.e., that extreme heat is a problem) and see the news source as credible. This is consistent with Russell’s (2002) study that finds that incongruency in multimodal content induces more elaboration than congruency. Tentatively, these findings suggest that O’Neill et al.’s (2023) concern about “fun in the sun” images did not translate into discernible effects on audiences, as incongruent images of beachgoers did not dampen audiences’ worry about climate change or their trust in news, at least in the short term. Future research should continue to probe the effects of positively valenced climate change images among news audiences.
Our post hoc analyses suggest that in terms of influencing heat concern and trust in Reuters, perceived congruency may matter more than manipulated congruency. Regression analyses found that perceived congruency demonstrated a moderate positive association with both heat concern and trust in Reuters. Recognizing that these relationships are only correlational and not causal, the findings are nonetheless suggestive. Regardless of which image they were assigned to, the respondents perceived higher concern about extreme heat if they perceived the image to be congruent, indicating the learning of the message. Unlike Gibson and Zillmann’s (2000) findings, which demonstrated that visual information can override dissimilar textual information when incongruent text-image pairs are present, it may be that audience perceptions of congruency, rather than scholars’ understanding of congruency based on the assumed mismatch, led to higher elaboration of the news content, resulting in greater concern. Future research should test elaboration as a mechanism of the effects of both manipulated and perceived congruency.
Similarly, given the positive association between perceived congruency and news trust, participants, regardless of the image condition, trusted Reuters if they perceived the images as congruent. This relationship may be attributed to the public’s appreciation for professional standards in news production, such as fact-checking and utilizing multiple sources, which aim to bolster audience trust in journalistic content (Moran & Nechushtai, 2023). When audiences perceive a “congruent” photograph that aligns with the news article’s message, they may view this as responsible and skillful news gathering and reporting. This finding holds implications for news practitioners who rely on trust as a foundational component in securing funding mechanisms, news production, circulation, and consumption, wherein the trust of individual news consumers is key in ensuring the stability and existence of news outlets (Moran & Nechushtai, 2023).
One explanation for different interpretations of congruency between scholars and news readers could be that although scholars see a mismatch between beach images and text about the impacts of climate change, the audience sees them as all pertinent to climate change and its consequences. This understanding may not resemble a mismatch in the eyes of the audience and does not impact their concern about climate change and trust in the news outlet. At the same time, it is also possible that some audiences do not see the connection between climate change and wildfires due to a lack of understanding and/or denial of the consequences of climate change, and in turn, do not perceive the wildfire images as any more congruent than the beach images. Thus, understanding what the online news audience interprets as (in)congruent and how this varies across social groups is essential, and this is an important direction for future research.
We also found that news format can have important effects on audiences’ attention to and interpretation of news images. Supporting H4, we found that participants pay significantly more attention to the image when the news is presented in a headline-only, social media style format versus when they are exposed to the full text of an online news article. Previous studies have highlighted the importance of visual elements in increasing attention to news on social media (Vergara et al., 2021), as well as their diminished significance in online newspapers (Leckner, 2012). Our study directly compares attention to visual elements in news on social media versus online news articles. This finding has implications for journalists and news editors in selecting images for different news modes and highlights the extra importance of visuals for sharing news on social media.
However, whereas news format affected attention to imagery, news format had no significant main or interactive effects on perceived congruency, heat concern, or trust in Reuters (H5). Given previous literature that finds that visual elements are central to users’ elaboration (Bode et al., 2017; Vergara et al., 2021), thus resulting in more effortful engagement, learning, and awareness of the issue (Eveland, 2001; O’Neill & Smith, 2014), we might have expected heightened attention to the image in the headline-only condition to lead to differential effects of news format on perceived congruency, heat concern, and trust. However, it is possible that other aspects of the news format counterbalanced any effect of image attention on the study outcomes. For example, the full-text condition may offer important informational context for interpreting image-text congruency and deducing heat concern and media trust that offsets the heightened elaboration that may occur due to higher image attention in the headline-only condition, thus resulting in null effects. Consistent with this idea, we found that participants who were in the no-image condition were less likely to choose the fire image over the beach image as the better match for the news story when they had seen the headline-only format compared to the full-text format; this suggests that the full-text format may have offered important context that shaped people’s interpretation of the images. This finding further reinforces the need for precision in photo editing practices on social media to improve the communication of the intended messages, especially given the “news snacking” (Sauvageau, 2012) prevalent among digital media audiences. Additional research is needed to better understand how news format interacts with imagery and its effects on audiences.
Taken together, our findings suggest that congruency is, at least partially, in the eyes of the beholder and that O’Neill et al.’s (2023) concerns about “fun in the sun” images may be less consequential than assumed. Our study demonstrated that showing people a “congruent” wildfire image versus an “incongruent” beach image did not significantly impact individuals’ concern for climate change consequences, such as extreme summer heat, or their trust in news sources like Reuters—and, if anything, the incongruent images increased concern and trust. This suggests that the responsibility for ensuring responsible visual discourse in news coverage, at least in this context, may depend less on congruency than previously thought. However, this does not diminish the importance of advocating for a more inclusive and responsible visual discourse in climate change reporting and in news media more broadly. Instead, our findings highlight the need for a deeper understanding of how audiences interpret visual elements in news and the potential for such knowledge to guide more responsible and ethical photo editing practices. Our findings also have methodological implications for research concerned with the effects of visual–text congruency. Some prior studies that have investigated the effects of image–text congruency have taken congruency as a given (e.g., Seo & Dillard, 2019; Tukachinsky et al., 2011; Zillmann et al., 2001); our findings suggest that audiences may interpret congruency differently and that these perceptions should be taken into account in research design.
As with any study, the present investigation has limitations that are important to acknowledge. The present study is only one study investigating this issue, and caution should be taken when generalizing. We used only two images in the congruent and incongruent conditions; it is possible that characteristics unique to these images influenced the pattern of effects. Selecting more photos or a different variety of them could change the results of this study. In particular, the beach images we used in the incongruent condition—although taken from actual news articles about climate-related heat—were not as vibrant as many of the exemplar images from O’Neill et al.’s (2023) study, which may have minimized differences between the congruent and incongruent conditions. Also, we did not stimulus sample the news headline and text; thus, it is possible that the effects observed here were idiosyncratic to the particular headline and text used in this study, which explicitly highlighted the dangers of climate change. Future research should consider different types of climate change news texts as well as different images. In addition, the observed effects may have been constrained given that we only exposed participants to a single message; the effects may be stronger following exposure to multiple (in)congruent messages, which represents a useful direction for future research. We also examined heat concern and trust in Reuters as our dependent variables. Image–text congruency could affect other learning and behavioral aspects of news consumption that were not considered in this study. Finally, we did not measure news elaboration, an important potential mechanism that should be investigated in future studies.
Subsequent research can expand upon the current study by further exploring how congruency is defined and measured in textual-visual multimodal news content and by more deeply exploring the factors that shape congruency in the minds of audiences. The present study emphasizes the need for researchers to examine environmental communication through the lens of audiences, gaining insights into their thoughts and perceptions in the current digital landscape rather than relying solely on existing theories, hypotheses, and assumptions. This approach can foster a more knowledgeable and engaged citizenry.
Footnotes
Appendix
Text of the news articles used in the stimuli.
Data Availability Statement
The data that support the findings of this study are available upon reasonable request.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
