Abstract
Despite calls for journalists and media agencies to address a disconnect between news audiences and news prose, content continues to increase in its difficulty to read and comprehend for the masses. While readability is often associated with audience comprehension and engagement, studies have neglected to assess whether readability is a factor in audience assessments of the credibility of content. Using an online experimental design, this study examines whether readability acts as a heuristic that helps news consumers make credibility judgments of news. Results show that readability tends not to be a predictor of credibility perception, regardless of partisanship strength or media use preferences. Theoretical and practical implications are discussed.
As described in their book, The Elements of Journalism, the news audience in the United States has slowly become less important to the news organizations that were built to serve them (Kovach & Rosenstiel, 2014). In the 1880s, the penny-press owners would publish content specifically meant to serve immigrants—stories written simply so immigrants could read and understand them, and use them as guides for becoming American (Kovach & Rosenstiel, 2014).
However, since that time, the journalism industry has moved from serving the middle class to something less publicly oriented. Nikki Usher (2021) reiterates this disconnect between newsrooms and the public in her book, News for the Rich, White, and Blue. The biggest problem with journalism’s choice to target specific demographics and move away from serving the masses is, of course, that advertisers no longer depend on newspapers or news outlets to reach their audiences (Usher, 2021). Instead, news organizations must again rely on audiences to sustain their operations. However, when news reporters are writing more for their sources than their communities (Hess, 1981), the ability for news operations to reconnect with their public is difficult.
The stories these outlets tell and the way they tell them can serve as indicators as to the value news organizations have for the people in their areas. If news articles are written in a way that is too hard to read or too complex to understand, it serves as an indication that the publication is really meant only for the well-educated and elite, not the average member of society. The use of words and the readability of the content act as cues to people that the news is not for them, but for someone else.
The way news content is written has bearing on its outcome. This study tackles “the readability problem” in news (Chall, 1987; Dalecki et al., 2009; Danielson et al., 1992; Flesch, 1948; Kincaid et al., 1975; O’Hayre, 1966). More specifically, this study aimed to determine whether readability, which is a measure of effort required to read and comprehend writing, is connected to one’s assessment of credibility in news content.
Based on the Heuristic-Systematic Model (HSM) as a theoretical framework, a readability cue in news stories should activate a credibility heuristic in content (Chaiken, 1987; Sundar, 2008). This study tests this previously understudied heuristic—news readability—which we suggest acts as a cue, “if a news article appears to be easy to read, then it must be more credible.” However, as the experiment indicates, readability does not seem to be related to credibility as much as other factors, including partisanship and previous media use.
Literature Review
Readability
The construct of readability has varied definitions across the literature, differentiating between an emphasis on comprehension of the audience (Mc Laughlin, 1969) or writing style (Klare, 1963). Chall (1987) defined readability using both concepts. For the purposes of this analysis, which focused mainly on readability levels as a heuristic, the writing style serves as the most efficient way to identify readability levels. For established scales such as the Flesch Reading Ease (FRE) scale and the Flesch-Kincaid Grade Level (FKGL) scale, readability is measured with objective, prescriptive measures such as word count, sentence length and word length (Flesch, 1948; Kincaid et al., 1975).
Readability as a Heuristic
Despite calls to solve the problem of declining readability of news content, news stories have become less readable over time (Dalecki et al., 2009; Danielson et al., 1992; Johns & Wheat, 1978; Kleinnijenhuis, 1991; Porter & Stephens, 1989; Tauberg, 2019). Journalists are trained to avoid complicated prose that would be difficult for readers to value (Danielson et al., 1992). Sundar (1999) considered content clarity and comprehensiveness as aspects of content quality, which would affect credibility measures. In addition, the use of complicated prose and complex structures makes content difficult to understand, such that only the most knowledgeable and educated are served by this work (Hess, 1981; Usher, 2021).
The use of simple language has been added as part of the strategy for garnering more audience appeal and understanding in politics (Bischof & Senninger, 2018; Cichocka et al., 2016; Kayam, 2018) and the medical field (Basch et al., 2020; Moore & Millar, 2021). In addition, audiences have recognized the value of more comprehensible content online. For example, restaurant reviews that were easier to read were also most likely to have higher “helpfulness” ratings (Fang et al., 2016), and health forums with articles that were easier to read (with less jargon) were perceived as more credible by nonexperts than forums with a lot of technical language (Zimmermann & Jucks, 2018). However, few recent studies on news content focus solely on the readability as it relates to credibility measures (Danielson et al., 1992; Johns & Wheat, 1978). In a related work of deceptive news practices, Dalecki et al. (2009) found that stories based on fabricated content (such as work by Jayson Blair at the New York Times) were easier to read than nondeceptive, authentic journalism. However, this study did not take into account audience perceptions of credibility in either type of content analyzed.
Heuristics, which are defined as mental shortcuts used to make quick judgments, can play a role in whether news content is seen as credible. According to dual-processing models such as the HSM, people use systematic and heuristic routes of thinking to make judgments of credibility (Chaiken, 1987; Petty & Cacioppo, 1986). The systematic route, which relies on complex cognitive effort to process information, is used most often when people are highly involved in the subject matter because the information is personally relevant (Perloff, 2021). The heuristic route, which relies on cues to trigger heuristics for quick judgment, is used when people are unmotivated or lack the ability to use cognitive energy toward a decision.
Readability measures such as language use and article length are commonly used cognitive heuristics (Fogg et al., 2003; Metzger, 2007). For example, Rennekamp (2012) found that readability in financial disclosures to the Securities and Exchange Commission was linked to stronger reactions from investors, but it was not linked to managerial credibility measures. Fix and Fairbanks (2020) found links between readability in U.S. Supreme Court decisions and citation frequency in state courts, suggesting that the clarity and comprehension of the decisions make them more likely to be cited and thus affect court precedent and decisions.
When readability is used as a heuristic, it gives readers the ability to make judgments on content based on the perceived difficulty of the syntactic and lexical complexity. This is especially important for the readability of news that readers are exposed to through online media because they are much more likely to need to rely on heuristics to make credibility judgments as a way to avoid information overload from the medium (Fogg et al., 2003; Metzger, 2007).
Credibility
In news, credibility is generally defined by how believable the information and the source of that information are (Hovland et al., 1953). Studies of credibility differentiate between three types of credibility: message, source and medium (Appelman & Sundar, 2016; Kiousis, 2001; Metzger et al., 2003; Sundar, 2008). While all forms of credibility are connected to each other, this study specifically focuses on message credibility, which is the way one judges the “veracity of the content of communication” (Appelman & Sundar, 2016, p. 63).
Shortcuts such as expertise cues, length cues, authority cues and design cues can help activate heuristics such as the readability heuristic or reputation heuristic to help news readers perceive news information as credible (Lin et al., 2016; Metzger & Flanagin, 2015). In online settings, heuristics can play a larger role. Sundar’s (2008) MAIN model identifies several heuristics that can be used to judge the credibility of news content in online settings based on four types of technological affordances. The modality affordance is most comparable with the medium of the content, but it is distinguished in the sense that a video might be offered online in the same space as a text-based story. While the agency, interactivity and navigability affordances play roles in credibility judgments, this study specifically focuses on the modality affordance as the message source is hidden from participants, and the experiment limited interactivity and navigability in the study environment.
The MAIN model suggests that cues (such as news readability) should serve as a prompt that affects perceived news credibility. If this is the case, then news should be seen as more credible in high-readability conditions, presumably due to the operation of a ease-of-reading heuristic (if the message is easy to read, then it must be credible). Given this theoretical logic, the following hypothesis is proposed:
Ideology and Partisanship Strength
Some studies have shown a correlation between conservative ideology, general partisanship and distrust in media (Culver & Lee, 2019), which could affect judgments of credibility. While trust and credibility are not the same thing, they are interconnected (though there are arguments as to whether trust is an antecedent to credibility or vice-versa; Engleke et al, 2019; Kiousis, 2001). Strength in partisanship has also been tested as a predictor of perceived bias in news. Political partisans tend to bring their own biases into their perceptions of news coverage, often viewing objectively neutral content as biased when it goes against their views and less biased when it aligns with their views (Gunther, 1992; Vraga & Tully, 2015). In addition, people prefer information that is consistent with preexisting attitudes and beliefs (Knobloch-Westerwick & Hastall, 2010; Stroud, 2011).
While people seek information that is consistent with their views, they also pay more attention to information that is relevant to their personal lives (Pennycook & Rand, 2019). In their study of hyper-partisan news use and its effects on cognitive and affective involvement, Peacock and colleagues (2021) showed some relationship between use of traditional news and cognitive involvement for adults. The reason, they suggest, is because news that does not align with preconceived beliefs and attitudes will drive partisans to seek news that does reinforce their beliefs and attitudes as a form of motivated reasoning. Motivated reasoning is the general tendency to assess information in a way that helps one achieve a goal other than accuracy (Kunda, 1990). In their discussion of partisanship and its association with belief in fake news, Pennycook and Rand (2021) point out that people are more likely to believe news content that is aligned with their personal beliefs, but “the effect of political concordance is typically much smaller than that of the actual veracity of the news” (p. 389). However, when identity-protecting motivated reasoning is combined with a high disposition to rely on heuristic processing, people are more likely to depend on cues, such as readability, to guide them on their evaluation of the content and its ability to conform with their partisan identity.
In addition, a recent study found that extremely partisan media outlets were more likely to use easier-to-read language and content than mainstream and nonpartisan outlets (Sparks & Hmielowski, 2023). If partisan-media readers become accustomed to easier-to-read content, there is a possibility that they will associate high readability as a cue of credibility.
While some studies demonstrate equal cognitive processing among partisans and nonpartisans (Pennycook & Rand, 2019), others suggest partisans are less likely to exert cognitive energy on information that is identity-conforming (Kahan, 2012). To account for this, we measured for partisanship as a moderating variable:
Media Preferences
As Sundar (2008) described in his study, some heuristics transfer across media platforms. For example, old-media heuristics are triggered with certain writing styles in news. Legacy media, which are defined as media outlets that started in a medium other than the web, have been shown to be distinct from digital native media (Vara-Miguel, 2020). Given their structural differences, it is not surprising that the audience perceives credibility differently among media platforms, as well. For example, audiences perceive visible sources such as a broadcaster reading the news on a television or the editors listed on a newspaper masthead in a different way than a technological source, which are distinct to the web (Sundar et al., 2007). This study specifically focuses on text-based news without the attribution of a gatekeeper or source. Given that, there is some evidence that text-based media holds higher credibility among users than television media or other more visual outlets, as Kiousis (2001) also found evidence that people held stronger credibility opinions about text-based media channels and local media than web-driven outlets.
Previous work in the area has also shown a connection between media use and perceived media credibility (Kiousis, 2001; Tsfati & Cappella, 2003). Tsfati and Cappella (2003) theorized that people use news to gain accurate information about the world, based on the assumption that people make rational decisions for maximum use. However, media use can also be built into ritual and habit (Rubin, 2009), in that it can serve personal and social identity needs (Rubin, 2009; Tsfati & Cappella, 2005). This complicates the notion that media use is based in reliance, as the relationship can be much more nuanced (Strömbäck et al., 2020). Kiousis (2001) found some evidence that perception of online news credibility was associated with web use, and newspaper credibility was related to newspaper use. However, there is not much research that examines whether media preferences would moderate the relationship between readability and media credibility judgments. Therefore, we explore whether media preference would affect the relationship with the following research question:
Method
An online experiment was conducted to test study hypotheses with a single-factor between-subjects design, testing the effects of readability (low vs. high) on perceived news credibility. An a priori power analysis determined a total sample of 278 subjects would be adequate to achieve an 80% power assuming a “small effect” of d = .30. Participants were recruited on Amazon Mechanical Turk (MTurk), which is a crowdsourcing labor market. MTurk workers were required to be at least 18 years old and U.S. residents to participate in this study. The study was posted in August 2021. To accommodate issues with excluded responses, an additional 20% (n = 56) of cases were collected to maintain adequate statistical power. About 130 cases were excluded for incompleteness and lower-than-needed time allotted to complete the study. The institutional review board approved the study protocol and deemed the study “exempt.”
With 253 cases left in the sample, the mean age of participants was 35.63 (SD = 10.84, median = 33) years, ranging from 20 to 72 years of age. The sample skewed slightly male (59.1%). The population was primarily White/Caucasian in self-reported racial identity: 83% reported being White/Caucasian, 8.7% reported being Black/African American, 3.6% were Asian American, 2.4% were Hispanic/Latinx, 2% were American Indian/Alaskan Native and 0.4% selected the “other” category. The sample also skewed high in education, with 75.3% of subjects reporting completion of a bachelor’s degree, 15.5% reported completion of a graduate degree, 4.4% completed an associate’s degree, 2.4% had some college education and 2% completed high school with a diploma. Regarding political ideology, the sample was fairly even between those who identified on the liberal end of the spectrum (44.8% cumulative) and those who identified on the conservative end of the spectrum (48% cumulative): 13.5% strongly liberal, 25.8% liberal, 5.6% lean liberal, 7.1% independent, 11.1% lean conservative, 20.6% conservative and 16.3% strongly conservative.
Stimuli
For the current study, the measure of readability is based on the FRE scale and the FKGL scale. The FRE scale is one of the most widely used scales for reading comprehension used (Flesch, 1948). Developed in the 1940s during a plain-language movement, Flesch worked with the Associated Press and other media companies to make popular media more accessible to the general population, which was generally less educated than the people in the newsrooms (Flesch, 1948).
Stimuli were analyzed in an automatic readability analysis application on Readable.com (a free, automated reading comprehension analysis tool) and altered in ways to create content that fit two conditions— “high readability” and “low readability” based on these measures. Most of the modifications included changing long words to shorter ones, and making sentences shorter. The stimuli were also constructed to match in content issue and length (see the appendix). No attribution for author or news source was provided for content to avoid any confounding variables associated with authorship (Choi & Lim, 2019; Weibel et al., 2008).
Two news stories with two conditions each (high and low) were used so as to expand the external validity of this study. The first story was a news report on the number of states in the U.S. reporting West Nile virus (WNV) in mosquitoes for the 2021 season, and the second story was a business story about Mercedes-Benz investing in electric vehicles (EV). Stimuli were controlled for length to avoid any effect caused by a length factor. Table 1 shows the comparative data such as number of words, sentence counts and readability scale scores for both high and low conditions. The stories were repurposed from the publication The Hill, which generally publishes difficult-to-read news content (around 14 on the FKGL scale; Sparks & Hmielowski, 2023), and then content was altered for the purpose of readability. Both held the headlines from The Hill, but bylines, branding and other content (such as images) were removed. These changes consisted of shortening sentences and substituting words that were more than 12 letters or longer than four syllables (full text of the stimuli is available in the appendix).
Stimulus Manipulation
To measure readability, we relied on three well-established scales: FRE, FKGL and Gunning FOG Readability (GFOG) scales. These three scales focus specifically on syntactic complexity in written content. The scale scores for the high and low readability conditions are listed in Table 1.
Flesch Reading Ease
The FRE scale (Flesch, 1948) scores text between 0 and 100, with higher numbers indicating easier-to-read content. The formula for the FRE scale is based on the average sentence length and average number of syllables per word.
Flesch-Kincaid Grade Level
The Flesch-Kincaid Grade Level index matches the FRE number to the expected number of years one should have in formal education in the United States to comprehend the material (Kincaid et al., 1975). For example, a 9.9 on the FKGL would mean the person reading for comprehension is expected to have an average of almost 10 years of formal education. Preschool readings score below a 3 and anything above an 18 is graduate-level reading.
Gunning FOG Index
Similar to the FKGL, the GFOG index calculates grade level based on the number of years of education one would need to understand the work. However, on this index, a 0 is equal to no years of education and a 17 would be a graduate-level degree.
Dependent Variable: Message Credibility
To measure credibility, this study utilized Appelman and Sundar’s (2016) message credibility scale, which uses three items on a 7-point scale (1 = not at all, 7 = extremely) to gauge perceived credibility including whether readers felt the content was accurate, authentic and believable. In the current study, the message credibility scale was shown reliable, α = .75, M = 5.57, SD = 0.87.
Moderating Variables
Partisanship Strength
A partisan strength scale was developed using questions adapted from Huddy and colleagues’ (2015) study on partisan identity. Two 5-point Likert-type items were used to measure partisan strength, asking participants whether they agreed or disagreed with the content of the article (strongly disagree to strongly agree) and whether the focal issue of the article was personally important (not important at all to extremely important). The scale items showed moderate correlation, r = .42, p < .001. Participants were also asked to identify their political ideological preference on a 7-point scale from extremely liberal to extremely conservative (M = 4.04, SD = 2.18).
Media Preference
To test for media use and preference, two scales were created based on media use frequency as an indicator of reliance on the platform for news. Participants were asked to rate the frequency of use of eight media platforms. The first scale combined items for use of broadcast television, cable television, newspapers and radio, as they all represent legacy media. This scale was found reliable with a Cronbach’s alpha of .76 (M = 3.42, SD = 0.91). The second index combined television news websites, newspaper websites, blogs and social media, as they all represent digital-native media. This scale was also found reliable with a Cronbach’s alpha of .69 (M = 3.51, SD = 0.78).
Results
To study the effect of readability on credibility (
To test the moderation effects of partisan strength (
With concerns that the results were affected by ideological differences, we ran the PROCESS macro (Model 1) to test for interactions for ideology with readability and credibility. However, results showed no support for an interaction of ideology (b = .07, SE = .05, p = .17), so again
To test the any possible moderating effects of media preference (
Finally, because the effects of readability were tested across two stimuli variants for the sake of stimulus sampling, a two-way analysis of variance (ANOVA) was conducted with the significance of the interaction term assessed to indicate the invariance of readability effects across news story type. Results revealed that the interaction effect of reading condition and story selection on message credibility was near the threshold for statistical significance, F(1,245) = 3.98, p = .05,
Discussion
The present study theorized high readability would increase perceptions of news credibility. However, this experiment showed that was not the case, as this study showed no difference in credibility ratings for hard-to-read and easy-to-read content across two sets of stimuli. As previous research has posited that heuristics could be used to judge media for credibility (Chaiken, 1987; Petty & Cacioppo, 1986; Sundar, 2008; Sundar et al., 2007), this study shows that readability might not work toward that judgment.
In determining whether partisan strength would be a moderator on the effect of readability on credibility, this study did not show effect across partisanship strength or ideological measures. Despite previous research suggesting that previous attitudes about content would influence credibility judgments (Gunther, 1992; Stroud, 2011; Vraga & Tully, 2015), the results of this study showed no difference in perceived credibility among people with strong or neutral beliefs regarding the nonpolarized topics.
In addition, media preference did not seem to act as a moderator in credibility perceptions of easy-to-read and difficult-to-read news content. While prior work has suggested people who use web-based news would view it more credible than those who do not (Kiousis, 2001; Sundar et al., 2007; Tsfati & Cappella, 2005), the results of this research did not see a statistically significant difference in credibility judgments based on media preference.
Past studies of complex language suggest that jargon and technical language can deter readers from engaging in the content and elaborating on the messages, thus influencing the extent people rely on cues to assess message credibility and validity (Hafer et al., 1996; Ratneshwar & Chaiken, 1991). This link between language complexity and comprehension has been the subject of news readability scholarship for decades. In addition, this connection of comprehension and content is valuable to news outlets for establishing themselves as part of their community. In past studies of news readability, researchers focused on the need for more comprehensible content without explaining why comprehension matters. In their study of news readability, Dalecki and colleagues (2009) aimed to address the “news readability problem” by eliminating influential factors such as deadlines, format and managerial concerns as possible reasons for hard-to-read content. Danielson et al. (1992) suggested that journalists should worry about difficult prose because audiences would grow impatient with time-consuming material. Flesch (1948) addresses the need for better comprehension to create more interest, “The real value (of his reading formula), however, lies in the fact that human interest will also increase the reader’s attention and his motivation for continued reading” (p. 226). However, what this study found is that credibility, as one aspect of people’s judgments of news, is not related to readability or comprehensibility. If it takes people longer to read or understand a news article, that’s not going to be the reason they deem work disconnected or unworthy of their trust. Chall (1987) is one of the few readability researchers to explicitly explain that easier-to-read content is needed for societal masses because as societal complexity increases, greater comprehension is needed to analyze and solve problems; however, she was more concerned with adult literacy than news readability. The current work adds to this long line of research by suggesting that language complexity, and its association to comprehension, might not be linked to credibility. There are several factors that go into news selection (Rubin, 2009; Strömbäck et al., 2020; Tsfati & Cappella, 2005), and despite previous findings that show readability’s impact on comprehension (Chall, 1987), the readability might not play a role in the news selection or credibility process.
An alternative explanation for the findings could be that readability is only influential when people are motivated to elaborate and engage with the news content. This is similar to variables such as argument strength (Areni & Lutz, 1988), which Lutz and Areni argue might only produce effect when people are most willing to use cognitive effort. Most dual-processing theories suggest that people tend to be cognitive misers, in that we reserve cognitive energy when we can use cues to make decisions and judgments; therefore, readability might not be an important aspect to people making quick decisions. If this is the case, then this study helps to identify another way in which content traits might influence the way people process information and elaborate on it. This finding is important for text-based news operations as research has indicated that reading behaviors have changed in digital environments, where more time is spent browsing and scanning than intentionally reading. For example, Liu (2005) found that online readers spend less time on in-depth and concentrated reading than people used with print media. Thurman, in his report on the print–digital gap for the World Printers Forum found that people below the age of 55 spend dramatically less time and attention to newspaper brands in print or digital format than their older counterparts (Thurman, 2018). To explore whether levels of cognitive power moderate the impact of readability, future research might include tests for level of need for cognition and level of involvement, which could help clarify whether readability interacts with high levels of central route processing.
As with all research, this study has its limitations. In studies of dual-processing, research has shown the effect of need-for-cognition on one’s use of central and peripheral cognitive routes (Cacioppo & Petty, 1982). However, while we tested for need-for-cognition, a shortened six-item scale (Lins de Holanda Coelho et al., 2020) yielded unreliable alpha scores in the current study (α = .57, M = 3.37, SD = 0.57); thus, analyses with the need-for-cognition scale were excluded from the report. Future research might extend this study to see whether perceptions of credibility differ among subject matter across readability levels. The inclusion of level of involvement as a moderator would further add to the literature in terms of possible research in this realm, as has been indicated in previous credibility research. Researchers might also examine whether source credibility is affected by readability, as the author of the content was not used in this study. We tested conditions across two stories of different subject matter—one on health news and one on business news—the differences in structure and expectancy in these news fields might play a role in perceptions of credibility. This could have implications for several contexts of news content including different types of news, such as political or education, and some platforms for getting news, such as television and social media. In addition, it is important to note that this study focuses specifically on the United States and English-based media; thus, media from other countries and cultures might not see the same results (Hanitzsch & Vos, 2018).
This study set out to show that audiences valued comprehensibility as a factor of credibility perception; however, the results here indicate that is not the case. In his study of media literacy, Wasike (2018) calls for use of automated readability tests as part of the editing process so as to make the news more engaging and less discouraging for readers, suggesting that difficult-to-read content might be exacerbating supposed declines in written news credibility and trust. To the contrary, this study suggests that might not be the case. The use of shorter sentences and easier-to-read language, as the current study measured, might help with more comprehension but it will not help readers perceive higher credibility in the content.
