Abstract
Can engagement with far-right social media communities socialize users into a new political identity? This study addresses concerns about the spread of far-right groups on mainstream social media platforms by examining how newcomers are affected by their engagement with these groups. I theorize that changes in social identity expression indicative of socialization will be measurable in the language users use to express themselves on social media, and that the magnitude of this linguistic change will intensify with more frequent far-right engagement. I develop a custom dictionary of far-right-relevant terms sourced from communities like The Daily Stormer and Stormfront. Using an original dataset of Reddit user posting histories from 2015–2017, I test for increases in the frequency of this far-right vocabulary. I find that users who engage often with a far-right community like r/The_Donald begin to sound more like white nationalists within three months. Socialized users also use far-right vocabulary more frequently in other spaces on their platform, contributing to the spread and normalization of far-right rhetoric.
Introduction
How much influence do far-right groups on social media exert over their members’ identities? Concerns about the strong hold social media communities might have on members have grown since 2016, when a far-right digital movement calling itself the “alternative right” (or, “alt-right”) broke out of cyberspace and onto the national stage. The alt-right challenged preconceptions of American political identity and digital mobilization by demonstrating that fringe groups with relatively few members can significantly change the narrative of a US presidential election, influence the rhetoric of prominent politicians, and shift policy debates.
The alt-right is one of many concerning far-right movements propagating their ideas on major social media platforms like Reddit, Facebook, and Twitter. Of course, far-right groups do not only influence the rhetoric and narratives of “in-real-life” (IRL) politics; they can also inculcate civil unrest and violence. Social media communities like these have repeatedly enabled acts of mass violence against religious and ethnic minorities. These communities provide radicalizing misinformation, peer validation, and in some cases a literal audience for these acts. For example, in May 2022, ten Black Americans were killed in a racially motivated mass shooting at a Tops Friendly Market in Buffalo, New York. The shooter, Payton Gendron, had spent months learning about racist conspiracy theories in far-right communities on 4chan and Discord. Gendron posted a seven-hundred-page diary to social media detailing his plan to commit this act in the city with the highest concentration of Black people within driving distance of his house. He then livestreamed parts of the attack to members of a private Discord community he belonged to (Frosch et al., 2022). Gendron’s attack serves as a stark reminder that interactions with racist or conspiratorial communities on social media have likely played a role in deadly attacks on racial, ethnic, and religious minorities (e.g. Collins, 2017; Kirkpatrick, 2019).
Acts of domestic terror are the most visible symptoms of far-right community influence, but they represent only the tip of a much larger iceberg. The proliferation of white nationalist and conspiratorial communities on major social media platforms has garnered a number of normative concerns about larger sociopolitical effects: the promotion of white nationalism (Romano, 2016), the normalization of hate speech against minorities and women (Southern Poverty Law, 2018; Ward, 2018), and the spread of misinformation and conspiracy theories (Klein and Dunn, 2019; Zadrozny & Collins, 2018), to name just a few. All of the above concerns are grounded in the assumption that social media users are at risk of being influenced by these groups and their content. Recent works have made inroads into understanding the characteristics of far-right audience members (e.g. Schulze, 2020; Klein et al., 2019), but what the initial radicalization process might look like—and whether it affects all viewers or just the predisposed—remains more mysterious. We also struggle to measure the short-term effects of these encounters on social identity expression.
To address these gaps, I extend theories of political socialization into the digital context, asking if online communities are providing political socialization experiences similar to those found in traditional civic groups. This paper focuses on three major questions: (1) Do digital communities socialize members in similar ways as churches or labor unions?; (2) If so, can they socialize newcomers into a social identity well outside of the political mainstream?; and (3) Do socialized individuals proceed to spread far-right notions outside of far-right spaces?
Using social media conversation and activity data from the period in which far-right social media communities became salient, 2015–2017, I find that users who have never interacted with far-right social media communities on the platform before “sound” more like white nationalists within a few months of engagement. When users interact with far-right social media communities full of racist, sexist, and conspiracy theory-laden content, they rapidly adopt terms that reference that content in order to signal their assimilation into the group identity. They make more references to anti-Semitic conspiracy theories, Islamophobic memes, and racial slurs. And in a result suggestive of socialization, the most active users become more willing to use far-right language outside of far-right spaces. I conclude that far-right social media communities teach members important markers of social identity such as group-specific language, and that socialized users contribute to the spread and normalization of far-right rhetoric elsewhere on the platform.
The Digitization of Political Socialization
Group membership is central to our understanding of political identity and socialization (Campbell et al., 1960). Party identification is best understood as a set of overlapping group or social identities initially transmitted from our families and early social context and later refined, as individuals sort themselves into parties based on group stereotypes (Green et al., 2002; Jennings et al., 2009; Layman & Carsey, 2002; Mason, 2018). Social-psychological perspectives on partisanship treat political identity like other central and emotionally loaded social identities (e.g. religious identity). The perceptual screen that forms between a strongly politically identified person and their information environment structures political behavior, attitudes toward the out-party, and ability to update opinions based on new information (Green et al., 2002; Lippmann, 1922). Recent theories of partisan megaidentities suggest that the perceptual screen is becoming ever more influential (Mason, 2018). Today, a change in an individual’s political identity may also induce changes in the many zones of their life currently structured by politics, including behavior on social media platforms, media selection, and elements of identity expression like rhetoric.
It is through a process of political socialization that young Americans acquire in- and out-group attitudes that shape other political attitudes in adulthood. Traditionally, socialization is understood as a process of top-down intergenerational transmission of political knowledge and norms from older adults to youths, occuring primarily in the home and traditional civil society. Youths learn about political and civic participation from their parents and early social context (Andolina et al., 2003; Verba et al., 1995, 2005). They also learn about politically relevant group attitudes such as partisan animosity (Tyler & Iyengar, 2022), ethnic and racial identity and prejudice (Sears & Levy, 2003), and gender attitudes (Cunningham, 2001).
More recent work questions the uniform fit of the social or family transmission models and suggests that socialization is a more contextually dependent and dynamic process than previously conceptualized. Community group involvement can take a more central role where political socialization in the home is weak (Jennings et al., 2009), or where participation in a civic organization is particularly meaningful (Frisco et al., 2004). Immigration, socioeconomic, or class status also affect the manner and site of political socialization. For example, the children of immigrants are less likely to learn about politics in the home than their third-generation peers (Humphries et al., 2013), and the children of working-class parents may learn more about the structural disadvantages they will experience in the political system than their middle-class peers (Lareau, 2018). Young people can learn about political behaviors and attitudes from perceived socially integrated peers in their social networks, suggesting that young people are socialized in peer contexts as well as family and community contexts (Settle et al., 2011). And finally, socialization may be a dynamic and recurrent process. Political identities can change in response to a change in social context or new attitude toward one’s group after the influence of early childhood fades, into our twenties or thirties (Franklin, 1984; Lewis-Beck et al., 2008; Niemi & Jennings, 1991). Even group attitudes like racial attitudes that were previously long understood as causally prior to political identity, have become more dynamic in recent years; for example, Americans increasingly demonstrate a willingness to change their racial attitudes to match their partisanship (Engelhardt, 2021).
The American relationship to groups and community has been changing for decades. Conventionally speaking, in-person civic memberships shape political participation by providing participation-relevant political skills, an arena for group mobilization, and reinforcement of exclusive identities (Brady et al., 1995; Hill & Matsubayashi, 2005; Jennings et al., 2009; Putnam, 2000; Rosenstone & Hanson, 1993). Yet in-person civic memberships and the social capital produced by them have been in a state of decline since about 1970, leaving behind a vacuum in social connectedness that has likely contributed to lowered political participation. Americans today possess fewer civic group memberships, participate in fewer voluntary associations, and in general are less likely to be civically involved (e.g. attend a town hall or join a union) than their forefathers (Brehm & Rahn, 1995; Putnam, 1995, 2000).
In recent years, social media has become a locus of political media consumption and political socialization. This sea change in venue for American political learning, communication, and mobilization has likely had political implications. By the 2000s, young Americans were likely to have their first formative political experiences outside of the home, in online spaces (Xenos & Foot, 2008). A majority of Americans now report getting news on social media at least some of the time, with Americans under the age of 30 even more prone (Pew, 2023). Individuals who find their news on social media are exposed to different (and often lower-quality) political information than their old-media peers. For example, during the COVID-19 pandemic, Americans who primarily relied on social media for news were more likely to have heard the conspiracy theory that the pandemic was orchestrated (Mitchell et al., 2020). Replacing legacy news media with social media may contribute to normatively undesirable trends such as lowered trust and heightened affective polarization (Christensen et al., 2022), lower political knowledge (Wattenberg, 2008), lower interpersonal empathy (Kiesler et al., 2011; Konrath et al., 2011), and “tuning out” political news in favor of entertainment (Prior, 2005).
The digitization of the contemporary political experience has changed the groups Americans belong to. Digital communities provide alternatives to our disappearing in-person civic groups. Like in-person groups, online communities can transmit peer culture (Wattenberg, 2008); individuals appear to choose digital communities based on their ideological preferences (Barberá, 2015). However, digital communities differ from IRL groups in some substantial ways. Studies of computer-mediated communication (CMC) have long found that digital communities lack the social norms that would exclude antisocial individuals in IRL settings. Weakened norms and a lack of social feedback in the form of facial or verbal cues reduce peer empathy, enable or reward antisocial behavior (Connolly et al., 1990; Edinger & Patterson, 1983; Turkle, 2015), and facilitate the entry of fringe ideas into public discourse (Davis, 2009; Hill & Hughes, 1998; Kiesler et al., 1984).
Digital political experiences are not a one-to-one transcription of the analog to the digital. Americans are replacing in-person, community-based political socialization experiences with digital alternatives. This replacement is transforming elements of the American political experience in unpredictable ways, potentially including the manner and site of political socialization. Previous models of socialization emphasizes its role in perpetuating the status quo—instilling existing social orders, values, and civic skills in succeeding cohorts (Gordon & Taft, 2011, p. 17). Newer perspectives on socialization espouse a diffuse horizontal model of socialization that accounts for peer-to-peer interactions and recognizes the role of context. In the digital age, the particular social media space in which formative peer-to-peer interactions occur matters; spaces that have not traditionally been considered part of the political become so (Andersson, 2015, p. 973; Pfaff, 2009).
Current studies of political activity on social media often concentrate on identifiably partisan individuals loyal to major political parties. We know less about how users react to novel political identities encountered accidentally on social media platforms. If social media communities provide a space for meaningful performances of political identity and a social context for peer-to-peer socialization, then it is possible that pursuing such peer-to-peer interactions drives any influence far-right communities might exert over members. I propose that such interactions would have socializing effects, even on users without preexisting far-right beliefs. Finding evidence of such an effect lends credence to normative concerns regarding far-right groups on popular social media platforms.
Measuring Online Socialization
When are individuals most likely to learn elements of social identity from social media interactions? In other words, when do individuals learn the language of their group? This project articulates a proxy measure of change in social identity: change of identity-relevant language in social media posts. Thus far I have discussed the central role of groups in political socialization, as well as the increasingly digital, dynamic, and peer-to-peer nature of that process. I propose that social media communities are filling the vacuum left by the loss of civic life in America, and that this change in venue is likely to have consequences for political learning, just as it has for our media diets and information searches.
However: online communities are still communities, and they possess similar features. Online communities dedicated to particular social identities maintain the boundaries of those identities in analogous ways to in-person civic groups. Many readers may be familiar with online communities and typical forum rules governing who can participate and what may be posted. Like any social group, an online community sets rules for acceptable behavior and language. Group members and moderators enforce the boundaries of the group’s exclusive identity and worldview. Social media posts that violate community rules are often taken down or flagged by moderators. This type of norm enforcement incentivizes users to conform to a community’s expectations in terms of the language they use and the topics they choose to discuss, even when the consequences for violating those norms are purely digital in nature.
People are especially likely to pick up new words when first entering a community (Danescu-Niculescu-Mizil et al., 2013). Linguistic learning accompanies drastic changes in social context; in the past, this was mostly associated with periods of adolescent exploration, or moving to a new region with an unfamiliar culture. Today, it may be associated with joining a new digital community. The field of Internet linguistics has explored Twitter users’ adoption of specialized slang terms found on the platform. Even while accounting for global linguistic drift, evidence suggests that 2014 Twitter users learned new words from the peers they followed very rapidly, especially if those words were only encountered online: “For rising words that are primarily written, not spoken, [such as] abbreviations…and phonetic spellings…the number of times people saw them mattered a lot. Every additional exposure made someone twice as likely to start using them” (McCulloch, 2019, p. 30, emphasis added). In other words, lingo unique to online spaces is adopted more rapidly than words encountered in everyday life. Peer-to-peer interactions have particularly strong socializing effects when the networked individuals perceive each other as demographically similar (McCulloch, 2019). Users are more likely to mimic strong ties in the social network than weak ties. Since Twitter users are more likely to follow individuals with demographic similarity to them (Centola et al., 2007), language tends to spread through strong network links among individuals who perceive a shared identity.
My approach relies on Giles and Johnson’s (1981) ethnolinguistic identity theory (EIT), which articulates the close relationship between social identity and language. EIT holds that “social identity and ethnicity are in large part established and maintained through language” (Gumperz & Cook-Gumperz, 1982, p. 17). In other words, language is not just a way to perform identity—it is inextricable from identity. People constantly make linguistic decisions based on social identity. “[W]e …align ourselves with the existing holders of power by talking like they do …Sometimes, we decide to align ourselves with particular less powerful groups, to show that we belong and to seem cool, anti-authoritarian, or not stuck-up” (McCulloch, 2019, p. 41). For example, Canadian teenagers make decisions about how to pronounce their alphabet: switching from the American “zee” to the Canadian “zed” signals pride in their national identity (Chambers, 2002). Group members are motivated to achieve a more positive status in their group by signaling assimilation into the group identity (Verkuyten, 2021).
How does in-group identity affect our language around out-group members? Consistent with the principles of social identity theory, identity performance can be dynamic and contextual (Tajfel & Turner, 1979). Individuals perceiving their in-group to be the object of discrimination often perform code-switching behaviors when in the presence of outsiders. Contextually reliant code switching has long been studied in the context of African American identity expression, a practice extending from survival mechanisms among enslaved peoples. However, code switching behavior can involve any number of social identities: in Labov’s classic linguistic experiments, New Yorkers keep or drop their Rs to match the social class performance of their conversation partner (Labov, 1962). Queer youth often employ coded language in certain environments, such as Facebook, as they seek to communicate with peers without being outed (McCulloch, 2019, p. 232; Wilson et al., 2016). Coded language conceals illicit ideas from outsiders and it is often deployed differently among in-group members than outsiders.
Important to note here is the narrative of personal transformation central to many contemporary digital far-right movements. Far-right communities emphasize the life-altering nature of the perspectival shift experienced by new members. They describe this shift in terms similar to a religious rebirth. One 2017 commenter in The Daily Stormer forums describes the process as follows: “I started my red pill with Holocaust revision videos on JewTube and went from there. Once))) people (((realize the whole (((system))) is cucked they come around pretty easy.” Interested readers will note that this individual uses several pieces of coded language to communicate much ideological content to others in the know. Nested triple parentheses ((())) denote Jewish identity or something under Jewish control, while their inverted form )))((( denotes the opposite. Thus this individual’s use of codes reveals their beliefs about who is in control of “the system,” and who isn’t.
Language is inextricable from identity performance. If language is often learned online from those that we perceive are similar to us and used to signal group belonging, then language used in social communities may be an effective proxy measure of social identity. Individuals participating in social media communities centered on a particular identity are likely to pick up group-specific vocabulary and use it to signal group assimilation. However, far-right-engaged individuals might choose not to deploy norm-violating language when posting in non-far-right spaces.
Linguistic Learning through Online Socialization
The social media communities we belong to serve as approximate substitutes for in-person civic memberships, and thus provide a similar opportunity for Americans to perform political identities in front of in-group members. Social media users can rapidly learn new words from peers on the platform, particularly if these peers are perceived to be similar to the user, or if the peers represent strong network ties. Novel terms that are only found in that online space are most likely to be adopted, and this acquisition is particularly rapid when users have first entered the space. Sites with anonymous user accounts allow members to experiment with socially unacceptable ideologies like white nationalism or anti-Semitism without family or work colleagues finding out—in other words, without any long-term consequences to their reputations. 1
These two expectations regarding acquisition of novel group-relevant vocabulary and uncensored performance form the basis of H1. I hypothesize that users will use more far-right vocabulary following their first engagement—in the form of posts and comments—with a far-right community. H1: Users who have engaged with far-right communities will use more far-right-group-relevant language after engagement than before engagement.
The social media communities we belong to reflect the political identities that we subscribe to. EIT predicts that an in-group member will use group-relevant vocabulary to signal the strength of his or her identification with the group. If the amount of engagement with a community is a good proxy measure of a person’s degree of socialization, there should be a positive correlation between a user’s far-right engagement and increase in far-right vocabulary rate over time. H2: Users who are more/less engaged with far-right communities will exhibit a greater/lower increase in the rate of far-right language across all platform activity.
If engaging with peers in far-right social media communities provides a form of political socialization, then highly engaged users would acquire a group identity that shapes their engagement on all parts of the platform and among both in-group and outside users. Highly engaged, and therefore heavily socialized, users will be likely to use far-right terms around in-group members as they signal group belonging. Thus, strongly identified users will likely use group-relevant vocabulary at higher rates in far-right spaces than weakly identified users. H3: Users who are more/less engaged with far-right communities will exhibit a higher/lower rate of far-right language use inside far-right spaces.
Concerns regarding the normalization of offensive ideologies on social media are based on the assumption that these ideas could spread beyond their dedicated spaces. Therefore, learning whether engagement with these communities makes spread more likely is crucial. If engaging with a far-right community encourages the adoption of an identity—something “sticky”—users’ self-expression should change globally, not just in dedicated spaces. I anticipate that code switching may occur.
2
However, given the centrality of anti-PC sentiments to most far-right groups, and their celebration of any behavior that shocks “normies” (outsiders), my primary prediction is that socialized users will use group-relevant terms more often outside of their communities, in spite of the terms’ norm-violating nature. H4: Users who are more/less engaged with far-right communities will exhibit a higher/lower rate of far-right language use among outsiders (outside of far-right spaces).
Research Design
This project analyzes text found in public posts from the social media platform Reddit. Reddit is a popular social media platform with about 330 million monthly users worldwide. Reddit now boasts a more American user base than Facebook: 48.6% of Reddit’s active users live in the United States, compared to only 12% of Facebook users (Statista, 2021). Reddit is composed of a networked system of sub-forums, called subreddits, each dedicated to a discernible topic (e.g. r/cats or r/programming) or social identity (e.g. r/conservative). Subreddits must declare a community description and a set of community rules governing member behavior and permissible content. Subreddit communities are internally moderated for the most part, meaning that in-group members police other in-group members and enforce community norms. Community descriptions, rules, and moderators together define the boundaries of permissible ideas and behavior within each group, making it likely that users who participate regularly are aware of the particular culture of that space.
All of the data for this project is scraped from the PushShift.io API, data scientist Jason Baumgartner’s massive archival collection of Reddit data. Any content available at the time of a PushShift scrape is captured, meaning that the Pushshift API provides access to data missing from Reddit’s own API such as deleted or banned posts, users, and subreddits. As Reddit began purging its platform and archives of its most infamous far-right communities in 2017, this data can now only be accessed via Pushshift. The dataset for this project contains records of over 700,000 Reddit posts from between January 2015 and December 2017. I perform text analysis on the full text of over 69,500 posts containing more than 2.3 million words.
Community and User Sampling
The average Reddit user is regularly exposed to posts from a variety of communities they did not subscribe to. This is because Reddit employs a flat, uncensored, and user-generated content promotion system that treats content engagement and quality as synonymous. Controversial or explicit content that garners high levels of engagement—even negative engagement—often rises to the top of the site. That engagement-driving content is collated at the top of r/all, a page that skims quality content from all topics and community types. Though Reddit has increasingly taken measures to hide non-work-appropriate content behind special tags and prevent rule-breaking or offensive content from reaching r/all, the site’s tendency toward uncensored content promotion often resulted in far-right content reaching the top of r/all prior to February 2017. This phenomenon led to many user concerns about the possible normalization and spread of far-right rhetoric. 3
I identify newcomers within a sample of posts made in r/The_Donald in January 2017. r/The_Donald was originally founded as a fan club for then-presidential candidate Donald Trump. However, it soon became far more ideologically extreme than the name suggests, eventually earning a reputation as a significant generator of fake news and misinformation sourced from “alternative news” sites (Zannetou et al., 2017), promotion of hate speech and white nationalist memes (Southern Poverty Law, 2018), and repeated rule violations regarding racist and violent content eventually leading to a 2019 quarantine and 2020 ban. 4 Although some smaller subreddits have claimed far more explicit far-right labels (r/altright being the classic example), the sheer size and activity level of r/The_Donald—peaking at almost 800,000 subscribers before its banning—combined with its long history of content violations recommends it for a study of far-right newcomers. 5 In fact, r/The_Donald content had such a nasty habit of reaching the top of r/all that Reddit changed its rules specifically to bar the group’s posts from sitewide promotion.
This project relies on identifying far-right newcomers: individuals who have never interacted with a far-right community prior to posting in spaces like r/altright or r/The_Donald and thus are unlikely to have known far-right vocabulary before the timespan of interest. Under Reddit’s pre-February 2017 structure, any user looking at the most-upvoted content on the site (found on r/all) would have seen content from r/The_Donald, whether they chose to or not. r/The_Donald’s big-tent nature, high activity level, and knack for landing on the trending page make it a likely site of first-time far-right engagement and thus newcomers.
Data and Variables
I ask if changes in identity expression correspond with changes in community engagement patterns for Reddit users. If a user who has never interacted with far-right subreddits in their previous years on the platform begins to post in a far-right subreddit like r/The_Donald on a regular basis, and the language they use to express themselves changes at the same time, we come closer to understanding the effects of joining a radical social media community.
It is not easy to observe individual behavior on Reddit. This is part of the reason so many far-right groups developed on the site in the first place. Records of the communities Reddit users subscribe to are not publicly available, meaning the only indicator of what a user reads is engagement data—records of posts and replies to other users. Luckily, records of where, how often, and what a user posts on the platform are publicly available and conveniently anonymized. And due to the site’s organization into subreddits, the data is prelabelled by topic or group identity as well. Archived Reddit posts are a rich source of long-form text data that provide a toehold on the observation problem.
Identifying Newcomers
Determining if the language users employ in their posts changes following engagement with a far-right community requires that I identify newcomers unlikely to enter the community pre-socialized, which I accomplish by sampling users from the most active subreddit in the far-right space at the time, r/The_Donald. However, I categorize user activity later based on a list of far-right communities provided in Online Appendix 2. Furthermore, as the project taps normative concerns over contagion and normalization of extremist communities, it is crucial to identify users who found far-right content like this when it was still being promoted by the Reddit’s algorithm. In February 2017, Reddit made a series of unannounced changes to its system of content promotion in order to prevent r/The_Donald content specifically from reaching the r/all page. These were changes intended prevent the average Redditor from stumbling on the content unknowingly. Therefore I look for far-right newcomers in a sample of r/The_Donald posters from January 2017 expecting that these users participate in other far-right spaces as well.
I generate random timestamps between midnight on January 1, 2017, and midnight on February 1, 2017. I scrape the first 100 posts—PushShift.io’s request limit—in r/The_Donald at each timestamp and note the username that generates each post. I continue this process until I collect 5000 unique username values. I select a random sample of 1600 usernames from that sampling and scrape the first 100 posts of each user’s monthly posting history from January 1, 2015, through December 31, 2017, recording which subreddits the user posts in each month.
Once I record where and how often each user in the sample posts each month, I find each user’s initial engagement point—the first month in which each user posts in a far-right community like r/The_Donald or r/altright (see Figure 1 for range of engagement dates). I record the month and year of every user’s first far-right engagement. I then use this information to filter my sample on two criteria: (1) three months or more posting history prior to a username’s first post in a far-right community, and (2) at least five posts per month in the three months following initial engagement (including the month in which engagement first occurred).
6
Range of initial far-right engagement dates.
I filter out users who had posted in a far-right community immediately after account creation in order to exclude “lurkers,” users who observe a community and learn its rhetoric before making a Reddit account. I also filter out accounts labelled as “throwaways” 7 and bots. 8 544 accounts remain after this filtering procedure. It is worth noting that out of the original sample of 1600 usernames from r/The_Donald, about a third belong to far-right newcomers.
Sampling Activity Before and After Far-Right Engagement
I measure change in identity expression by comparing the way Reddit posters expressed themselves three months before posting in far-right groups to how they express themselves after three months of engagement. I design my pre- and post-test sampling procedure following the conventions set by other studies of group behavior on Reddit (e.g. Buyukozturk et al., 2018). Here I measure changes in user behavior over a period of about seven months of Reddit engagement. The pre-test corpus is sampled from three months before the user’s first far-right engagement, and the post-test corpus is sampled four months after this point, with the three months between engagement and post-test comprising a treatment period. For an example of this individualized sampling procedure, see Figure 2. Individualized sampling procedure.
Keeping three months of cushion between pre-test and initial far-right engagement minimizes the risk of gathering data contaminated by lurking. 9 Measuring change after three months of engagement in far-right spaces gives a user time to be socialized. I generate pre- and post-test corpuses for each user from the beginning and end of that seven-month period. This sampling timeline allows enough time for a user to engage with far-right communities for several months before the post-test.
IV: Far-Right Engagement Percentage (FREP)
This model treats the three months between first engagement and post-test as the treatment period for each user, in which more engagement constitutes more treatment (i.e. socialization experiences). To determine the strength of the treatment I record the far-right engagement level of each user in the months between initial engagement and post-test. I disaggregate this engagement by community type: far-right versus all others. The far-right engagement percentage (FREP) represents a volume-insensitive measure of how the share of a user’s engagement occurs in far-right spaces in that treatment period. The FREP measure ranges from zero (posts were made during the treatment months but none of these were in far-right spaces) to 100 (posts were made during the treatment months, all in far-right spaces). I expect that a higher FREP will predict higher rates of far-right term use in the post-test corpus. For more detail on subreddit classification, see my discussion of precursor subreddits below.
Custom Dictionary of Far-Right Terms
I form a dictionary of group-relevant terms based on the expectation that engagement with far-right subreddits may change member linguistics by teaching words referring to ideological elements of contemporary far-right groups: white nationalism and racism, misogyny, conspiracy theories, anti-Semitism, and anti-political correctness. I source my far-right-relevant terms from several projects. First, I rely on my own term frequency analysis of 2017 forum threads from The Daily Stormer (N threads = 3, 632). I supplement this with other linguistic analyses of the alt-right (Sonnad & Squirrell, 2017; Squirrell, 2017). 10 My far-right term dictionary includes words that were rare and unusual in the context of 2015–2017. Some of these words have since been popularized by the alt-right and journalists reporting on the alt-right (e.g., pilling has since entered the common parlance), while others remain unusual. Dictionary terms are selected to minimize ambiguity and reduce the risk of false positives. I choose words that tap central ideological elements of the far right, such as globalist/m, which refers to a set of conspiracy theories about Jews’ allegiances to a secret global order. I also select references to obscure memes that reference in-group culture: for example, kek refers to the satirical worship of a frog-headed Egyptian god resembling the Pepe the Frog meme. “All hail Kek,” or “Praise Kek” are popular identity-signaling rallying cries on sites like The Daily Stormer.
I source other terms from a linguistic analysis of the “manosphere,” a collection of misogynistic communities that share a number of linguistic traits with far-right groups like the alt-right (Sisemore, 2020). This includes terms like red pill: similar to Neo in The Matrix, users who take the red pill “wake up” from a false reality and perceive the world as it truly is. In spaces like r/Incels (a contraction of “involuntary celibates”), women are often referred to with dehumanizing terms like femoid, meaning “female humanoid”. These are only a few of the many pieces of coded language shared across far-right spaces. For further examples of group-relevant terms and their sourcing, see Online Appendix 1.
DV: Rate of Far-Right Terms per 1000 Words
After stemming and removing stop words from my corpuses, I determine the total number of words in each and calculate the term rate by dividing the number of dictionary terms used by the number of words in the corpus. For the sake of intelligibility this measure is expressed in terms of thousand-word units. I use these rate comparisons to calculate the change in term rate between pre- and post-test as well. I distinguish between identity expression within far-right spaces and other spaces by disaggregating the post-test corpus by community type. Using term frequency analysis I can determine the rate of group-relevant words and phrases before a user engaged with far-right spaces, after they did so, and how they spoke within and without those groups.
Controls: Precursor Subreddits
Recent work on conspiracy theory-centered social media communities like r/conspiracy (Klein et al., 2019) finds that user engagement with fringe communities is determined by an interaction between individual and social factors. A preexisting conspiratorial disposition, say, could cause an individual to self-select into Reddit communities dedicated to the discussion of conspiracy theories. Just in case any preexisting propensities could make a user more likely to engage with far-right subreddits (though as described above, my sampling procedure minimizes this possibility), I control for users’ past engagement with far-right-related content in the months or years they spent on the platform prior to the sampled time period. Users possessing precursor traits such as a preexisting interest in components of far-right belief systems such as racist, misogynistic, conspiratorial, violent, or other offensive content may have visited subreddits associated with these topics prior to engagement with far-right political communities. For example, a user discussing the QAnon conspiracy theories on r/The_Donald might have originally learned those theories on r/GreatAwakening.
In order to control for precursor traits that might provide an alternative explanation for a user’s knowledge of far-right-relevant vocabulary, I compile a list of the largest communities associated with (1) conspiracy theories, (2) offensive humor, (3) explicit racism, (4) misogyny, (5) violence, and (6) gaming. 11 I rely particularly on lists of subreddits recently banned for these kinds of content violations, 12 filtering out left-wing communities where relevant. 13 The resulting lists of precursor subreddits provides me with a more holistic understanding of each user’s preexisting propensities. For a complete accounting of far-right and precursor subreddits, see Online Appendix 2. I generate a series of dummy control variables indicating whether each user had engaged with one of the precursor community types before engaging with far-right groups.
Evidence of Socialization in Far-Right Communities
Date Variable.
Continuous Variables (Expressed as Terms per 1000 Words).
Nominal Variables.
Note: Descriptive statistics for all variable types. Tables were generated using the stargazer package (Hlavac, 2022).
Term frequency analysis using my far-right dictionary reveals the most common terms found in the dataset (see Figure 3). To test if the terms captured in the post-test corpus were novel relative to each user’s pretest corpus, I test for per-user term overlap and find that overlapping terms (pictured in Figure 4) were relatively rare. Only 85 out of 2767 (3.07%) of terms found in users’ post-test corpuses had also been found in their pre-test corpuses. Only seven words overlapped more than three times; these words were slave, plot, SJW, cuck, Hitler, and Nazi—some of the least specialized terms in the dictionary. Top 25 far-right terms found. Within-user term overlap.

H1 predicts that users who engage with far-right subreddits will use more group-relevant terms after that engagement. To test this hypothesis, I find the difference in rate of term use between users’ pre- and post-test corpuses. I run a two-tailed dependent t test and find that across the entire sample, users use group-relevant terms at a significantly higher rate three months after engaging with a far-right subreddit than three months before (μ = 3.99 terms; t [544] = 8.49; p < .01), with an average increase of 3.99 terms in per 1000 words. A density plot of the increase in term use rate after those 6 months is displayed in Figure 5. This significant increase is interesting given the specialized nature of these terms and their rarity in general. Taken together, the results of the t test and overlap analysis suggest that users were using more group-relevant terms three months after engagement than three months before, and that in most instances these terms were novel to the user. On average, users across all engagement levels had acquired a significant number of in-group-relevant terms three months after their first exposure to r/The_Donald. Increase in term rate after far-right engagement.
Effect of Socialization on Change in Global Term Rate.
Note: *p

(a) Effect of Engagement on Change in Term Rate.
The coefficients of the linear model reveal a positive relationship at traditional levels of significance between the FREP score and change in term rate (ß = .06, p < .01). Users whose FREPs near 100% use an average of 6.36 more terms per 1000 words in their post-test corpuses than users with FREPs near zero. Spending a greater percentage of some given total of engagement with far-right communities in the three months prior corresponds with greater increases in far-right term use in the post-test.
Effect of Socialization on Term Rate in Far-Right Spaces.
Note: *p
Effect of Socialization on Term Rate in Other Spaces.
Note: *p
H4 predicts that users who spent the treatment period being highly engaged with far-right communities will be more likely to use far-right terms freely wherever they are in the post-test point, even when posting in communities that could be hostile to far-right references. Therefore I estimate a linear model regressing FREP on term use in other spaces. The linear model finds a positive relationship that differs from zero at a statistically significant level (ß = .04; p < .1). The most engaged users employed approximately 4 more terms per 1000 words around potentially hostile outsiders than the least engaged users (see Table 5 and Figure 7). Taken all together, these models indicate that the most engaged users use far-right terms more often around outsiders than less engaged users. Effect of engagement on term rate in other spaces.
This project tests a series of hypotheses in an attempt to determine if engagement with far-right social media communities like r/The_Donald constitutes a digital socialization process, whereby users with no history of far-right involvement could learn a far-right group identity. I predicted that a change in identity expression should be measurable as group members learn and use group-relevant terms in their posts. My study examines change in language use over time, over different levels of engagement, and in friendly and unfriendly spaces. I find evidence of a global increase in term frequency following the first three months of engagement. I test for a relationship between socialization, as measured by the percentage of a user’s monthly engagement in far-right spaces (FREP), and term rate. This measure that allows me to compare users with differing amounts of free time or talkativeness. FREP is a conservative measure, in the sense that any effects I find should be weakened by not distinguishing users by posting volume. Despite this conservative research design I find higher term frequency among users with higher socialization (FREP) scores. Thus the FREP measure may successfully capture higher levels of group involvement associated with adoption of group political stances, internalizing norms of behavior, and learning new markers of social identity such as a group’s unique coded language.
Based on the preexisting literature on linguistic learning on social media, H1 predicts that users engaging with far-right communities for the first time would rapidly adopt the group rhetoric in their first few months of posting. The more unusual and internet-specific the terms encountered are, the more rapidly a newcomer should learn them. I compared term rates in sampled posting activity from three months before first engagement to a sample taken after three months of engagement. My results indicate support for H1. Across the entire sample, users use far-right terms more frequently after three months of far-right engagement than they had six months prior—approximately 4 more terms per 1000 words, or twice as often. Given the specialized nature of many of the words in the dictionary and the heteroskedasticity in activity levels and date ranges, this is a notable change in identity expression. Almost all (97%) of the terms detected in the post-test corpuses were novel to the Redditor who used them, suggesting some level of learning had occurred within individuals.
If engaging with a far-right community repeatedly over several months is analogous to attending many meetings of a traditional civic group, then more frequent engagement should result in more rapid socialization and a greater change in identity expression. H2 predicts that users who started posting often in far-right spaces after never doing so before would exhibit more language change than their less engaged peers. I find evidence in support of H2. Highly engaged users exhibited larger increases in far-right term use between pre- and post-test corpuses—an increase of about 6 and a half terms per 1000 words—than those who had rarely engaged with far-right communities in the months previous. This result was robust in spite of variables controlling for alternative explanations, such as previous engagement with conspiracy or manosphere communities.
By examining in- and out-group behavior, H3 and H4 tap the central question of this paper: whether the changes in behavior I observe among far-right newcomers are best explained by a change in social identity, or by a simple change in posting location. If the former, I expected to see changes in behavior across all of sampled engagement with the platform, whether occurring in far-right spaces or not. If the latter, I expected to see changes in language use within far-right communities only (though, as previously noted, this could have also been explained by codeswitching). While my results do not reveal a significant difference in rate of term use within far-right spaces between users with very high and very low FREP scores (H3), they do reveal a surprising effect on posting behavior outside of far-right spaces (H4). This is particularly notable due to the controversial nature of far-right language, though I am unable to measure how easily outsiders could identify far-right identity expression in 2015-2017. Highly engaged users spend more of their time posting in places like r/The_Donald or r/alt_right, so they have less time and energy left over to post in other places. However, when they do post around outsiders, they are more likely to write about topics like anti-Semitism, racism, sexism, etc., than less engaged users, even if this risks offending outsiders.
There are two major takeaways from this last finding. First, higher rates of far-right term use in unaffiliated spaces among users who spend a lot of time in those communities indicates that the change in behavior is not accounted for by simple location effects. Highly engaged users are not just using more far-right terms because they are posting in a community that rewards this kind of content. Instead, the language they use to express themselves changes in general, and self-censorship around critical outsiders is not performed. Put another way, highly engaged users aren’t just talking about Hitler around other Nazis—they are talking about Hitler to everybody, whether they want to hear it or not. Changing the language one uses by adopting terms significant to one’s community fulfills the expectations of socialization into a group identity, not a simple topic effect. Second, these findings reinforce oft-expressed concerns about the systemic effects of far-right communities on large social media platforms. Highly engaged users may self-quarantine to some extent by not posting in many places other than r/The_Donald. However, they are more likely to spread their group’s ideas when they do choose to post around outsiders. Thus, even removing a far-right community’s ability to be promoted by a platform’s algorithm is not a perfect solution—users spread these ideas themselves.
Putting Term Rates into Perspective
In order to test the generalizability of my dictionary and term rate measure, I apply my text analysis procedure to a well-known alt right community, The Daily Stormer, as well as a long-infamous white supremacist community, Stormfront. I perform the same term rate analysis on these corpuses: a sample of 3614 posts scraped from Stormfront, and a sample of 3631 conversation threads scraped from The Daily Stormer. I expected to find more dictionary terms in the Daily Stormer corpus due to the theorized linguistic transmission from standalone alt-right groups to Reddit far-right groups. Both samples were collected from activity occurring between March and June 2017. The comparison of results is displayed in Figure 8. Term rate baseline comparisons.
Readers will note that the site with the highest incidence of dictionary terms is The Daily Stormer. This is hardly surprising—much of the dictionary was sourced from the same site. Additionally, there are few checks on behavior or language on the site. It is independently operated and at the time of data collection experienced little censorship. Readers will note that the post-test corpus exhibits a term rate twice as high as the pre-test corpus, and that the term rate in far-right Reddit spaces is only 4 terms per 1000 words lower than that of the Stormfront corpus.
Discussion
Popular concerns over the danger far-right content poses to unaware social media users rely on a few key assumptions: First, that far-right communities do not just come into contact with preexisting adherents, but are capable of converting new members to the cause. Second, that far-right concepts are referenced outside of far-right-dedicated spaces on large platforms like Reddit. And third, that regular users without a preexisting interest who unknowingly stumble onto far-right content might find it appealing. Examining these assumptions requires that social scientists address new and uncomfortable questions regarding the degree to which political experiences have been changed by the digitization process so accelerated in the last decade. The range of political identities Americans come into contact with is more diverse, and more likely to include noxious fringe elements, than ever before. Anonymous spaces allow people to perform extreme political identities with little risk to their reputations, and to change their political milieu with little effort. Political identity is increasingly divorced from ‘‘real’’ life, and long-term costs for shifts in identity expression are reduced, meaning that identity itself is becoming more plastic and fragmented.
The results of this study suggest that ethnolinguistic identity theory is useful to the measurement of far-right identity on social media. In the absence of data on all of the communities users are subscribed to and the posts in their feeds, researchers must rely on what and where users post to determine what the user is thinking. In this case, users displayed a significant increase in use of group-relevant slang terms within a few months of engaging with a far-right subreddit. The increase was greater among users who spent a higher portion of their time on the site engaging with far-right peers. In most instances, users were using terms three months after their first engagement that they had not used three months prior, a result suggestive of learning. Rate of term use in external spaces was higher among highly engaged users—meaning that once users spend enough time engaging with far-right in-group peers, they cease to code switch. The far-right group identity becomes more central to these users and is more likely to be expressed around outsiders. Finally, all of these results were robust to precursor community control variables. Even when I take into account previous activity in sexist, racist, or otherwise offensive subreddits that might contain similar terms to far-right communities, it is engagement with the latter that tracks with term acquisition most consistently. Overall, it is highly possible for a Reddit user with no prior activity in far-right groups to develop an interest in these communities and adopt group behaviors in a manner suggestive of a social identity shift. Such a shift appears to occur when those users start engaging with far-right communities and intensifies with higher engagement.
Why should we care about mass consumption and normalization of the rhetoric of a few white nationalist or conspiratorial groups on social media, especially if they do not have plans to form a national party or run for office? First, regardless of membership count, these groups have repeatedly demonstrated an ability to outpunch their weight class in terms of influencing well-known politicians. White nationalist and conspiratorial rhetoric shared on social media is frequently transmitted upward: content produced by these fringe groups is reshared around social media, eventually reaching the feeds of prominent political figures who subsequently repost or discuss that content, thereby broadcasting it for the entire nation to see.
To name just a few instances of this phenomenon: former President Trump retweeted content created by alt-right activist Jack Posobiec in August 2017 (Golshan, 2017), a racist (mis)infographic about black crime originally sourced from a UK white supremacist account (Farley, 2015; Greenberg, 2015; and PolitiFact, 2015), and a video originally posted on the far-right subreddit r/The_Donald (Pearce, 2017). After Hillary Clinton’s “Basket of Deplorables” speech in September 2016, Donald Trump Jr. reposted an alt-right meme to his Instagram account, in which the faces of the alt-right cartoon character Pepe the Frog, two “alt-lite” media figures, and the Trump campaign team were photoshopped onto a poster for The Expendables 2 (renamed “The Deplorables”) (Nguyen, 2016). Other prominent DC politicians have revealed the presence of white nationalist or conspiratorial content in their feeds via their public rhetoric: for example, before and after gaining office, Marjorie Taylor Greene publicly promoted the Parkland shooting “crisis actor” theory, expressed support for the QAnon conspiracy theory (Greene, 2021), and referenced multiple anti-Semitic conspiracy theories involving Soros and the Rothschilds (Coleman, 2020; Hananoki, 2021). Whether these politicians shared white nationalist or conspiratorial content unwittingly or not, their actions demonstrate how easily fringe groups can sidestep media gatekeepers and make an end run to the Oval Office just by sharing a meme on Twitter.
Recent scholarly work confirms that the extremist-to-elite pipeline issue remains a serious problem: for example, Kennedy et al. (2022) find that 2020 election conspiracy theories about Dominion voting machines were popularized by a combination of elite amplification and activity among a relatively small number of repeat spreader accounts. Researchers in American politics are still teasing out the long-term consequences of this pipeline for public trust in mainstream media and information. Thus far, research is not reassuring. Christensen et al. (2022) find that the mainstream media respond to short-term market pressures by disseminating trust-reducing content to audiences outside of social media, participating in the very media phenomenon leading to the long-term delegitimization of the mainstream media.
It is my hope that future work addresses the limitations of this study. More research will be needed to determine just how lasting the influence of social media communities is on political identity. Political socialization may less embedded in civil society than before, but it may also have less long-lasting effects on our identities and behaviors. Identities can be adopted and shrugged off more easily in the digital world. Individuals with multiple accounts can fragment their identities by spending time in a variety of spaces and performing a variety of personas. Experimental work may be needed to comprehend the full scope of one individual’s online behavior from posting to viewing. Such a study would likely be necessary to fully eliminate the analytical challenges presented by throwaway accounts, multiple accounts, or other means of splintering one’s activity online. Additionally, with each act of deplatforming, studies of far-right activity on the internet will become more difficult. Most far-right groups banned from mainstream platforms like Reddit have migrated into invite-only spaces on platforms like Discord. What data scholars can still access—that which has not already been erased by tech companies seeking to erase ugly pasts—are ever more precious.
Americans are increasingly likely to have their formative political socialization experiences in digital communities. Political learning is qualitatively different in digital spaces. Users are more likely to encounter political identities that are outside of the political mainstream, and they may be more willing to experiment with socially unacceptable ideologies with low-cost, anonymous activity. Socially unacceptable ideas that violate social norms can be incubated in insular social media communities and broadcast to other spaces, contributing to ongoing issues of spread and normalization of extreme rhetoric.
Supplemental Material
Supplemental Material - r/The_Donald Had a Forum: How Socialization in Far-Right Social Media Communities Shapes Identity and Spreads Extreme Rhetoric
Supplemental Material for r/The_Donald Had a Forum: How Socialization in Far-Right Social Media Communities Shapes Identity and Spreads Extreme Rhetoric by Vivian Ferrillo in American Politics Research
Footnotes
Acknowledgements
I thank my advisors Edward Carmines, Kevin Banda, and Christopher DeSante, whose support and guidance made this work possible, as well as my colleagues Matthew C. Lucky, Volker Schmitz, and Sam Bestvater, who provided invaluable feedback.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Ethical Statement
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
