Abstract
Social media communities are understudied due to the difficulty of obtaining data with any amount of generalizability. This research proposes a method of hybrid probability sampling of social media ecosystems (collections of communities) using network theory, stratified sampling methods, theoretical saturation, and widely available social media activity statistics. Because of the unique venue of research, these methods create an opt-in sample, where every member of a community is given an equal chance to participate. The proposed methods were then used to obtain a sample representative of a digital political ecosystem with a population size of 8 million. The results illuminate stark polarization resulting from algorithmic opinion aggregation, with implications in online extremism, media literacy, demographic representation in public discourse, and more.
Introduction and literature review
The use of social media communities in social science research is still underutilized and underappreciated as a source of distinct topical cohorts and influential information sources. While some papers have given noteworthy accounts of the utility of social media for non-probability sampling (Stern et al., 2017), this research proposes a novel approach to representative random sampling of social media communities. This methodology seeks to create a hybrid probability sample from a digital ecosystem. Digital ecosystems for the sake of this paper are defined as a network of multiple social media communities with shared topical interests (e.g. politics, hobbies, interests, demographics). Probability sampling of these networks will allow researchers to draw generalized conclusions about them and their dynamics.
Current methods for social science research within digital spaces are underdeveloped. With the growth of social media, has come the need for a robust field of study focused on the development of best practices for the research and analysis of social media communities. Current methods have found strong success with ethnographic approaches (Beran, 2019; Donovan et al., 2022), but representative sampling remains elusive. Initial literature on the digitization of social interactions focused heavily on the use and utility of big data as well as the embrace of computer science methodologies for studying said data (Burrows and Savage, 2014; Savage and Burrows, 2007). While the contribution and subsequent expansion of digital social science epistemology introduced new avenues of study and methodological tools, more traditional methods were only partially utilized to research these newly established social media communities. The dismissal of digital survey research and traditional probability sampling methods was based on technology of the time and such concern went largely uncontested as information access evolved (Coomber, 1997; González-Bailón and Xenos, 2020; Savage and Burrows, 2007). Some research has gone as far as dismissing the possibility of probability samples within social media communities, with little justification (Stern et al., 2017). The flaw in this thinking comes from the inability to frame social media communities as we would physical communities in our efforts to gain representative samples. Framing social media communities not as a proxy for real world communities, but as their own stand-alone populations, allows us to think of them as a cohort that can be sampled with robust statistical accuracy. Current literature suggests that social media networks, while valuable for research purposes, cannot be sampled representatively. This new method demonstrates how representative sampling within these networks is possible.
At the beginning of the mass adoption of the internet by the global community, researchers noted that it was impossible to know who used the internet, who had multiple accounts, and who just “lurked” without engaging in content exchange, and as such concluded that the only valid use of the internet for social research was as a focus group, or a way to generate mailing lists (Coomber, 1997; Fisher et al., 1996). In the decades since these conclusions were reached, the adoption of the internet has grown rapidly, and the information conveyed by internet communities is far more sophisticated in terms of user metrics and statistics. Unfortunately, the utilization of digital communities in social research has not changed to take advantage of this evolution. Today, the statistics required to model the populations of digital networks are widely available. Despite this, a brief synopsis of current methods employing social media communities for social research reveals continued reliance on the internet as a focus group tool and little more. Most disciplines simply use web scraping tools for big data analytics (Kobayashi and Ichifuji, 2015; Wang et al., 2017), real world surveys about social media community use (Abney et al., 2019; Cassell and Tversky, 2005; Colliander et al., 2015; Lowry et al., 2016; Tosun, 2012), or simple observational/social media community focus group techniques (Abidin, 2013; Oosterhoff, 2014; Skelton et al., 2020).
Digital populations
One key development born of the past two decades of internet growth is the collation of distinct social media communities. These social media communities, like those found in the physical world, are built around a topic, interest, ideology, or identity. Identifying the unifying theme of the community of interest can either be implicit or explicit depending on the social media platform utilized. For example, Reddit communities are explicitly created by users and given the distinction “subreddits”. These subreddits are given rules, moderators, and topical guidelines by the users who created the community. On Twitter, topics can be tagged with “hashtags” to allow others to find the current public discussion on the topic. The platform is structured such that users are meant to “follow” individuals they find interesting, and users’ natural interests and similarities create implicit communities. On Facebook, groups can either be implicit or explicit depending on the behavior of the community of interest.
The utility of social media communities as a tool for social science research is wide ranging. Social media communities can have direct impacts on multiple aspects of society including social support, psychological adversity, radicalization/extremist recruiting, and multi-million dollar financial movements. During the coronavirus pandemic the widespread success of the videogame Animal Crossing: New Horizons provided a natural experiment on the use of digital spaces simulating social interaction and the psychological impact that community provided (Zhu, 2021). In contrast, digital communities on Instagram have been found to have a direct negative impact on the mental health and body image of teenagers (Cohen, Newton-John and Slater, 2017; Kleemans et al., 2018; Ridgway and Clayton, 2016). Radical groups utilize social media communities in order to discuss ideology, mobilization, and public perception, creating an easy-to-reach community for the study of radicalization and extremism (Scrivens et al., Frank, 2020; Wojcieszak, 2009, 2011). Social media communities have also been used to organize massive socially motivated, counterintuitive financial activity as evidenced in the GameStop short squeeze in 2021 (Chohan, 2021). The use and future utility of social media community research is likely to increase as newer generations adopt social media at near-universal rates. According to Pew Research Center, 86% of Americans born after 1981 report using social media (Vogels, 2019), and as such are members of one or more social media community.
Social media allows populations to sort themselves according to their own agency as opposed to their geography. While plenty of research has been done on the actions and impacts of these social media, the populations within these spaces remain unstudied as discrete monolithic groups.
Digital ecosystems
Social science as a field must think of social media communities not as a supplement of a physical community but as independent populations. Social media communities create self-sustaining digital ecosystems that must be studied as we have done with physical communities (Kleineberg and Boguñá, 2015). For the purpose of this research, digital ecosystems will be understood as a collection of social media communities with a shared organizing feature (e.g. networked communities for a specific demographic group, political ideology, or interest in discrete topics). While the concept of social media ecosystems has been raised by other researchers (Abney et al., 2019; Del Vicario et al., 2016; DeVito, Walker and Birnholtz, 2018; Hanna, Rohm and Crittenden, 2011), the methods for representative research of these ecosystems remain somewhat sparse. Most recent research pertaining to digital ecosystems focuses on ecommerce or accessible education (Gill and Germann, 2022; Harrisson-Boudreau and Bellemare, 2022; Márton, 2022). Meanwhile, sociology remains interested in methods of online community research, but has seen limited innovation in the methods for researching these communities (Hampton, 2017).
For the purpose of probability sampling on social media, social media communities and the digital ecosystems they make up are the online spaces in which a researcher’s target population “resides”. Digital ecosystems are made up of networked communities designed for the consumption, analysis, and creation of content and opinions, related to a distinct interest or trait. The boundaries of a given digital ecosystem are determined according to how granular a researcher wants to get on the topic of choice. Much like survey research conducted in a city, researchers must make decisions about whether to draw their sample strictly from those within the city limits, or if they are to include the suburbs as well. Of course, researchers will also learn about true digital topical boundaries by the responses given to the sampling method discussed later. Though in theory a digital ecosystem could include several interconnected platforms and their features, this preliminary proposed method has only been tested within a discrete platform.
Because most social media websites involve some level of indexing according to identity, recommendations, interests, total website activity, or subscription, creating an appropriate sampling frame simply means identifying the scope and breadth of the networked communities that make up the digital ecosystem of interest. This can be achieved with network theory.
Well-established digital ecosystems are a realization of the Watts–Strogatz small world network model (Watts and Strogatz, 1998), in that the probability (P) of random connectivity (K) between random nodes is greater than 0 (regular fully connected networks) and less than 1 (fully random networks), while the clustering coefficient (C) is always greater than C random thanks to algorithmic and topical community sorting. Digital ecosystems are made up of clusters that are connected relationally, but not completely, positioning them between P(0) full networks and P(1) random networks.
Given this, we can expect topically similar communities to have only a few degrees of separation between them, in addition to communities being aware of communities with which they share similar contextual questions (those within their ecosystem network). This small world network nature of communities was demonstrated by Stanley Milgram in the 1960’s (Milgram, 1967), and remains intact today (Ferreira et al., 2021). Community connections that make up small-world networks are not hindered by the digital nature of social media communities, as they rely on statistical probabilities and social psychological principles. The added complexity of social media characteristics to the legitimacy of small world networks would only heighten community connectivity thanks to technological enhancement of this social inclination, namely, algorithmic connection recommendations. On top of normal social psychological small world network characteristics, the design infrastructure of social media incentivizing community connectivity only enhances the applicability of the small world network model to digital ecosystems. This suggests that building a sampling frame of a digital ecosystem can be done by consulting communities about their knowledge of other similar communities (ask each node about its edges). While doing so does involve navigating homogeneous networks, that same homogeneity is what allows us to draw conclusions about the representativeness of our sample. Given appropriate sample sizes for the size of the ecosystem, any remaining unknown discrete community would not be representative of the ecosystem as a whole, and as such, leaving them out of one’s sampling frame would not hinder the strength of one’s arguments for representativeness of the whole. Though one user may be unaware of the existence of an adjacent community, an entire subpopulation will not, so long as the adjacent community is prominent and durable.
This network theory model for probability sampling is not without caveats, however. Because of how quickly some online fads move in and out of the mainstream, some digital ecosystems are at risk of changing too fast to form durable small world networks. As such this model is only appropriate for use within ecosystems well established on the discrete social media platform of choice. Caution should be exercised when doing hybrid probability sampling on fad topics. The creation, or dissolution, of individual communities is less important, however, as the ecosystem is made up of many communities, and the most prominent will remain substantial nodes from which to determine edges for further stratified representative sampling. In addition, platform consistency is important when focusing on digital ecosystems. Much like biological ecosystems in different hemispheres, an ecosystem on one social media platform, does not necessarily share properties of interconnectedness or inter-community reliance with those on another. This sampling technique is appropriate for work within a platform, not between. Finally, the utility of these methods is theoretically limited to social media platforms with good topic indexing, and an active user base with sub communities. An example of appropriate use cases would be Reddit, Discord, Facebook or 4Chan.
Methods
This method of probability sampling on social media employs the use of exponential non-discriminative snowball sampling of social media communities (snowball sampling on the community level as opposed to the individual level). This is done to the point of strata sample saturation. By relying on social media communities to construct a networked digital ecosystem, concerns about strata sample validity and non-random sampling are alleviated. Because the population of interest (digital ecosystem/ a subset of social media) meets the parameters of a small world network, a stratified sample of all communities within an ecosystem can be found with minimal responses. Sample size will be determined by the size of the digital ecosystem population. Knowing when to stop snowball sampling is a balance between obtaining an appropriate overall sample population size for the digital ecosystem, and appropriate allocation of that overall sample across strata (communities) according to the ratio of their active population to the overall ecosystem population. For example an ecosystem with 10,000 members would require a sample size of 515 (given 98% confidence level and 5% margin of error), and a community within that ecosystem containing 5,000 members (50% of the total) would require 258 responses. See Table 1 for more details.
Activity statistics for different communities should be obtained to adjust for the possibility of inactive accounts. Combining community member statistics with activity statistics to create active community estimators will provide robust sample size estimates with underestimated statistical power. For example while a community may boast 20,000 subscribers, additional statistics such as comments per subscriber or average unique page views can be used to identify a true estimate of the active community.
Once a stratified sample of the ecosystem of interest has been identified, social media metrics can be utilized to organize strata according to community size and intended sample population. This allows researchers to weight sampling according to particularly active/large hubs within the network. Decisions about research goals and overall sample representative power can be made and sample sizes can be adjusted accordingly.
The target population for this methodology is a random subset of active users of the digital communities of interest. Surveys are opt-in, and advertised by bulletins, posts, or announcements to the entire community of interest, at varying times throughout the day and week. Because this survey strategy calls for an opt-in participation opportunity given to every member of an active community within the data collection period, this creates a representative sample of the digital community. Because individuals are only allowed to participate once, activity metrics are based on individual accounts, and sampling benchmarks are in place, hyperactive members will not be overrepresented in the sample. A long enough data collection window will allow users of all activity level to participate. Further, though opt-in surveys could create a selection bias, extending sampling windows to receive the mandated ratio of active users to community population alleviates the initial selection bias. This creates a hybrid probability sample (as opposed to a nonprobability sample) because every member of a population (a distinct digital community) is given an equal chance to participate using volunteer methods (a hallmark of probability sampling), and strata participation rates are set by their activity estimator as it compares to the ecosystem as a whole.
The digital ecosystem is a networked population made up of distinct communities (strata). To survey the entire ecosystem, calculate a total sample size for the entire ecosystem. Sort the strata by size and determine what percentage of the total ecosystem is made up of members in each stratum. Use this percentage to determine how many survey responses from each stratum will be included in your final analysis. Calculate stratum sample sizes according to representative research goals (data representative at the ecosystem level or down to the community level).
Procedures
Step 1. Identify the parameters of the digital ecosystem (population) of interest. Setting clear guidelines for what does and does not fall into the ecosystem of interest will be crucial for setting sample size goals and determining representativeness during analysis. Some digital communities (population strata) are not clear about what unites them as a community, and as such it is important to establish protocol for what does and does not fit the research interests. Best practice would suggest updating protocol if the users responses point toward some missing inclusion criteria.
Step 2. Identify a few key communities (population strata) in which to begin sampling frame construction. The frame will grow quickly as the community sampling method implies. Include any available metrics for the size and activity of the community of interest in sample frame data tables.
Step 3. Draft, pilot, and distribute a survey to the communities (population strata) identified in step 2, inquiring about their knowledge of communities that meet the parameters identified in step 1. For example “Please list 3-5 similar communities (individuals, groups, subreddits, hashtags, etc.) on (digital platforms of interest) dedicated to creating, discussing, or sharing content about politics or politicians”.
Step 4. Add any unique responses (previously unidentified communities/strata) to the community sampling frame created in step 2. Distribute snowball sampling survey to any newly identified communities.
Step 5. As the sampling frame grows, sort communities by size/activity metrics gathered and calculate the total population of the currently identified ecosystem (collection of strata). Keep a running calculation of required sample sizes for each community (strata) according to its size/activity relative to the total ecosystem population. Community sample sizes and representativeness preferences should be determined according to the researcher’s goals.
Case study
Reddit was chosen for this research for its focus on indexed communities with readily available activity statistics. Having content and social functions based on communities instead of individual brands/expression makes it an ideal tool for administering surveys. Reddit’s political ecosystem was chosen as a target for this exercise in probability sampling of digital ecosystems, as it contains many distinct and well-known sub cohorts. As a note: Reddit is made up of millions of communities known as “subreddits”. This nomenclature is used throughout this paper.
For this research the top 50 political subreddits (in terms of active members) were identified using snowball sampling techniques and subscriber metrics which were then sorted into the top 25 according to community activity metrics. An activity score was computed for each of these top 25 communities by multiplying the number of comments per subscriber by the number of subscribers. These activity scores were then used as estimators to create a proportional sample size target per strata and total sample size for all strata. In effect, these research design decisions make the findings representative of the most active communities, as described by the most active number of users. The target functional sample size for this research was 2662. The total population of this ecosystem was 8577000. It should be noted the true ecosystem population size is likely smaller as some of the users in the subreddit user metrics are likely to be inactive accounts.
All subreddits were contacted before the start of the surveying to obtain consent from moderators. All subreddit moderators were given copies of the survey questions and informed of the intent behind the survey project. Identifying information was required to participate in the survey. This was done to avoid repeating participation, and to obtain consent for the project.
After notifying moderators, posts promoting the survey and the broader project were made on a weekly to monthly basis depending on how many responses were needed in each community. Posts were made on varying days throughout the week and hours throughout the day to capture the attention of different users available online at each time.
As a note, one of the most popular communities in the strata sample experienced a significant departure from the apparent topic of interest during data collection. The moderators of this community decided against enforcing the guidelines for community topics. After monitoring the community for several weeks, it became clear the community’s namesake was no longer relevant to the topic of discussion and generally not political - outside of the interest of this research. As such, this community was removed from the sample list and the next most popular subreddit took its place.
The final raw sample included 3,000 responses taken proportionally across the top 25 most active political subreddits (our ecosystem). Although sample stratification only involved the top 25 communities, users from over 200 subreddits were represented in the sample. Post-hoc random sampling of strata responses was completed in order to avoid oversampling of any distinct communities, for a total working sample of 2,669 participants, seen in table 1. The survey administered to the sample contained questions about demographics, a user’s 3 favorite politics subreddits, and a user’s 3 favorite news sources.
Communities that served as digital ecosystem strata, and their subsequent sample weighting based on activity estimators
Note: It should be acknowledged that the absence of a particular community from the list provided does not imply its exclusion from the overall sample. Our study encompassed a broad spectrum, identifying over 200 distinct communities. However, due to practical constraints, we set sample size requirements for a selection of 25 communities. Furthermore, it is essential to highlight that the subreddit populations and comments per subscriber metrics utilized in this research are founded on statistical data from May 2021
Case study results
Results from this demographic and news source preference survey found that the digital ecosystem on Reddit is overwhelmingly younger than 34, male, white, and mostly democratic in their political preferences. Table 2 shows the numerical and proportional breakdown.
Numerical and proportional breakdown of the demographics of Reddit’s political ecosystem
Note: This survey did not discriminate by country of residency. “Other” refers to many different political parties around the world such as the Canadian liberal party and Switzerland’s SVP. Additionally, “socialist” means something quite different in the US in contrast to Spain for example, so interpretations should take this into account.
Additionally, Reddit’s political ecosystem has preferences for the following news sources: Associated Press, BBC, CNN, The New York Times, NPR, Reddit, Reuters, The Economist, Washington Post, and YouTube. For this question respondents were asked to give open responses regarding their 3 favorite news sources. The numerical breakdown in table 3 should be interpreted as inferential of the entirety of Reddit’s political ecosystem, not as a discrete preference between these 10 choices. For example 20% of all users of Reddit’s political ecosystem name the New York Times as one of their favorite news sources.
The top 10 most preferred sources for news content according to Reddit’s political ecosystem
Review & Discussion
This research set out to find a theoretical way of obtaining representative hybrid probability samples using social media. With some caveats, the application and theoretical statistical validity of these methods is reputable and replicable. It is entirely possible to obtain hybrid probability samples from a digital information ecosystem under the correct conditions. The data in the case study is representative of the Reddit political ecosystem, but not each individual community. With more surveying, the methods could be used to obtain a representative sample of each community as well.
Although these methods are novel in their reliance on network theory, activity metrics, and cohort stratification, they are not without limitations. The most prominent being the dynamic nature of digital community evolution. While these methods are representative of the political digital ecosystem on Reddit as of 2021, these communities experience changes in population size with some regularity. The sample size estimates remain robust for some time following the data collection, as some unknown number of accounts in each community population count are indefinitely inactive and these methods assume every member of a community population is a potential participant, therefore underestimating statistical power of the stratified sample collected and giving the methodological validity buffer room as the statistical power decays with the changing of the community. Despite this, some population turnover is inevitably expected, even disregarding overstated community population growth. A community cohort will not stay consistent over the course of several years, especially in online spaces. As such, although these methods are sound for observations and inferences within a distinct period, the underlying statistical power does become obsolete faster than statistics on a physical population. Future research on how quickly digital populations turn over would help to identify how long statistical outputs from these methods are valid for.
As always, there are also bound to be individuals within the community who do not wish to participate for an unknown reason. Non-response bias is a problem in every field of the social sciences. Interactions with the community can help to alleviate these concerns to some extent. For this research regular communications with community members and moderators were made in an attempt to earn the trust of the potential participants. There is no reason to suspect, however, that non-response issues would be greater in digital communities than in physical ones. Here too, future research could shed light on the difference in non-response of digital communities versus physical ones. Although it should be noted that there is some evidence to suggest individuals feel more comfortable sharing information online, than in person (Oosterhoff, 2014; Tosun, 2012).
This method and case study focused on representation at the ecosystem level. While this has its obvious research benefits and interests, it was not tested at a fine granular level in order to gain community level representation, for inter community comparisons. While this is likely possible, it surely comes with unique challenges not identified by this paper. Likewise, it is very plausible that ecosystems do expand beyond a single platform in some cases. It is plausible to imagine using similar methods for creating networks between platforms from which to sample even broader ecosystems, but that is a project suited better for future directions.
Future directions for this research could involve more robust sampling of distinct communities in order to obtain enough community-level statistical power to compare differences between cohorts. As digital spaces become more popular and grow in their natural ideological stratification, it will be important to understand how different sub-cohorts consume and understand information through ideological lenses.
Footnotes
Declaration of conflicting interests
The author declares that he has no conflict of interest.
Funding
The author declares that he has not received any funding for this research.
