Abstract
Social media are becoming more popular as a source of data for social science researchers. These data are plentiful and offer the potential to answer new research questions at smaller geographies and for rarer subpopulations. When deciding whether to use data from social media, it is useful to learn as much as possible about the data and its source. Social media data have properties quite different from those with which many social scientists are used to working, so the assumptions often used to plan and manage a project may no longer hold. For example, social media data are so large that they may not be able to be processed on a single machine; they are in file formats with which many researchers are unfamiliar, and they require a level of data transformation and processing that has rarely been required when using more traditional data sources (e.g., survey data). Unfortunately, this type of information is often not obvious ahead of time as much of this knowledge is gained through word-of-mouth and experience. In this article, we attempt to document several challenges and opportunities encountered when working with Reddit, the self-proclaimed “front page of the Internet” and popular social media site. Specifically, we provide descriptive information about the Reddit site and its users, tips for using organic data from Reddit for social science research, some ideas for conducting a survey on Reddit, and lessons learned in merging survey responses with Reddit posts. While this article is specific to Reddit, researchers may also view it as a list of the type of information one may seek to acquire prior to conducting a project that uses any type of social media data.
As response rates fall, costs rise, and the desire for timely data increases, social scientists are turning to alternative data sources such as social media (Murphy et al., 2014; National Research Council, 2013; Schober, Pasek, Guggenheim, Lampe, & Conrad, 2016). This includes the most popular social media sites such as Facebook and Twitter (Schober et al., 2016), but it also includes other platforms such as Reddit. Reddit.com is the self-proclaimed “front page of the Internet.” Registered users may post text, a link, images, or video onto any of thousands of subreddits, pages dedicated to a given topic (e.g., German news, recipes, or cat photos). Other individuals may read, vote them up or down, or comment on the post.
Reddit is an attractive data source for several reasons. First, its popularity means that there are a lot of data available for analysis. It is the 18th most popular website in the world and ranks even higher in some countries such as the United States (#6) and Germany (#9; Alexa, 2018). (As a point of reference, Facebook is the fifth and Twitter is the 33rd most popular site worldwide at the time of this writing [November 2019].) Second, most subreddits (exact statistics are unavailable) are public and those that are private can often be accessed on request. Public subreddits can be viewed by all, posts may be commented and voted on by all registered users (registration is free and anonymous), and the data may be downloaded for free by any registered user. Third, researchers may also collect data by posting links to surveys that may be completed by anyone (see https://www.reddit.com/r/SampleSize/ for current examples), giving researchers access to a low-cost survey mechanism and, if the Reddit username is provided, the ability to link survey and social media data. Fourth, there are over 138,000 active subreddits, allowing researchers of rare populations to easily identify and study eligible participants. For example, research into the emotional needs and concerns of expectant fathers has previously been limited due to issues in study recruitment. However, content analysis of subreddit r/PreDaddit allowed clinical psychologists to gain a better understanding of their concerns and develop perinatal education programs to address them (Pilkington & Rominov, 2017). Finally, Reddit provides an anonymous environment that redditors may use to express their true beliefs. For example, an analysis of Reddit posts was conducted to determine the reasons that some individuals are ambivalent or opposed to seeking treatment for their eating disorders and the types of support they get to reinforce their beliefs (Sowles et al., 2018). It is unlikely that these individuals would have been forthcoming with this information to survey researchers in a doctor, biasing conclusions drawn from traditional research methods. While still rare, some social scientists have capitalized on benefits of Reddit data and used them to study a host of topics including gender roles (Ammari, Schoenebeck, & Romero, 2018), news consumption (Wasike, 2011), sexual identity (En, En, & Griffiths, 2013), and mental health (Choudhury & De, 2014). As more researchers use Reddit, we expect that the findings drawn from the data will become more widely acceptable and understood.
When deciding whether to use data from social media, it is useful to learn as much as possible about the data and its source (Murphy et al., 2014). Social media data have properties quite different from those with which many social scientists are used to working, so the assumptions often used to plan and manage a project may no longer hold. Information about the amount of data, type of users, how to access the data, and more can be used to compare data sources, determine whether the data are sufficient for a given purpose, create a budget, determine what skills are required and how to staff the project, develop a time line, identify hardware and software needs, and brainstorm and mitigate the risk of negative events over the course of the project. For example, most social scientists work with data sets that may be opened in SAS or Excel, often stored in a generic file format such as American Standard Code for Information Interchange. Social media files may use different file formats such as JavaScript Object Notation (json), which takes different programming languages and skills to open and analyze. They may also be much larger than traditional data files, so large, in fact, that they cannot be processed on a single machine. A researcher may opt not to use social media data because they do not have the resources required to read and analyze the files. However, if the challenges of opening and analyzing the data are not known ahead of time, significant time and money may be devoted to collecting these data only to encounter this impasse later.
Similarly, knowing what data are accessible, how the data are generated, and for whom they are generated is useful in understanding the most suitable analytic, sampling, or weighting techniques or what error sources one may need to consider when drawing conclusions and making inference. While not a social media example, Google Flu Trends may help illustrate this point. In 2008, Google Flu Trends was touted as being able to successfully predict the number of individuals suffering from the flu (Ginsberg et al., 2009). However, in 2009 and again in 2013, this same data source severely misestimated flu prevalence. In one instance, the error occurred because Google’s user interface had changed, and users could now see and click on others’ searches, and in the other instance, people’s search behavior changed as news coverage of the H1N1 outbreak increased (Butler 2013; Cook et al., 2011). Understanding how individuals interact with the search engine and how that behavior may influence the analysis may have allowed researchers to avoid the observed bias.
Unfortunately, the information required to make informed choices and create management and analysis plans is frequently not included in the methods section of peer-reviewed papers. For example, in a study using tweets to predict the country’s mood, Bollen, Mao, and Pepe (2011) only mention that they analyzed over 9 million tweets between August 1 and December 20, 2008. No information on how they obtained, cleaned, or stored these data was included. These details may not be relevant to their research question, but they are helpful to other researchers attempting to use similar data sets.
Some information may be gleaned from IT sites such as StackOverflow.com or blog posts, but it is not comprehensive and often requires that the researcher has a good idea of what questions she should be asking. Some researchers may learn this information through their professional networks. However, young researchers may not have well-developed networks, and much of this knowledge is dispersed across multiple fields, making it difficult to form a clear picture of the data before the start of a project. Others may use methodology reports published online to get details of the practical implementation of a project, but these types of reports are frequently limited to federally funded or large-scale surveys for which social media data have yet to be used. Finally, researchers may learn on the job. While this is effective, it is not efficient, as is partially evidenced by the works cited; the information for this article comes from Application Programming Interface (API) documentation; programming books and blogs; literature found in computer science, sociology, survey methods, and epidemiology; preexisting knowledge from a strong multidisciplinary team; e-mails to numerous people at several institutions; and a lot of on-the-job trial and error.
A few Twitter and Facebook data users have taken it upon themselves to compile practical information about conducting social research on these platforms (see Annice et al., 2013 or Jürgens and Jungherr, 2016, for a summary of using Twitter data and Baker, 2013 or Wilson, Gosling, & Graham, 2012, for Facebook). Relatedly, De Vreese and his colleagues (2017) have compiled a set of considerations when linking survey and social media data. However, a summary has yet to be written for Reddit. In this article, we attempt to fill this void by providing descriptive information about the Reddit site and its users, tips for using organic data from Reddit for social science research, some ideas for conducting a survey on Reddit, and lessons learned in merging survey responses with Reddit posts. While this article is specific to Reddit, researchers may also view it as a list of the type of information one may seek to assemble prior to conducting a project that uses any type of social media data source. In describing what this article is, it is also important to note what this article is not. We do not make any judgment on whether Reddit is an appropriate data source. That is up to the researcher and her specific use case. Our aim is only to provide researchers with the information required to make those judgments.
The Reddit Population
In this section, we describe the type of people found on Reddit and how those individuals interact with Reddit. Given the speed with which social media demographics are changing, the statistics cited in this section will be outdated by the time of publication. However, we have included them here because we believe the main conclusions that may be drawn from these statistics will hold for several years: Reddit is large and will likely have sufficient data for several types of research objectives and questions. The sociodemographic characteristics of redditors are not comparable to the general population. The way in which redditors use Reddit (e.g., engagement vs. consumption of content) may affect the way in which researchers wish to use Reddit data.
We encourage readers to use the statistics cited in this section as a starting point to investigate the current landscape.
Reddit has a vast scope with over 330 million active users per month (https://www.redditinc.com/). Table 1 shows statistics about the site as of November 2017 (https://www.redditinc.com/; https://www.digitaltrends.com/social-media/reddit-ads-promoted-posts/). These numbers are worldwide, but the U.S. accounts for 41.4% of the screen views while the UK, Canada, and Germany, collectively, account for an additional 18.9% of screen views (https://www.alexa.com/siteinfo/reddit.com, accessed November 2019). As mentioned above, these data will be outdated by the time of publication. To be sure, the U.S. share of screen views has dropped by 29% (from 58.6%) since October 2018 (https://www.alexa.com/siteinfo/reddit.com, accessed October 2018). Data on growth and trends over time are available on GitHub through 2015 (Reddit’s 10 year anniversary; https://github.com/drunken-economist/reddit-10-year-data), and the cites/sites listed above are updated periodically with more recent statistics.
Reddit Usage Statistics.
While Reddit is popular in several countries, English appears disproportionately represented. In our analysis of 367 subreddits that originated in German-speaking countries (Germany, Austria, Lichtenstein, Luxembourg, and Switzerland), 1 we expected less than 5% of submissions to be written in English based on the proportion of native English speakers in each country (British Broadcasting Corporation (BBC), 2014). (The chosen subreddits were identified from a list on the DACH Reddit wiki page [https://www.reddit.com/r/dach/wiki/k] that was meant to include all subreddits originating from German-speaking countries. This list is maintained by users, and there was no method to confirm its completeness.) However, using R package CLD3, we found 68.7% were written in German, 10.6% were in English, and 20.7% were in an undeterminable language or language other than English or German. Unfortunately, we were not able to identify any additional information from Reddit or other sources to determine the generalizability of this finding to other countries or languages. Moreover, we expect some error in these estimates. Upon manual review of a subset of submissions classified as undeterminable or other, we found most were written in German but used nontraditional language, abbreviations (e.g., “LOL” instead of “laughing out loud”), or emojis.
Reddit is anonymous, and anyone (even without registration) can view most subreddits. Those who wish to post, vote, or comment on Reddit (and those who want to view/subscribe to private subreddits) must register. Registration only requires the creation of a username and password. While many accounts are also associated with a validated e-mail address, this is neither required nor are e-mail addresses accessible to researchers. Because no personal information is collected, the demographic makeup of the 330 million active users is unknown. However, evidence suggests that users are more likely to be male and younger than the general population. Pew conducted a telephone survey of the U.S. general population in which they asked whether individuals used Reddit. They estimated that 6% of all U.S. adults used Reddit, but men were twice as likely as women to use it (8% vs. 4%, respectively; Duggan & Smith 2013). The overrepresentation of men on Reddit has been corroborated by our own research and that of Singer, Flöck, Meinhart, Zeitfogel, and Strohmaier (2014), though the magnitude of this overrepresentation widely varies. We estimated that men were 19 times more likely to use Reddit while Singer and his colleagues estimated they were 3 times more likely. Reddit users also tend to skew younger. Reddit estimates that 79% of worldwide users are 18–34 years of age (https://www.digitaltrends.com/social-media/reddit-ads-promoted-posts/). This estimate is slightly lower but relatively consistent with our estimate (84%) and Singer and colleagues’ (2014) estimate (90%).
While gender and age are highly skewed, the device type on which Reddit is accessed is relatively consistent with expectations; 40% of web traffic in the United States and 51% worldwide occur on a mobile device compared to approximately 49% of Reddit traffic (Drunken_Economist, 2016; StatCounter, n.d.). 2 Unfortunately, post- and comment-level information do not include the type of device on which the entry was made, so this cannot be used in analysis of mined data. However, if researchers are interested in launching a survey or posting other content, they will likely need to consider device optimization so all Redditors can access the information.
All estimates of the demographic makeup of the Reddit population are likely to contain some error. One reason for the discrepancy across sources is the difference in who is being counted. Because individuals are not required to login to read Reddit content and because, even if they login, they are not required to post or comment, there are significantly more lurkers than participants. Reddit’s CEO, Steve Huffman, estimates that two thirds of new Reddit users are lurkers (over 6 million users per day; Protalinski, 2017). This is consistent with our findings that suggest 82% of Reddit consumers read subreddits in which they do not subscribe (and likely, though we do not have the data to be sure, do not post). Unfortunately, neither of these statistics differentiate how people use Reddit’s various features. Singer and colleagues (2014) addressed this limitation and found that while 79% of respondents reported never or seldomly posting original content, the vast majority reported commenting on others’ posts or voting on posts at least occasionally (71% and 83%, respectively).
Reddit Organic Data
Most researchers use Reddit for the organic content created by users, which consists primarily of posts, comments, and votes. However, a plethora of other metadata also exist including date of submission, username, and “flair” (a tag that users attach to their post to describe the type of content—think of it as a keyword—or to themselves to describe their role or interests). There are two ways to access Reddit content—via download from http://files.pushshift.io/reddit/ (maintained by redditor Stuck_In_The_Matrix) or via the Reddit API. While this may seem like splitting hairs—the same data for both come from the same place—there are some significant differences in what data are available. These differences are summarized in Table 2.
Differences Between Two Reddit Data Sources.
Note. API = Application Programming Interface.
Perhaps the most important difference between the two data sources is the size. The downloadable data sets are stored as large compressed json files organized by month and by type of content (e.g., post or comment). The compressed comment files we accessed averaged 6.1 GB each and were as large as 9.0 GB once unzipped. Given the size and depending on the research needs, researchers may not be able to open the entire file at once and will need to subset it by reading one line at a time and determining whether to keep or discard it before moving to the next line (see Appendix A for Python code to do this). Whereas the downloadable data sets are large (making them much more difficult to work with), the Reddit API can be used to target specific subreddits, users, variables, time periods, types of content and more, limiting the data extracted and stored. This benefit may also be a limitation as the API restricts the amount of data that may be pulled down at any given time or by any given search.
Additional documentation about the downloadable data sets may be found with the data sets themselves (http://files.pushshift.io/reddit/). Most variable names are consistent across data sources. Formatting information and variable definitions for most variables, regardless of source, may be found here: https://github.com/reddit-archive/reddit/wiki/JSON. Individuals using the Reddit API may also consult the PRAW (Python Reddit API Wrapper) package (https://praw.readthedocs.io/en/latest/getting_started/quick_start.html) for a quick start guide to using the Reddit API and example code. As mentioned above, any registered user may use the Reddit API, but she must fill out a special request and be authenticated. (This will take less than an hour.) The PRAW documentation outlines the steps for authentication.
Regardless of the source, data may be structured in multiple ways. For example, data may be downloaded by subreddit—one row per subreddit and a list of information about each subreddit such as the number of subscribers. Alternatively, and likely of more interest to social scientists, one may download the list of posts, comments, or both. A post is an original submission, while a comment is a response to a post or a response to another comment. This can create a response tree with many branches that look like this: Post Comment 1 Comment 1.1 Comment 2
Each submission (post or comment) contains a unique ID and up to two additional IDs: one that allows the researchers to link the comment to the original post and one that allows linkage to the submission to which the comment was a direct response. In our analysis of 367 subreddits originating in German-speaking countries, many posts (30.5%) had no comments, but others had many. On average, we observed 11.3 comments per post (including comments on comments). The text was, on average, 2.0 words long for posts and 38.0 words long for comments. Posts and comments can include links and images. In our analysis, 10.3% of submissions were accompanied with a link. While this information may be useful as a baseline, it will likely vary by language, country, and over time.
Posts and comments also contain user IDs, so they may be analyzed by user. Popularity information is available at the submission level, including number of views and number of up and down votes. An up vote is similar to a “like” on Facebook, and a down vote is its complement. Votes can only be viewed in the aggregate and are not linked to individual users. Other variables are also available at some levels but not others. For example, the number of subscribers is available for each subreddit, but one cannot access to which subreddits an individual subscribes.
So far, we have discussed these data as if they are clean and ready for analysis immediately after download. This is not accurate; these data are messy, regardless of their source. There are duplicates (mostly the result of individuals editing their own posts) that must be removed. Duplicates may be identified using their unique ID. Some submissions have been edited. Depending on the source of the data, one may need to decide whether to use the original or the edited submission. Finally, some posts have been removed by the user who posted it, by the subreddit moderator, or by Reddit Inc. These data may still appear in the data set, but data users will find “[DELETED]” where submission content would typically be found. In our analysis, we removed 5.5% of all submissions prior to analysis due to these events. Submission content itself is also messy. When analyzing non-English data, researchers will need to confirm that all accents downloaded and formatted correctly. UTF-8 is the most common format that should be compatible with languages that use Roman letters. Languages that use other alphabets such as Asian languages will require a different formatting scheme. There will be typos and social media language (e.g., LOL). In many cases, there will also be multiple languages used. Language may change across and within subreddits, threads, or authors.
In addition to the Reddit data themselves, other supplemental data are available. We already mentioned above that all users are anonymous, so sociodemographic information is not available. However, sites such as https://snoopsnoo.com/ analyze all posts from each user and put together a list of attributes to describe the user. More active redditors and those that post on more diverse topics have more information about them than others, but the accuracy of these data are unknown. Moreover, we have not tested the ability to automatically download and merge data from these websites nor have we found anyone else who has conducted this type of data append. In addition to the technological limitation the lack of automation may pose, there is an ethical one. Based on the Reddit privacy policy (https://www.redditinc.com/policies/privacy-policy), redditors likely expect anonymity and privacy. While all data used by SnoopSnoo was posted by the user and are publicly available, redditors likely did not consider how individual posts may be analyzed in aggregate to create a user profile. Researchers should consult their institutional review boards and laws (e.g., General Data Protection Regulation [GDPR]) before attempting to link Reddit organic data with other data sources that may threaten the anonymity of the redditor.
Reddit Surveys
In addition or as an alternative to using organic data from Reddit, researchers have the opportunity to conduct a survey on Reddit. The benefits of using Reddit as opposed to a crowdsourcing platform such as Amazon’s Mechanical Turk or an Internet panel are 2-fold. First, fielding a survey on Reddit is free, unlike other sources which could range from about US$2 to US$50 per interview, depending on the frame and incentive. Second, using Reddit provides the potential to link the survey responses to users’ Reddit content, enriching the data and increasing the number of research questions that may be addressed.
There are three ways in which to launch a survey on Reddit. The first is to post to https://www.reddit.com/r/SampleSize/, a subreddit specifically designed for surveys. While Reddit does not have a built-in survey tool, researchers may create a post that contains a link to the survey that has been programmed on another platform (e.g., SurveyMonkey). Any registered user may post to this site, and visitors (registered or not) may click on the link within the post and be directed to the survey (see Appendix B for an example). This approach must follow the rules of the subreddit (e.g., include the survey topic in the title) but is otherwise straightforward and easy to use. It requires a single post and has the potential to elicit several responses over a short period of time. Based on an analysis of the most recent posts, an average of 70 posts are made to /r/samplesize each day. Due to this volume, the life span of a survey may only be a few hours. After that time, it drifts to the bottom (or off) of the subreddit’s front page where it is unlikely to be seen (Chandra, 2018). Not all posts are survey links (e.g., some share results), and the average number of responses is not available.
The second method is to post to other subreddits. Posting to subreddits with substantive topic areas may be more efficient in getting responses from a rare population for whom a list is not available (e.g., boaters: https://www.reddit.com/r/boating/). It may also be necessary to target a given geography or language. (The samplesize subreddit is predominantly in English.) Operationally, posting to other subreddits is the same as posting to the samplesize subreddit. A registered user posts to the subreddits of her choice and waits for individuals to view, click, and respond. However, in practice, this is very different. Reddit has a list of rules designed to prevent spamming. If the researcher wants to post the same content to multiple subreddits, she may be flagged as a spammer and privileges may be revoked. Some anti-spam safeguards are also built into Reddit. For example, one may not make more than one post in a 9-min period, or 160 per day. Each subreddit also has a set of rules. For example, some subreddits do not allow posts if the user is not subscribed to that subreddit. Of 358 subreddits to which we attempted to post the survey link, 12 bounced back for this reason. Some subreddit moderators will remove off-topic posts (we witnessed this occur to others but did not experience this ourselves); some do not allow links so a post to a survey is not possible (n = 59 of 358). These rules make scaling to a large number of subreddits difficult and time-consuming. In our experience, subreddit rules prevented us from posting to 22% of targeted subreddits and posting took 5 days. (Given 9 min between posts, it should have taken less than 2 days to post to all available subreddits. It took 5 days for several reasons. We automated the posts using PRAW in Python and factored in 10 min instead of 9 min to ensure posts did not get blocked. The system crashed twice due to human error, and there was a learning curve in understanding all of the rules and adjusting the code to adhere to them. Most of the subreddits to which we could not post were only identified by attempting to post to them. Given the way in which we built our automation tool, this counted as a post attempt, and we had to wait another 10 min before attempting the next post. Example code on how to do this is in Appendix C.)
To further minimize the risk of being labeled as spam and to attempt to lend legitimacy to our project, we also conducted an experiment using prenotification e-mails to moderators of a random half of the subreddits in which we wanted to post our survey. These e-mails were autogenerated using PRAW, using our Reddit account to post to each moderator’s Reddit inbox (not a personal e-mail account). The code to generate these e-mails is in Appendix D, and the content of the e-mail is in Appendix E. Unlike posts, there were no limitations on the number of e-mails we could send. The prenotification e-mail had minimal effect. None of our posts were removed from the sites, regardless of the experimental condition. A few posts received comments, but a qualitative assessment revealed little difference in the number or content across experimental conditions. Of the 170 e-mails that were sent, 37 responded via e-mail (8 explicitly requesting that we not post to their site).
Another challenge in conducting a survey via multiple subreddits is user karma, a Reddit gamification tool. Users receive karma when their submissions are upvoted and lose it when they are downvoted. Users could see our karma (which started at 0 but was at 53 at the end of the survey) and use it to make a judgment about our legitimacy. (Based on an analysis by redditor Hilburn, the median karma is 8, and the mean is 633. Half of all karma is owned by 1% of all redditors [https://www.reddit.com/r/theydidthemath/comments/5yf8hm/request_average_karma_of_all_reddit_users/]. 3 ) Similarly, anyone can view any user’s history. Individuals viewing our survey post on one subreddit could click our username and review all submissions. They would see that the account was new and we had repeatedly posted the same survey on 279 subreddits. Researchers and social media platforms have identified this type of repeat posting behavior as indicative of spammers (Lin et al., 2013; Robertson, 2018). One may hypothesize that users may have observed this repeat posting behavior and perceived that we are spammers, though we do not have any data to test such a hypothesis.
In our experience, our 279 posts elicited 746 clicks and 75 completed interviews over 15 days, yielding a completion rate of 10%. Among the 671 individuals who launched the survey but did not complete, 509 (76%) did not get past the first screen (informed consent) and an additional 107 (16%) broke off at the first question following informed consent (Reddit username). The remaining 55 (8%) of breakoffs occurred intermittently throughout the rest of the survey. We hypothesized that individuals may be more willing to click on the link if they were on a subreddit for which the topic was relevant to the survey. Our survey included questions about general social attitudes and policies such as political ideology, attitudes toward immigration, and belief in climate change. Given our hypothesis, we expected subreddits that contained content on politics, news, and public opinion would elicit the most response. While our ability to test this hypothesis was limited given the small sample size, we failed to find any support for it. However, these data may not be indicative of others’ experiences, and researchers should use this information with caution when planning their own surveys.
A final method for conducting a survey on Reddit would be to sample redditors. Under this approach, the researcher would construct a frame of currently active redditors, draw a sample, and invite them to the survey by sending an e-mail to their Reddit inbox. In the most general approach, a frame may be constructed by downloading the most recent month’s data set from http://files.pushshift.io/reddit/ and identifying unique user IDs. Samples could be stratified based on user attributes (e.g., karma scores) or subset to users who had posted on a given topic area, and e-mails could be generated using code similar to that found in Appendix D.
In our review of the literature, we were able to identify only one example of this approach. Jhaver, Appling, Gilbert, and Bruckman (2019) sought to study how individuals reacted when their social media posts were removed from Reddit. Using PRAW, they randomly sampled a subreddit and extracted new posts. After 3 hr, they rechecked each post to identify whether it had been removed. Authors of removed posts were sent unique survey invitations to their Reddit inbox. The researchers achieved an 8.2% response rate and received several (count unknown) messages back asking clarifying questions about the invitation. They considered this to be a high response rate and contributed their success to their customized message, varying the time of day the invitation was sent to ensure adequate response from individuals in all time zones, and their prompt response to inquiries from sampled individuals. We also hypothesize that this higher-than-expected response rate may be the result of topic saliency, the topic of the survey was of unique interest to the invited individual, increasing the likelihood of response (Groves, Presser, & Dipko, 2004). While the response rate may be high relative to expectations, 91.8% of invited redditors chose not to respond. Unfortunately, we do not know why. We are unaware on how often individuals check their Reddit inbox. If their account is not linked to their personal e-mail or if they have opted out of e-mail notifications sent to their phone or personal e-mail, they may not be aware of a Reddit message unless they choose to visit their Reddit inbox. We also do not know how often individuals receive unsolicited Reddit messages or whether they consider this type of contact a violation of their privacy or “personal space.” Finally, we do not know the turnover rate among user IDs (how many were active last month that are inactive this month) or how many user IDs belong to businesses or similar entities. Given these unknowns and given that we have only one example from which to estimate response, additional research into this survey method will be necessary before determining its viability.
In addition to response rates, the authors provided information on the types of redditors who responded. Survey respondents had a median of 3,412 karma points, had been on Reddit for a median of 436.6 days, and posted a median of 35 times. Given these statistics, survey respondents were more experienced and engaged than the average redditor (comparable statistics in The Reddit Population section). Nonresponse bias analysis was not performed (or, at least, not reported) by the authors. Therefore, we cannot determine whether the difference between survey respondents and the redditors is a function of the type of individuals for whom posts are likely to be removed or in the type of individuals who are likely to respond to survey requests sent to their Reddit inboxes.
While much is still unknown about this methodology, there are several benefits. First, the approach makes it easy for researchers to link survey and social media data. To link these data, researchers will need the redditor’s user ID (see the next section for details). Under this approach, the user ID is on the frame and would not need to be requested in the survey, eliminating the risk of item nonresponse. Despite this benefit, researchers will need to consider the ethical and legal (e.g., GDPR) considerations of linking data without the individual’s explicit consent. Second, if the goal of the survey is to make inference to the Reddit population, this approach creates a probability-based design and many weighting and analysis methods used on probability surveys may be appropriate. Third, and relevant to Jhavar and colleagues' (2019) research, this approach allows researchers to identify and target invitations to rare subpopulations by identifying their eligibility through their posts. Finally, this approach may avoid the negative perception that may result from posting to multiple subreddits since messages to redditors’ inboxes are private, and users would not be able to see that multiple messages with similar content were sent.
Regardless of the method used, researchers must consider survey standards and ethical and legal regulations. Any Reddit survey should inform individuals about the survey topic, length, purpose, and sponsor. To further legitimize the survey, it should also include information on how to contact the principle investigator.
Combining Reddit Survey and Organic Data
If researchers may use Reddit’s organic data and may conduct a survey, they could also do both. This involves linking survey data to organic data via the Reddit user ID. (An alternative to micro-, user-level linkage is to conduct analysis on each data source independently and then merge the statistics from each source. This type of analysis is very common in economics research to compute indicators such as economic strength. We have not included a discussion about this approach here because it does not have any strengths or challenges unique to social media or Reddit.) Survey data can help minimize the limitations in the organic data such as the lack of covariates available and the scarcity in the amount of content on some topics. The organic data may better measure trends over time (as opposed to a single-point-in-time survey) and provide data that individuals rarely retain or ever know (e.g., enumerating their social network). Combined, the amount of information available for each individual increases, allowing for more complex, multivariate analysis. While we found limited examples of researchers conducting data linkage with Reddit data, some researchers commented on the value-add this approach could have. For example, Kilgo and her colleagues (2016) conducted an analysis of Reddit posts to determine whether Reddit opinion leaders can maintain their anonymity. In her summary, she discusses the benefits (and drawbacks) that a survey could have had in limiting the weaknesses of her study. Specifically, gaining additional information on users’ personality traits and online media usage through a survey could inform researchers on how individuals become opinion leaders and decide who to follow.
However, data linkage is not without its own challenges. First, it requires the Reddit username to be on both the survey and organic data sets. This requirement limits what organic data may be used since not all variables are available with an appended username. For example, data on voting are only available by post. Researchers may download the amount of up and down votes that a given post received, but they cannot access who voted. Similarly, researchers can identify who posts on a given subreddit, but they cannot access a complete list of subreddit membership. Individuals may post to subreddits to which they are not members and may be members to subreddits on which they do not post. Reddit voting behavior and subreddit membership may improve political scientists’ ability to predict of voting behavior in national elections, but without the ability to link these data, this information cannot be exploited.
Second, researchers must obtain the username. If the researcher has sampled individual redditors as described in the third survey method above, the username is available on her “frame.” However, if she has used one of the first two methods and posted a link to a survey on a subreddit, then she must ask the respondent for his or her username. In our experience, asking for the respondent’s username results in high breakoffs and item nonresponse. Among the 237 individuals who completed the informed consent and were directed to provide a username, 107 (45%) broke off at that request. Among the 75 individuals who ultimately completed our survey, 17 (23%) refused to fill in a username, 8 (11%) filled in a false name (e.g., “ichgebeuchnichtmeinennamen”—English translation: I am not giving you my name), and 24 (32%) provided usernames for which we did not find any submissions over the past year. The 65% item nonresponse rate occurred despite the inclusion of help text (Appendix F) that explained the importance of providing a username.
A third challenge is the lack of control over the content. Linkage can only occur for cases for which both survey and organic data are available. However, survey respondents may not have posted on the subject matter of interest or may not have posted as frequently as required for sufficient power. In our research, we were looking to identify posts on seven topics (immigration, climate change, European Union (EU), gay rights, political ideology, interest in politics, and trust in people). Two researchers manually coded the topic of 22,084 submissions (posts and comments) identified across 26 users over 1 year. Only 685 (3%) were on a topic of interest. One user posted a total of 3,465 times over the course of the year but only one post was on a topic of interest. The least number of submissions by a user was two, while the median was 103.
Summary
The plethora of social media data offer a rich opportunity for social science researchers to answer research questions in new ways. However, these data are not perfect. Accessing and analyzing these data come with their own set of challenges that need to be considered prior to using them. In some instances, it may be more expensive to use Reddit data than to conduct a traditional, probability-based survey. In other instances, the cleaning and models required to code text data may be too time-consuming. Our goal is not to convince researchers that social media data are too difficult to use, but that they should be aware of the opportunities and challenges that they may encounter to allow an informed choice.
Table 3 summarizes how the above information about conducting research using Reddit may affect design choices or may be used to determine if Reddit is an appropriate data source for any given research project. The lists of questions and considerations are a starting point. Each researcher will need to adapt and expand this list to her specific research project.
Questions to Consider in Combination With Reddit Information.
Note. API = Application Programming Interface.
Footnotes
Appendix A: Python Code to Open and Subset a Large json Data Set on a Computer With Limited Memory
Appendix B: Submission Text for Reddit Survey
Title: Bitte nehmen Sie sich 5 Minuten Zeit, um Ihre Meinung zu Reddit und wichtigen sozialen Themen zu teilen? Dies ist Teil eines Forschungsprojekts. Wir werden keine persönlichen Daten erheben und wir haben auch keine politische oder finanzielle Agenda. Content: [URL]
Title: Please take five minutes to share your opinion on Reddit and important social topics as part of a research project. We do not collect personal data and we do not have a political or financial agenda. Content: [URL]
Appendix C: Python Code to Automate Subreddit Submissions
Appendix D: Python Code to Automate Subreddit Reddit E-mails
Appendix E: Prenotification Materials for Reddit Survey
Hallo,
Mein Kollege und ich sind Sozialwissenschaftler und suchen nach alternativen Wegen für Meinungs- und Sozialforschung, u.a. mittels Reddit. Unser Ziel ist es zuverlässige Daten für die akademische Meinungs- und Sozialforschung zu sammeln. Das heißt, wir haben weder finanzielle noch politische Absichten.
Wir würden uns sehr freuen, wenn wir einmalig einen Link zu einer kurzen Umfrage (weniger als fünf Minuten) posten dürfen, die subscriber vom (subreddit name) subreddit, aber auch alle anderen Nutzer beantworten können. Da wir uns an die Regeln von Reddit und diesem subreddit halten wollen, möchten wir vorher kurz bei Ihnen als Moderator dieses subreddits nachfragen. Auch wenn unser Post vielleicht etwas off topic ist, wollen wir dennoch sicherstellen, das wir Einschätzungen von möglichst vielen verschiedenen Reddit-Nutzern bekommen – und nicht nur von Nutzern, die sich für ein bestimmtes Thema interessieren.
Falls Sie mehr über uns oder unsere Arbeit erfahren wollen, finden Sie hier mehr Infos über uns oder können uns per e-Mail erreichen.
Ashley Amaya: https://www.rti.org/expert/ashley-amaya (
Ruben Bach: http://sswml.uni-mannheim.de/Team/Ruben%20Bach/ (
Danke,
Ashley
Hello, –
My colleague and I are social science researchers who are looking into alternative ways to measure public opinion on everything from Reddit to immigration reform. Our goal is to find affordable ways to collect accurate data that can be trusted. We have no political or financial agenda.
All we’re asking for is permission to make a one-time post with a link to a short (less than 5 min) survey that [SUBREDDIT NAME] subscribers (or anyone) could take. We want to be respectful of what is acceptable practice on this subreddit and check in with you as the moderator before posting. We know this post may be a bit off topic, but we have a diverse set of questions and want to make sure we get opinions from all types of people—not just people interested in a specific topic.
If you want to learn more about us, feel free to check out our websites or e-mail us:
Ashley Amaya: https://www.rti.org/expert/ashley-amaya (
Ruben Bach: http://sswml.uni-mannheim.de/Team/Ruben%20Bach/ (
Thank you,
Ashley
Appendix F: Help Text Accompanying Survey Request for Reddit Username
Warum fragen wir Sie nach Ihrem Usernamen? Wir würden ihn gerne aus vier Gründen wissen: Damit jede Person den Fragebogen nur einmal ausfüllt. Um uns sicher sein zu können, dass sie ein registrierter Redditor sind. Um dazu beizutragen, zukünftige Fragebögen wie diesen kürzer zu machen. Um Ihre öffentlichen Reddit-Beiträge und -Kommentare mit unseren Umfragedaten zu verbinden.
Why are we asking for your username? We would like to know it for four reasons: So that each person fills in the questionnaire just once. To make sure that you are a registered redditor. To help us shorten future questionnaires. To connect your public Reddit posts and comments with the survey data.
Data Availability
Most of the statistics referenced in this article are from secondary sources. In those cases, we have cited the article and data source, when available, within the text. There are two types of data that we analyzed ourselves. All data used in analyses of Reddit organic data may be found here:
. Unfortunately, any data used in the analysis of the Reddit survey cannot be shared because it would be in violation of our confidentiality and privacy clause. In order to maximize the replicability of our findings, we have included information on the methods used to conduct the Reddit survey and have included syntax to replicate some of the procedures used to scrape and/or subset the data. Given the purpose of this article is to provide researchers with guidelines and some considerations when working with Big Data, replication of some of the descriptive statistics cited from the Reddit survey should not change the conclusion.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported, in part, by the German Research Foundation (DFG) through the Collaborative Research Center SFB 884 “Political Economy of Reforms” (Project A8) [139943784 to Annelies Blom, F. Keusch, and F. Kreuter].
