Abstract
Social media platforms have become very popular these days among individuals and organizations. On the one hand, organizations use social media as a potential tool to create awareness of their products among consumers, and on the other hand, social media data is useful to predict the national crisis, election polls, stock prediction, etc. However, nowadays, a debate is going on about the quality of data generated on social media platforms, whether it is relevant for prediction and generalization. The article discusses the relevance and quality of data obtained from social media in the context of research and development. Social media data quality issues may impact the generalizability and reproducibility of the results of the study. The paper explores possible reasons for quality issues in the data generated over social media platforms along with the suggestive measures to minimize them using the proposed social media data quality framework.
Introduction
Social media is an internet-based network platform that enables content creation like reviews, discussions and feedback by the users (Rodriguez et al., 2012). It is an important platform for interaction among users and other stakeholders. Users are spending a significant amount of time on these platforms, and several public and private organizations are engaged in social media usage and taking advantage of the user base. Examples of social media platforms include Facebook, Twitter, Wikipedia, blogs and e-commerce websites, which give the provision to customers to connect and exchange information. A massive amount of data is getting generated on social media platforms. Owing to the presence of such a large amount of data, researchers and practitioners are fairly interested to know the contribution of social media data towards academics and society. The literature suggests the use of the social media data for various research purposes like for studying behavioural aspects of the users, finance, products’ demand forecasting, health care, customer acquisition, value creation, disaster management, etc. (e.g., Chakraborty et al., 2013; Dong et al., 2018; Iftikhar & Khan, 2020; Kallinikos & Tempini, 2014; Meire et al., 2017; Pihl & Sandström, 2013; Wu & Cui, 2018). Further, nowadays, both private and public sector organizations leverage social media platforms to their benefit. Fischer (2009) argued that social media has the potential to generate participative cultures, learning and knowledge-sharing among people. People freely communicate their opinions and ideas over the groups and communities on social media platforms. E-commerce companies use these platforms as consumers have the provision for sharing information of the products over the e-commerce websites (Alalwan et al., 2017; Meire et al., 2017). Not only E-commerce companies but also government and political organizations take advantage of the social media (Dwivedi & Kapoor, 2015; Kapoor et al., 2018). In the 2008 and 2012 U.S. presidential elections, political parties used social media as part of their campaign and generated interest towards their parties among young students and adults (Adjabeng, 2014). Vakeel and Panigrahi (2018) found that social media usage has a positive impact of on e-participation and e-government development of a country. Social media is a platform where individuals can connect and exchange their thoughts and opinions about a subject or product. As per Nielsen (2009) report, nearly two-thirds of the world’s population spends almost 10 percent of their internet usage time in visiting social sites. Currently, there are around 3.6 billion social media users worldwide, and the count is continuously increasing. The use of social media platforms is a popular way to reach out to the opinions of others and share one’s own experiences. With the advent of new technologies, social media platforms are easily accessible to most individuals. Consumers use social media not only to share their personal experiences of the products but also to share information for the benefits of others. Data generated from social media is widely available and provides numerous research opportunities to the academicians and practitioners (Lewis et al., 2008). However, the huge amount of social media data comes with several data quality issues (please refer to the section ‘Analyzing Issues with Collected Social Media Data’ for a detailed discussion). In the current work, we explore various data quality issues related to social media-generated data along with an analysis of the related issues. Specifically, we argue that following the proposed well-defined measures during data collection of social media can provide a potential breakthrough in ensuring the data quality issues. Issues in data quality badly impact the generalizability of the research (Bender & Friedman, 2018). Researchers face hardships in ensuring data quality and reliability of collected data using existing social media data metrics (Zahedi & Costas, 2018). The current study is a step towards providing a standard of social media data collection in a presented checklist and social media data collection framework. There is a dearth of studies that investigate or provide remedial measures towards data quality issues of social media data (Vial, 2019). The literature suggests that big data (e.g. social media data) quality plays a vital role in an organization’s decision-making process (Ghasemaghaei & Calic, 2019; Hargittai, 2020). We explore social data quality issues through two perspectives. First, technology-related issues like handling and making sense out of the volume of data, aka big data which is getting generated on social media platforms every second. Second, user-generated content related to issues like political and religious bias, rumours, etc., on social media platforms. By doing so we assume that social media is a socio-technical platform. While using social media data, the researchers not only face the technical challenges of controlling the data (Erevelles et al., 2016; Vial, 2019) but also need to take into account the social context and biases of the data (Baeza-Yates, 2018; Hargittai, 2020) that is inbuilt into the data. Hence, it is indispensable to investigate data quality issues and their remedial measures. The rest of the paper is organized as follows: The next section talks about data types and data collection techniques in social media research. The following section analyzes data quality issues of the data collected through social media platforms. Remedial measures for social media data quality issues are discussed in the following section. The data quality framework for social media data is presented in the section after that. The conclusion of the work with limitations and future scope is produced in the final section.
Data Collection and Data Types in Social Media Research
Use of social media is a cost-effective medium to access groups of people and obtain information from them to improve upon and devise methods for making better decisions (Lee & Kim, 2011). As discussed in the previous section, a huge amount of data is generated on social media every day. Given the rising popularity of social networking sites nowadays, we argue that social media is a rich repository of data for obtaining many important insights (Still et al., 2014). There are mainly three types of data collection methods in social media research:
The literature suggests that there has been a focus on investigating social media data for demand forecasting, impact on morale enhancement, customer empowerment, intention to purchase and disaster management (Iftikhar & Khan, 2020; Johnston et al., 2013; Nisar & Whitehead, 2016; See-To & Ho, 2014; Wu & Cui, 2018). Social media data reflects the actual behavior of users rather than the declarations by the respondents (Bharati & Chaudhury, 2019; Zhang et al., 2020). Nowadays, both primary survey data and secondary data directly crawled from social networking sites have been used by the researchers. Both Chakraborty et al. (2013) and Zhang et al. (2017) studied social media users’ behavioral responses related to privacy issues of the users. Chakraborty et al. (2013) used data from Facebook pages of profiles of elderly adults to study their privacy-preserving actions. Zhang et al. (2017), on the other hand, used data from the online surveys of respondents who use social networking sites. Amazon Mechanical Turk was used to recruit respondents (social media users). Dong et al. (2018) used financial social media to detect corporate fraud detection. They argue that traditional corporate fraud detection methods rely on analyzing financial statements, reports, etc. that are available only on release and indicate the past operations of an organization. On the other hand, social media include a wide variety of stakeholders of the organization and have information related to the currentoperations of the organization. Nowadays, social media researchers/practitioners are extensively using all three mentioned types of data collection techniques to capture the behaviour of the users. The varied types of social media data and data collection techniques give rise to the debate among academia regarding data quality and research reproducibility issues in social media-based research. Subsequent sections will throw light on several such issues and will discuss potential remedies wherein we tried to explore and analyze the issues and remedies based on critical assessment of the literature of social media research.
Analyzing Issues with Collected Social Media Data
In this section, we define and discuss two categories of data quality issues related to social media data: (a) technology-related data quality issues and (b) user-generated content-related data quality issues. Nowadays, the credibility of data is the key issue prevailing in an online web environment. We define high-quality social media data as the data that is credible and accurate (Metzger & Flanagin, 2013). Before delving deep into social media data quality issues, let us discuss characteristics of social media data, which may give rise to data quality trade-offs related to social media data. Social media generates a huge amount of data every second. This data is generated at a large pace, is in various formats like structured, unstructured, etc. is termed as big data (Erevelles at al., 2016). The complex nature of big data is platform-dependent and is categorized as technology-related characteristics of social media data. The big data generated on social media is hard to handle and poses several information or data-quality issues (Singh et al., 2017). Secondly, data quality issues may arise out of user-generated content on social media. A huge amount of user-generated content may provide users with lots of valuable information; however, users find it difficult to find and take advantage of the relevant information (Kapoor et al., 2018), which could be categorized as user-generated content characteristics of social media data. In the following subsections, we will discuss social media data quality issues related to both types of social media data characteristics.
Technology-related Data Quality Issues
Pertaining to technology-related characteristics of social media data, the issues studied in past data talked about the four Vs’ of big data: (a) volume (b) velocity (c) variety and (d) veracity (Erevelles et al., 2016). The four Vs’ of big data pose limitations at data collection and pre-processing stages. They create hardships in harnessing meaningful data as it is difficult to create standard filters that can cope up with ever-increasing, multidimensional, and ever-changing data. Further, different social media platforms require different crawling algorithms to extract the data. For instance, Facebook mostly contains unstructured data. On the other hand, data from Twitter is more structured in comparison. Twitter is a micro-blogging platform and has a word limit of 280 characters. The character limitation may pose a restriction to the users in terms of posting detailed content. This limitation, however, proved to be beneficial for the quick exchange of information, which is helpful during an emergency (Wu & Cui, 2018). The issues related to big data poses hardships related to the technical side like the architecture of social media platforms and processing of social media data (Vial, 2019). Social media platforms have huge data repository, both structured and unstructured, and is rich in providing attitudinal information of the users (Meire et al., 2017).
User-generated Content-related Data Quality Issues
Other issues that may creep in with social media data pertains to its user-generated content. These types of data issues are very subjective and vary from user-to-user and context-to-context. User-generated content are influenced by a wide variety of reasons, which limits the authenticity and reliability of the social media data. The reasons for influence may include users’ political bias, peer influence, social influence and so on (e.g., Baeza-Yates, 2018; Chou & Edge, 2012; Georgarakos et al., 2014; Hargittai, 2020; Wilcox & Stephen, 2012). Further, content on social media could be sponsored or commercial posts, which again is a threat to the validity of the content. Nowadays, several organizations are making use of social bots to post contents for popularising their products and ideologies (e.g., Boshmaf et al., 2011; Elyashar et al., 2013; Erevelles et al., 2016; Wagner et al., 2012), which in turn pose hardships to arrive on the true and unbiased picture of the actual information. The current body of research also continues to discover the limitations and possibilities of exploring social media data at its best. Dong et al. (2018) analyzed financial social media platforms ‘Yahoo Finance’ and ‘Seeking Alpha’, and they argued that social media data might contain rumors and irrelevant information. The authors developed a rumor detection model for identifying the rumors and irrelevant information in the mentioned social media platforms. Further, a quick change of nature of social media data (e.g., Twitter) poses limitations on the generalizability of results; for example, if the study is done on data collected during a particular time period or event (e.g., Hopp & Vargo, 2017). Given the rising popularity of social media, the volume of social media data is expanding at a large pace, which results in information overload on the users (Singh et al., 2017; Zha et al., 2018). The information overload may also pose hardships to the researchers to identify the relevant data for their studies. Also, nowadays, social media users are concerned with various privacy invasions in online social networks, which lead to leaking of personal information, online harassment such as stalking and sexual abuse, misrepresentation, public embarrassment, hate speech, racial bias, etc. (Al-Qurishi et al., 2018; Bender & Friedman, 2018; Choi et al., 2015; Jiang et al., 2013; Sap et al., 2019). This demotivates users from sharing true information on social media. Subsequently, it impacts the reliability of the contents posted on social media (Chakraborty et al., 2013; Feng & Xie, 2014; Jung, 2017).
In line with the aforementioned discussion and extant literature, we categorize quality issues of social media data as summarised in Table 1.
Social Media Data Characteristics and Challenges
Raising an important and somewhat under-researched issue of replicability/reproducibility in information systems (IS) research, Marsden and Pingry (2018) argued that data quality and transparent data collection are the important factors to ensure reproducibility of research. The authors discussed different data collection types prevailing in IS research literature. They emphasized that controlled laboratory experiments are the best way to ensure high data quality standards and hence to ensuring research reproducibility. However, manipulating experimental conditions in social media data is difficult (Choi et al., 2015), which limits perfect laboratory experiments in case of social media. Consistent with Marsden and Pingry (2018), several researchers emphasize data used in research to be one of the important factors for research reproducibility. For example, Buckheit and Donoho (1995) and Fomel and Claerbout (2009), as cited in Wandell et al. (2015, p. 1) define reproducible research as ‘the idea that the product of scientific research is not only the paper but also the data and software needed to reproduce the results’. Hence, high data quality is deemed to be a foremost criterion for research to be reproducible. In the subsequent section, we discuss the reproducibility of research in the context of social media data quality and propose a data quality framework in line with the extant literature.
Discussion on Remedy, Reproducibility, and Generalizability in Social Media Research
Social media data is a rich and open source of information content. Holistic predictive models can be built using social media data along with operational data that augments the scope and performance of predictive models. However, certain limitations are always there irrespective of any data source under consideration. Cross-sectional data collection poses issues with the generalizability and reproducibility of the research. The time period of social media data collection is also an important aspect that impacts the results. The accuracy of the prediction can be improved by wisely choosing the data collection time period. Stieglitz et al. (2020) found that data collected from social media before the event have more predictive power than the data collected after the occurrence of the event. Most of the researchers who studied social media mostly use Twitter and Facebook data for data collection purposes. Twitter, due to the character limitations on its post, is more structured in comparison to Facebook data. Facebook data is more complex but has high-quality information. Meire et al. (2017) studied prospect identification in the B2B business scenario. The authors studied three types of data: commercially purchased data, data from websites, and data from social media. Data collection from social media was done through the Facebook platform. The authors conclude that social media data (Facebook) is the most informative data source out of the three data sources used for the study. Further, we observe the diverse nature of the dataset on social media networks ensures that the collected samples from social media are a close representation of the population (Casler et al., 2013). This might not be the case if data is collected by posting online surveys (to the respondents who use social media), which lacks data quality (Marsden & Pingry, 2018). We argue that on social media, comments from users reflect their actual behaviour naturally. In contrast, in the case of responding to the survey, users might be attentive and an element of bias may creep in the responses. We, therefore, can safely assume that social media data contains unbiased data (in terms of intrinsic behaviour of the user), which is also supported by the studies (e.g., Bharati & Chaudhury, 2019; Meire et al., 2017; Zhang et al., 2020). However, a recent article on a survey of the social media analytics literature revealed that social media data is a source of high-quality information. Still, the information is not present in a filtered form (Ghani et al., 2019). We suggest that a standard operating procedure should be followed uniformly across all the social media research streams, both at the time of extracting social media data and while preparing the same for the study. Also, as mentioned, social media users generate a huge amount of data; this ease of generating the data attracts lots of malicious and irrelevant information, which contaminates the useful information (Bindu & Thilagam, 2016). The irrelevant and malicious information is termed as noise, while the relevant and useful information is known as signals. To enhance the quality of social media data, a data filtering scheme can be introduced in the data preparation stage to filter out the noise (Stieglitz et al., 2018). As a standard practice, the data filtering scheme should be documented by the researchers so that it can be shared further with academicians/practitioners to ensure the data quality and reproducibility of the studies. Further, as social media represents a socio-technical system (Boyd, 2015), we urge while studying/analyzing social media data, researchers/practitioners should approach data collection procedures with any of the social theories as to the basis. Many studies have given due consideration to this practice (e.g., Bharati & Chaudhury, 2019; Chakraborty et al., 2013; Church et al., 2017; Mohamed & Ahmad, 2012). The social theories that form the basis of the mentioned studies are social capital theory, social role theory, social exchange theory, and social cognitive theory, respectively. The mentioned social theories list is not an exhaustive one but gives an idea about the trend of social theories being adopted by the researchers who use social media data for their studies. We suggest that social media data collection on the lines of relevant social theory will improve social media data quality as it will give the direction to the researchers to fetch the relevant data at the time of data collection. This ensures that if the replication study also follows the same theory as the basis for data collection, then there are high chances of reproducibility of the results. Also, as we argued that social media possess socio-technical characteristics, improving on the social/relational side, social media users should be encouraged to generate high quality of data (Zha et al., 2018). Further, as discussed in the previous section, there is a huge amount of user-generated content that is increasing day-by-day on social media, and subsequently there us an increasing information overload on users/researchers; improving on technical/information technology (IT) platforms side, we suggest that certain standard social media information filters could be designed to filter noise or low-quality data in the data extraction stage (Stieglitz et al., 2018; Zha et al., 2018). Finally, we propose the guidelines based on one HOW and six WHYs of data collection (Marsden & Pingry, 2018) to ensure the quality and usefulness of social media data. Table 2 representsthe guidelines/questions that researchers may ask herself/himself during pre, pos, or during data collection from social media platforms. We will further map each parameter listed in Table 2 with individual characteristics of social media data. Further, we will arrive at a framework that can be used in conjunction with the checklist to ensure social media data quality issues.
Checklist During Social Media Data Collection Lifecycle
Based on the checklist obtained from the seven parameters and questions (refer to Table 2), we propose a framework containing innate features of the data during data-collection and data-processing lifecycle. ‘What’ parameter tells about ‘content’ of the data like whether data is collected from school students, from elderly people, from an organization’s employee, etc. ‘When’ and ‘Where’ parameters could be attributed to the spatial-temporal characteristics of the data, say ‘When’ parameter could be mapped to the ‘temporal’ aspect of the data like when the data is collected like whether data is collected during some election campaign, during some festive season and so on. Different timings of data collection may induce different types of biases in the data. Likewise, ‘Where’ parameter could be mapped to the cultural dimensions of the social media users from whom the data has been collected. ‘Who’ and ‘Which’ parameters could be attributed to the ‘access control’ like ‘Which’ data is available and ‘Who’ is authorized to access the data. Last but not least, the ‘Why’ parameter gives an idea about the objective of the data collection. The objective of data collection should be crisp and clear, which will further enhance the research reproducibility provided the future studies also have similar or related objectives. Further, we propose that the ‘How’ parameter resides at the top level throughout the entire lifecycle of data collection and should take into consideration the steps to be followed while following the steps related to the remaining parameters. Based on the checklist (refer to Table 2), we arrive at a social media data quality framework that can be utilized by the researchers or IT managers of the firms. The academicians and managers can introduce the framework as a formal standard operating procedure to be used for ensuring social media data quality. Ensuring social media data quality issues put forward huge implications and research opportunities to the academicians and practitioners fraternity. The recent literature suggests that although the availability of huge amount of data provides everlasting opportunities to the researchers, there is always a severe issue associated with the unverified open datasets (Alsudais, 2021). The presented checklist and social media data quality framework addresses the issues of reliability and quality of the datasets by suggesting an approach at the very initial level of data collection itself. The current work suggests the data quality checks and framework that handle uncontrollable social media data (big data). Apart from addressing technical challenges of the social media data, the presented checklist and framework also address the potential human and contextual bias in the data.
Data Quality Framework for Social Media Data
Based on the earlier discussion, we propose and explain our data quality framework for social media data in reference to the data collection process over social media (e.g. Twitter) during the spread of the coronavirus disease (COVID-19) pandemic, which could be helpful to study the pandemic disaster management capability of Twitter. The overall structure of the proposed framework is given in Figure 1. The objective of the data collection is to identify the location of COVID-19 victims in different countries of the World. Location identification is done based on the tweets posted by the Twitter-users. Further, we expect that through data collection from social media, the researchers can follow and document the various steps that can be used to ensure the data quality and research reproducibility. This will benefit future researchers who may be re-using the data to investigate different phenomena.
As discussed in section 2, the social media data can be retrieved in three ways, that is, primary data collection, data collection from social media-specific web pages and user-generated content posted on social media websites (e.g., Twitter, Facebook, etc.). To demonstrate the use of the social media data quality framework represented by Figure 1, we take example of social media data collection from Twitter. In the mentioned context, our social media data quality framework can be utilized as demonstrated further.

The step ‘How’ is the overall data collection process that consists of ‘What’, ‘When’, ‘Where’, ‘Who’, ‘Which’ and ‘Why’. Figures 2, 3 and 4 represent three sample tweets accessed via the Twitter website, which helps in explaining our steps of social media data collection framework. The sample tweets are from three different locations across the world, namely India, the United States and Europe, respectively. Let us assume that one has to collect COVID-19 related data from Twitter. Now, going by the mentioned social media data quality framework, ‘What’ consists of the tweets with the keywords ‘Covid’, ‘Corona’ or any other adjectives/words related to the; COVID-19’ pandemic (refer to blue-highlighted rectangles in Figures 2, 3 and 4).



Further, ‘When’ and ‘Where’ refer to the spatial-temporal characteristics of the collected data. In Figures 2, 3 and 4, the location of the collected tweets are from India (e.g.: Uttar Pradesh), the United States (New York) and Europe, respectively. Let us say that time period of the data collection is selected from ‘Jan-2021’ to ‘March-2021’, that is, for a period of three months (refer to the section enclosed in green-highlighted rectangles in Figures 2, 3, and 4 referring to the location and time-frame). ‘Who’ and ‘Which’ refers to the access control, that is, the origin of the source of data. Let us assume that in our case, we are interested in filtering the tweets based on the source. For example, we are interested in collecting the tweets that originate from hand-held devices like android phone, iPhones etc. (refer to the section enclosed in red-highlighted rectangles in Figures 2, 3, and 4). Further, we assume that data (tweets) originated from hand-held devices provide a more dynamic and real picture of the citizens suffering from the pandemic. We discussed and explained the components of the mentioned social media data quality framework with the help of an example of data collection process from Twitter. On the same lines, various stages of data collection from a social media platform can be mapped to the components of the social media data quality framework, which in turn will perform and ensure the data quality checks.
Conclusion
In the present work, we tried to touch upon the data quality and research reproducibility issues in the context of social media data-based research. As already mentioned, social media data is rich in information, but it also possesses much noise like irrelevant and biased information; it is, therefore, necessary to explore the issues surrounding social media data quality. We discussed the increasing trend of studying social media data citing the relevant literature, followed by a description of methods and types of social media data collection referred by various studies. The paper also provided a detailed explanation of potential data quality issues of social media data, which may result in the hampering of the reproducibility of social media research. From the past research on social media data, we find that social media data issues can be attributed to the unique characteristics of social media data. We propose two broad categories of data quality issues based on social media data characteristics. One class of data quality issues could be attributed to big data characteristics (technology-related issues) of social media data, while the other class of data quality issues could be attributed to the content created by a large and diversified set of users. Social media data quality may hamper due to technical challenges to handle volume and variety of data. Further, content created by social media users may be biased due to a variety of reasons; for example., social media content may contain rumors, the content may be influenced by users’ personal bias, the content may be sponsored for the benefit of some person or organization and so on. We also discussed the potential ways out to avoid/reduce threats related to social media data quality. Based on Marsden and Pingry (2018) we proposed a checklist of questions that serves as a guideline during social media data collection. Further, we proposed a data quality framework for social media data, which could be properly documented by the researchers during the data-collection process to ensure the quality of the collected data. The theoretical contribution of the study is that following the checklists and data quality framework, future researchers can make use of data from existing studies in its entirety. We know that existing data can be used to study various phenomena, and hence the researchers would only need to concentrate on identifying and working towards new ideas without being concerned too much about the availability of the data. This will decrease time, effort, and cost to a greater extent in various industrial projects. Since data plays an important role in enhancing the chances of reproducibility of research, implementing methods to improve the social media data quality will improve the chances of research reproducibility. Further, there is no study without limitations, and ours has some as well. We demonstrated the functionality of our social media data quality framework with the help of only one social media platform, that is, Twitter. Future works can try fitting our proposed framework over different platforms. Also, future studies could compare the efficiency of our proposed framework across various social media platforms. Last but not least, we did not empirically test our proposed framework. This would be interesting to investigate, and we suggest the same as future scope of the current work.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
