Abstract
In this article, we analyze the rise of data visualization in social and political contexts. Against the background of the COVID-19 pandemic, we consider a case in Shenzhen, China, and demonstrate the impact of visualization as intermediation for data on policies and society. We propose an epistemology of visualization as infrastructure. In China, visualization has been developed as a technological system to support information dissemination and collective action in public crises. Three characteristics are proposed to describe China’s data visualization politics during the pandemic, in particular, to summarize the power relationships between visual brokers, policymakers, and public users in data visualization. Further, based on the conceptualization of the links between visualization, infrastructure, and data politics, visualization serves as an important component of information infrastructure, and is an easily overlooked technical process. The creation and use of data may seem like normalized actions, but data visualization in fact supports the perception, participation, proposal, and critique of data politics issues. Via a focus group discussion with members of the Shenzhen public health sector and in-depth interviews with Shenzhen residents, we developed a qualitative effectiveness measurement framework highlighting the potential for the involvement of visualization in politics and activism.
Keywords
Introduction
Recent decades have witnessed the exponential growth of the amount of available data, and people are attracted to collect, retrieve, and analyze them (Breiter and Hepp, 2018; Cukier and Mayer-Schoenberger, 2014). However, datasets without interpretation may be difficult for the audience to understand (Gray et al., 2015). As Hilbert and López (2011) pointed out, data may be much more quickly produced and stored than they are analyzed. Therefore, data visualization as a means of efficient analysis and display has attracted widespread attention, and can be defined as a structured presentation of data using various graphical formats to promote findings and understanding (Foucault and Meirelles, 2015; Kennedy, 2015). These graphical formats include bar charts, line charts, scatter charts, maps, visualization platforms/systems, etc. (Aparicio and Costa, 2015). Data visualization is used by researchers, and is also used for the public dissemination of information (Bucchi and Saracino, 2016; Kennedy and Allen, 2016), such as in newsrooms, governments, and non-governmental organizations (NGOs). Professional practices such as information design, and data journalism have made data visualization increasingly popular and attractive (Knight, 2015).
Data visualization operates in the real world where politics and power work (Nærland, 2020). For instance, one of the purposes for which data are visualized is to detect trends or outliers (Foucault and Meirelles, 2015), which in turn leads to business or government decisions (Ruppert et al., 2015). Several studies have pointed out the significance of Open Government Data (OGD) (Graves and Hendler, 2013; Hilbert and López, 2011), claiming governments are not only one of the providers of public data, but should also be producers and advocates of data visualization to support internal evidence-based policy decisions (Robinson, 2016), as well as audience-oriented persuasion (Emerson, 2008). The importance of data visualization in political issues has evidently attracted attention. Our initial research interest was in how visualization plays a role in political issues by mediating the relationship between data and the real world.
Ruppert et al.'s (2017) elaboration on “data politics” offers a promising path. They conceptualized data politics as being “concerned with political struggles around data collection and its deployments, and how data is generative of new forms of power relations and politics at different and inter-connected scales.” The authors proposed to shift the focus from datafication to social practices and agents, claiming that data can be deployed as an object of knowledge generated through structured fields. Agents in various fields and their interests generate expertise, concepts, and methods, which are then used in the context of power and knowledge. This article investigates an emerging element of the datafication of society, namely data visualization. We claim that data visualization mediates the relationship between data and the real world, employs visuals to recode the power and knowledge fields of data, and may participate in the " [reconfiguration of] relationships between states and citizens” (Ruppert et al., 2017).
This work analyzes data visualization during the COVID-19 pandemic to shed light on the emergence of data visualization politics. As one of the most crucial public health emergencies faced by the world, the COVID-19 pandemic has elevated the need for information, as well as its visualized demonstration. To meet the public demand for information, many countries and regions have permitted the collection and analysis of epidemiological data (Xu et al., 2020). As such, COVID-19-based data visualization is presented in a variety of forms to describe the overall situation of a political or geographic area and to generate public health advice and policy. In this article, we present the data visualization practices in Shenzhen, China, which has a population of 17.56 million, as a case study. By 31 August 2021, Shenzhen had a cumulative total of 431 confirmed local COVID-19 cases and 143 imported cases. The city’s public health sector was acknowledged by other cities for its excellent performance and ongoing creation of COVID-19 data visualizations (Southern Metropolis Daily, 2020; Zou et al., 2020). In this article, our understanding of “politics” includes both narrow and broad dimensions, starting with the understanding that data visualization is relevant to the operations of government organizations and policy decisions. From a broader perspective, data visualization may be a component of a wide range of resources, power distribution, and ideologies. We pose the following two main research questions: Taking the COVID-19 pandemic as an example, how did data visualization work in China during the public health emergency? What does the new definition of the effectiveness of data visualization from the perspective of politics and activism include?
The remainder of this article is structured as follows. First, discussions on the political significance and effectiveness of data visualization spanning the literature of the politics and cognition of data visualization are reviewed to clarify the theoretical location of this research. Second, a case of how data visualization mobilized collectives during the COVID-19 pandemic is presented. Third, we provide insights into the practice of visualization as an infrastructure for Chinese society, analyze the characteristics of China’s data visualization politics, and then go beyond the COVID-19 case to conceptualize “visualization as infrastructure.” We then highlight the theoretical implications of the proposed analytical structure for data politics (Ruppert et al., 2017). By conducting a focus group discussion (FGD) with three visual brokers (a department director/manager, a designer, and an editor, respectively), and in-depth interviews (IDIs) with 17 Shenzhen residents, we discuss the factors of engagement and the definitions of the effectiveness of data visualization, and provide a preliminary qualitative framework that furthers the understanding of visualization as a toolkit of political involvement and activism in the context of a public emergency.
Literature review and research questions
This article investigates big data in relation to the field of power and knowledge, that is, what people can do with data and who is using data, rather than on the scale and complexity of information (Beraldo and Milan, 2019). Therefore, in the literature review, the focus is placed on the function of data visualization in the policymaking process and the significance of assessing the effectiveness of visualization in terms of persuasiveness.
Data politics and policy decisions
Many scholars agree that data have the potential to aid decision-making and enhance accountability (e.g. Aparicio and Costa, 2015; Park and Gil-Garcia, 2017). However, who can collect, analyze, and present data is a political issue. As Bourdieu (1993) pointed out, data has a performative power to represent political life, and Ruppert et al. (2017) proposed the concept of data politics that focuses on how data becomes an object of power, and how power critically interferes with data deployment as an object of knowledge. This concept is based on the context in which datasets are not easily understood by public users, and in which data professionals act as gatekeepers and custodians, thereby shaping competitive struggles not only for the definition of social facts, but also for principles regarding how to understand and intervene in data politics, ultimately allowing data to constitute power from fields of knowledge.
The growth of the information society has brought about a staggering amount and scope of available data. The use of data by governments or platforms engenders the emergence of new subjects. Data such as social accounts, geolocation tracking, and health monitoring can be collected to identify individuals. However, the inertia, atomization, and immediacy (Lake, 2017) of data politics are revealed when data accumulate and conflicts of rights emerge. Balibar (2012) and other scholars promoted a theorization of the collective political subject, for example, using the presence and actions of the public to coordinate the data politics brought about by platforms or governments (Ruppert and Isin, 2020; Gray et al., 2018). Efforts to translate data into action may be complex, and may bring about ways to actively engage with data (proactive data activism) and strategies to resist massive data collection (reactive data activism) (Milan and Van der Velden, 2016; Gutiérrez, 2018). Beraldo and Milan (2019) pointed out that data activism supports and illustrates an emerging “contentious data politics” in which data is understood as a field and tool of political contention.
Visualization is a way of communicating data. A large body of literature has noted the abilities of visualization to improve cognition and solve problems (e.g. Tufte and Graves-Morris, 1983; Card, 1999: xiii). Several studies have reported the application of data visualization in the political field. In general, the relationship between data visualization and policy decisions can be divided into three levels. First, data visualization itself is an act of public policy. Increasing numbers of government agencies have promised to share data in the OGD framework (Dawes and Helbig, 2010), and data availability is the prerequisite of visualization. Second, data visualization provides evidence for knowledge discovery prior to policymaking. Beginning with Snow’s (1854) Broad Street Pump Map and Nightingale’s (1858) Rose Map, data visualization has been believed to have the potential for knowledge discovery (Grinstein and Wierse, 2002; Drucker, 2010) and reducing uncertainty among decision-makers about best practices (Robinson, 2016). Although policy formulation is often considered to be nonlinear (Shaxson, 2005), i.e. based on a broader set of information, it is acknowledged that data visualization provides support for policy formation or reform, at least for policy advocates (NGOs, journalists, researchers, etc.), and can provide “possible solutions to bridge knowledge gaps between stakeholders involved in the policymaking process” (Ruppert et al., 2015). Third, data visualization, to a large extent, is used for policy persuasion and mobilization. Policy advocates transfer knowledge and evidence to help policymakers (Williamson, 2016) spread new ideas and norms, and to make social deployments; during this process, the recognized functions of data visualization are subsequently manifested, for example, reducing the difficulty of understanding (Kellehe and Wagener, 2011), speeding up understanding and saving time (Chen et al., 2014), improving the accuracy of information (Otten et al., 2015), and providing open-ended means of discovery and exploration for the people of interest (Foucault and Meirelles, 2015). Data visualization has the potential to deploy effective communication and persuasive advocacy campaigns to facilitate policy implementation (Hagen et al., 2019). From the preceding review of related studies, the chain of the relationship between data visualization and policymaking can be identified as follows: knowledge discovery → policy-shaping → policy persuasion and mobilization. This chain involves three groups, namely policy advocates (researchers, experts, NGOs, journalists, etc.), policy makers (government officials), and public users (residents, etc.).
The studies reviewed previously have helped identify the role of data and their visualization applications in society, and motivated our interest in the investigation of how visualization has played a role in the broader distribution of power and knowledge based on data politics. The data and visualization work that spurred during the COVID-19 pandemic provides context for our analysis. More importantly, the persistence of the pandemic highlights critical issues that focus on the “everyday” practice of “living with data” (Kennedy, 2018), e.g. how emerging data visualization practices support the action of collectives, and how the visualized body becomes a site of data politics. Implicit in these issues is the question of how social groups shape power relations regarding data visualization and policy decisions. Thus, the first research question is posed. RQ1: Taking the COVID-19 pandemic as an example, how did data visualization work in China during the public health emergency?
The persuasiveness and effectiveness of data visualization
“A picture is worth a thousand words.” Although the presence of data itself brings persuasiveness (Pandey et al., 2014), some empirical studies on human–computer interaction (HCI) have demonstrated that texts with data visualization may be more persuasive than those without data visualization (Pandey et al., 2014; Hullman and Diakopoulos, 2011); texts with data visualization can provide insights to the general public (non-experts) who do not have access to conclusions from complex datasets, and can give credence to the stories they tell. As Williamson (2016) pointed out, data visualization amplifies the rhetoric of data, thereby allowing it to generate explanations about the real world. Therefore, visualizations are hardly neutral, and can be accompanied by implicit or explicit standpoints. Allen (2018) noted the knowledge and power agency issues of visual brokers, and stated that the selection and creation of visualizations can be subjective. Simultaneously, their credibility can influence public engagement. Given this characteristic, some activism studies have advocated the importance of the persuasive power of data visualization (Rall et al., 2016; Boy et al., 2017) by using “emotionally powerful, morally compelling, and rationally undeniable” data to enhance advocacy work (Tactical Tech Collective, 2013). On the other hand, some critical analyses have pointed out the ethical issues behind data visualization, and have advocated for increasing awareness of the effect of design features and the biases behind visualizations (Beer, 2013: 118–119; Correll, 2019; Gough et al., 2014; D'Ignazio and Klein, 2016).
Effectiveness and persuasion are closely related. Persuasion is one of the main goals of effectiveness, which points to a more important ontological question about why we need data visualization, namely how the work of data visualization has significance and how its impact is measured. Many studies have focused on the user’s interaction with the visualization, that is, the measurement of specific elements of the engagement process, such as memorability (Bateman et al., 2010; Borkin et al., 2013; Huang et al., 2009), aesthetics (Cawthon and Moere, 2007), consistency of comprehension (Haroz and Whitney, 2012; Otten et al., 2015), and efficiency of comprehension (Chen et al., 2014; Kelleher and Wagener, 2011; Mason and Azzam, 2019; Zhu, 2007). Based on the definitions provided by previous studies, some measures of effectiveness have been discussed, for example, the common or unique visualization types (Borkin et al., 2013) and the addition or removal of visual embellishments (Bateman et al., 2010). However, Kennedy et al. (2016b) criticized this research trend, arguing that the definition of effectiveness provided in these studies is narrow, and suggested the inclusion of factors beyond data visualization. After re-distinguishing possible definitions of data visualization effectiveness, they proposed to focus on the third aspect that affects effectiveness, namely engagement, which includes six factors: subject matter, beliefs and opinions, emotions, confidence and skills, source/media location, and time; they claimed that these factors influence engagement, and suggested new definitions of effectiveness.
The persuasiveness and effectiveness traits of data visualization highlight its potential political issues. The second interest of the present research is how visualization helps public users make sense of data, and how it supports the perception of and activism towards datafication. Indeed, the work of Kennedy et al. (2016b) provided inspiration for the present study. However, the existing definitions of effectiveness, including the factors of engagement, focus on cognitive and design science. There remains room to develop a framework to outline the indicators of the effectiveness of data visualization and ultimately reveal which properties are valid and appropriate in the social and political contexts. Therefore, the second research question is posed. RQ2: What does the new definition of the effectiveness of data visualization from the perspective of politics and activism include?
Methodology
To understand the function of visualization from the social and political perspectives, and to explain its role during the public emergency, a qualitative mixed research approach was used. For RQ1, the city of Shenzhen, China, was chosen for a case study. In the early stage of the COVID-19 outbreak in China (January–June 2020), the public health sector of Shenzhen used data visualization to interpret information in a timely manner, which achieved prominent social feedback both in Shenzhen and the entire country. The practices in Shenzhen during the COVID-19 pandemic were quickly popularized in other Chinese cities; thus, this case provides homogeneous information that also applies to the whole of China. We also examined quantitative and qualitative data, including public health announcements and policies of the Shenzhen government, official technical documents involving data visualization, and national news reports during the COVID-19 pandemic.
For RQ2, the primary data collection methods were a focus group discussion (FGD) and semi-structured in-depth interviews (IDIs). Unlike audience studies based on quantitative methods, FGDs and IDIs allow for the acquisition of a large amount of data on users' attitudes, feelings, and beliefs (Gibbs, 1997), which allows for the distillation of the abstract socio-political factors associated with the effects of data visualization. Drawing on Gregg’s (2015) terminology, the groups involved in data visualization are divided into those “backstage” and “onstage.” The backstage is for visual brokers (Allen, 2018), referring to how they manage, analyze, and give rules and meaning to data visualization in the invisible production process. The onstage is aimed at public users, focusing on how they view and debate data visualizations. This sets the basis for the later discussion on visualization and data politics. Regarding the backstage group, visual brokers from the Shenzhen public health sector, including a department director/manager, a designer, and an editor, were evaluated via an FGD; they were inquired about how they, as a team, use data visualization as a mediator to engage with their audiences, and about their experiences and ideas in creating data visualizations. The FGD lasted three hours.
For the onstage group, referring to public users, 17 IDIs were conducted. Regarding the selection of interviewees, the sampling was targeted, and recruitment was conducted with reference to the gender, age (between 20 and 57), and education structure of the Shenzhen population to conduct rationing. Interviews continued until data saturation (Francis et al., 2010). Moreover, before the formal interviews began, discussions were held with data journalists and public health experts to identify six types of data visualization that were frequently adopted during the COVID-19 pandemic as stimulus material (see Figure 1 Six types of data visualization frequently adopted in China during the COVID-19 pandemic.
Findings
Before answering the research questions, this section summarizes how China’s public health sectors announced the pandemic and its management, during which data visualization acted as an important political toolkit, and how the Chinese population accessed and understood different layers of data visualization.
Developing visualization for the pandemic
The COVID-19 pandemic has encouraged innovative data practices in China. Beginning on 20 January 2020, after nearly 300 new COVID-19 cases had been confirmed in China (National Health Commission of the PRC, 2020), the public health sectors in Chinese cities released public announcements with epidemiological data on COVID-19 patients, including their gender, age, hometown, locations before and after infection, symptoms, and the time of onset of COVID-19. These data were anonymized. For example, on January 20, the Guangdong Provincial Government made the following announcement with regard to its first confirmed case: A 66-year-old male, currently living in Shenzhen, went to Wuhan to visit relatives on December 29, 2019; developed fever and fatigue on January 3, 2020; returned to Shenzhen on January 4 for medical consultation; and was transferred to Shenzhen designated hospital for isolation on January 11.
Such announcements were usually investigated and written by government public health officials and distributed through traditional and digital media channels, then organized into datasets, which were available for further use by the media, researchers, and public users. In February 2020, public panic over COVID-19 prompted the government to further refine data granularity. By 12 February 2020, 77.5% of all Chinese cities were reporting the epidemiological data of new confirmed cases (NanDu Institute of Big Data Research, 2020), including whom the infected people had been in contact with before and after infection, the places they had visited, and the transportation they had used.
Textual descriptions and datasets may not be the most effective presentation method for a polity that wants its residents to quickly understand the situation and comply with policies. Thus, since February 2020, developed Chinese cities, represented by Shenzhen, have pioneered the practice of visualizing epidemiological data. In addition to local governments, the public health sector directly under the control of the central government (the National Health Commission) has led visualization efforts by presenting the national trend and the status of each province. Further, via the use of data collected and shared by the government, scholars, NGOs, newsrooms, and citizen data scientists have been encouraged to join the visualization effort- to act as brokers and advocates of data to heighten the social awareness of the public health emergency. They have continued to analyze epidemiological data and create various visualizations to provide evidence, and even solutions, to policymakers; thus, different visualization works may have originated from different producers, resulting in a hybrid system of data and its agents. Of the six main types of visualizations presented in Figure 1, the production permits for (A) and (B) are issued by the government, as the data are linked with personal IDs; and the output for (A) and (B) is smartphone-accessible only by the specific personal ID holder to protect privacy. In contrast, data visualization types (C) to (F), which may originate from the government or are produced by other actors, are produced to be compatible with television, newspapers, or digital media channels. In particular, location maps (F) are the result of a collaboration between platform enterprises, citizen data scientists, and government sectors, and are often embedded into other Chinese mega-apps (e.g. WeChat, Alipay) as a feature.
China is becoming a highly digitized society with over 1 billion Internet users, 99.6% of whom use smartphones to access the Internet (CNNIC, 2021), and most real world activities (e.g. shopping, public transportation) in China today are often linked to a smartphone. Access to pandemic-related visualizations is omnipresent and necessary for all Chinese populations, except for children and the elderly, and is not limited to those in developed Chinese cities such as Shenzhen. Mobile versions of data visualization represented by health QR codes and location maps were found to accumulate tens of billions of mobile page views within two months after the outbreak (China Weekly, 2020; CCTV, 2020). Although the combination of visualization efforts with government public health policies implies that the Chinese population has had no choice but to comply, social surveys at the time also attest to fairly high acceptance rates in China (Zhao, 2020). Echoing Ruppert et al.’s (2017) discussion of the “world” condition in data politics, data visualization completes the connection between the virtual (data collection and analysis) and the real (residents’ actions during the pandemic).
In February 2020, more than one month after the fight against the COVID-19 pandemic began in China, different visualization works were given the opportunity to interconnect. A visualization system was presented in a four-stage model that covered geographic areas from the street level to the whole country, thereby playing a pivotal role in guiding and conversing with the population (Figure 2). This system was first implemented in developed Chinese cities including Shenzhen and Hangzhou, and was applied in cities at all levels across the country around March 2020. The four-stage visualization system used in China during the COVID-19 pandemic.
In the first stage, namely, the individual stage, when an infected person was detected, his or her trajectory data were collected and used to produce geographic visualizations, and residents near the infected person’s location were warned when they visited the location map. In the second stage, namely the region/community stage, the presence of an infected person changed the assessment and regulation of people in the same region/community. In combination with the geolocation data provided by GPS and network carriers, a non-transparent algorithm from the government measured the spatio-temporal relationship of each resident to the infected person, and generated a visualization result after a semi-automated data analysis. The analysis yields a green (low-risk), yellow (medium-risk), or red (high-risk) health QR code. The risk level determined a person’s freedom of movement. In the third stage, namely the city stage, changes in the risk level of the region/community further influenced the city’s risk level, which was also visualized with a color-coded result (green, yellow, or red) based on a non-transparent algorithm. In the last stage, namely the country stage, the numbers of infected persons in all cities were tabulated and made into a common bar chart or line chart based on a timeline, thereby presenting the overall status and trends of the country in a general and abstract manner. Of these, Stage 2 and Stage 3 were part of the quarantine and isolation policies, and also represented the results of the policies. When an individual’s health QR code or city’s risk level was green, only the recommended policies, such as maintaining social distance, wearing a mask, or going out less, needed to be followed. However, when the visualization result became yellow or red, mandatory policies were implemented, such as 14 days of isolation.
Visualization as infrastructure
Based on the review of the early stage of China’s fight against the COVID-19 pandemic (January-June 2020), we claim that, in this case, visualization practice formed an infrastructure of Chinese society. While elements such as cables and hardware, satellites, microchips, data center facilities, and semiconductors constitute the information infrastructure of the Internet itself, this simultaneously encourages the fragmentation of other infrastructures (Plantin et al., 2018; Van Dijck, 2021). We take inspiration from scholars focusing on “infrastructure studies” (Edwards et al., 2007; Star, 1999; Star and Bowker, 2006), and suggest considering Star’s (1999) proposal to view infrastructure from relational and ecological perspectives, rather than as “things.” The discussion of visualization as infrastructure in this case study is a prerequisite for our further conceptualization of the link between visualization, infrastructure, and data politics.
First, data visualization for pandemics is oriented toward the public interest, and it creates an important counterweight to public health emergencies, such as the COVID-19 pandemic. China’s data visualization work has adopted a hybrid agency model, in which the government has been responsible for the development and operation of visualizations to provide clear standards and social commitment, and also for encouraging other actors to analyze and present data. Second, data visualization is ubiquitous. Data visualization has been embedded in the highly digital lives of Chinese people as an important element of social services during the pandemic. By January 2021, visualizations amassed 40 billion page views by 900 million Chinese users (CNNIC, 2021) on mega-digital platforms such as WeChat and Alipay.
In the fight against COVID-19, data visualization has provided mandatory and continuous access to content and data streams in the invisible backstage. More importantly, data visualization is embedded into existing communication networks and physical infrastructures as data visualization technologies that support access to open protocols (State Market Regulatory Administration of China, 2020), which is the classic case of gateways (Klose, 2015). Gateways are a component of social technologies in infrastructure, providing interfaces that allow new systems to access frameworks (Edwards et al., 2007). For example, the visualization of health QR codes operates based on virtual and physical facilities, which can be integrated into QR codes conneted to physical surfaces (e.g. counters and cue boards in public places) via scanning with smartphone, and algorithmic devices that upload and analyze individual data in real time, thereby reflecting hybrid systems of people and machines throughout the urban space. Finally, the dependence of Chinese society on data visualization should also be noted. For the government, the returned visualization results are used for policy decisions and advocacy. For public users, data visualizations are analogous to traffic signals and subway maps; the public must learn and understand the visualizations to update their alert level toward urban spaces, and remind themselves of their own cognitive maps regarding what they need to do and which public places are best avoided, thereby constituting a unique type of technological rematerialization. If the visualization system or related algorithm disappears or breaks down, social disruption may arise.
Based on the demonstrated indicators, we propose that data visualization acts as a significant component of today’s information infrastructure in the face of a public emergency. In the next section, we will further argue that this relationship encompasses the important community of actors and, in particular, how this community of actors is mobilized.
The data visualization politics in China
With visualization practice as an infrastructure, China’s data visualization politics is proposed to have three characteristics, namely common action, spatio-temporality, and tri-partite power relations. While China’s data visualization politics originated from the COVID-19 pandemic, they are likely to continue in the future of domestic governance, and may have referential value for other countries.
First, the pandemic encouraged the need and willingness of the population to engage in common actions. We note the rise of proactive data activism, in which public users can upload and appeal data, thereby changing individual and community visualizations, or can use open data for secondary analysis and discussion on social media to check the accuracy of official data (Zhao and Wu, 2020). This implies that public users in China formed a collective political subject with a common goal to fight the pandemic, and their data practices were not conducted in an individualized or atomized manner. Contrary to Ruppert et al.’s (2017) concerns about the atomization direction of the “subject,” in the case of data visualization in China, individuals ceded data anonymously to a collective political subject to combat a global pandemic.
A range of data visualization practices demonstrates that visualization can facilitate a new spatio-temporality of data that is not limited to the immediacy and temporality proposed by previous researchers (Beraldo and Milan, 2019). Via this generally acknowledged visualization system, policies are formulated or improved. Public users can understand the risk presented by historical or real time data, and visualization is thus characterized by a shared memory and collective mobilization of original temporality (Beraldo and Milan, 2019; Ruppert et al., 2017). Further, data originating from the body and collected from the real world do not remain only in numbers, but connect real and virtual spaces, thereby returning the results of data visualizations to the physical level; they re-profile the spatial meaning of cities and communities (e.g. which areas are in crisis and which are not), serving as a guide to action for public users and decision-makers. Based on the conclusions and experiences generated by these data visualizations, people interact with the visualizations (e.g. by viewing information adjacent to themselves, submitting feedback, exchanging their understanding of the visualizations in the community, and modifying their traveling actions).
Figure 3 presents a framework on how data visualization can be configured in politics, particularly as applied to the power relations between the three groups in policy decision-making. The power relations framework emerges under the conditions of information and communication technology (ICT) development and users’ increasing need for information. Policymakers will shape policy, provide data to visual brokers to complete visualization creation, or disseminate the new policy directly to public users. Public users provide feedback to both policymakers and visual brokers, and may receive persuasion and mobilization. We argue that when most data visualization plays an infrastructural role in society, policy persuasion and advocacy will be mediated by visual brokers. The premise of this argument is the assumption that data visualization carries standpoints and rhetoric aimed at social deployment in security, health, and emergent contexts, such as the COVID-19 pandemic; simultaneously, the roles among power relations are flexible and dynamic, and may overlap in different contexts. For example, public users may also become visual brokers for policy advocacy in communities and organizations, and visual brokers may also emerge within policymakers to achieve internal discussion and reform. The framework of tri-partite power relations in data visualization with policymaking.
As a national narrative of “fighting COVID-19,” it promotes community motivation and consensus on risk perceptions in urban spaces (Ball-Rokeach et al., 2001; Matei et al., 2001), which are difficult to achieve with raw data (Zurkowski, 1984). Via the mediation of visualization, the body and the city have become sites of the political shaping of data, and have been made more visible. This also means that the “subject” has been transferred from an individual to an anonymous collective political subject. Consequently, the privacy of one’s activities and identity, and the “power” of collecting, storing, and modifying data, are temporarily set aside in favor of prioritizing the right to survival in an emergency.
Moreover, moving beyond the narrative in the case study, we consider the role of visualization as an infrastructure in the stage of data politics. Because visualization provides a notion and ecology that brings the public closer to data conclusions, the implication is that it may support the operation of emerging data politics issues. In Ruppert et al.'s (2017, 2020) discussions of data politics and data publics, which include dataset agency and activism, we realize that visualization is an important member of information infrastructure and an easily overlooked technical process. In all this, the availability of this infrastructure is presumed: this infrastructure is necessary to support the perception, participation, proposal, and critique of data by subjects in the “onstage,” and coordinates competition and power configurations through professional practices such as data analysis and data journalism; this is difficult to achieve via raw data, without the interpretation and connection of visualizations. This explains how, in Chinese practice, visualization has become an integral part of people’s actions, as common as traffic lights. We are certainly concerned that visualization is imbued with a quality that operates power, and further hides the transparency of the “backstage” of the data, thereby reducing the disclosure of data sources and processing details, and making it more difficult to reveal the essence of things in the agents of different actors. From practice to the notion, we show how visualization works as infrastructure, and the resulting data visualization politics, to expand the discussion of “data politics.”
Factors of data visualization engagement
The uniqueness and potential of China’s data visualization politics are illustrated by the COVID-19 case, but the feelings and attitudes of public users in the face of data visualization practices must be collected and analyzed to discuss which factors might influence them as data subjects to engage in datafication and activism. An FGD with members of the Shenzhen public health sector and IDIs of Shenzhen residents were conducted to expose this point.
We agree with Kennedy and Allen (2016) that the engagement factor of data visualization is important, and goes beyond the visualization text, but has implications regarding how the “effectiveness” of data visualization is defined. During the FGD and IDIs, questions were asked to detect which factors may influence the engagement of visualizations, and these findings were compared with the six factors proposed by Kennedy et al. Overall, four factors considered in the present research were similar to those considered by Kennedy et al., while two new factors that may influence visualization engagement were identified. The four validated factors are subsequently described. 1. Subject matter. Data visualization is accompanied by storytelling, and the subject matter around which the story revolves influences engagement. The COVID-19 pandemic is the subject matter by which data visualization became meaningful and thus engaging. The data collected from the IDIs reveal that most public users were interested in the outbreak and related topics of the COVID-19 pandemic in the early stage (January-June 2020). Public users reported frequently visiting related platforms to learn about new trends, and tried to figure out what they saw in the visualizations.
IDI: “Of course, I'm very interested because it's about my own health. I would be curious about how many new cases were added each day and how many people died each day” (Male, 57, high school education).
Visual brokers corroborated this factor. They noted public users expressed strong curiosity about the way in which the visualizations interpreted data, and offered their own insights.
FGD: “We get condemned by users who are interested in being updated if we don't post the latest visualization work on our official social media accounts” (Female, visual broker – department director/manager). 2. Beliefs and opinions. The results demonstrate that when a data visualization matches a public user’s beliefs, or challenges his or her original beliefs, the individual will either be satisfied with the visualization, or will remain curious to see more relevant visualizations.
IDI: “What I didn’t expect was that the number of new cases in the United States would grow so much. I was concerned about that, so I would wonder, what will tomorrow’s data be like?” (Female, 25, master’s degree). 3. Emotions. Emotions include how users feel after understanding a data visualization, but can also appear as first impressions that influence whether or not they will continue viewing visualizations. For example, some people reported that they would stop viewing a data visualization when certain visualizations became confusing and scary to them. Furthermore, the user’s own mood can also affect engagement.
IDI: “When I’m in a bad mood, for example, I would rather watch some videos or comics than these charts, which are very much like what I would only engage with when I'm in a math class” (Female, 24, master’s degree). 4. Confidence and skills. Half of the interviewees expressed concerns about their ability to understand the data visualizations, and most of these interviewees were older with lower levels of education. The common explanation they provided for their lack of confidence was that decoding visualizations required data skills and data literacy, without which they could not necessarily draw the right conclusions.
Visual brokers also mentioned this factor. They said that users were uncomfortable with charts that were too professional and complex. When they created this type of content, the number of views usually decreased, and fewer people would want to discuss it.
FGD: “Our feeling is that basic statistical graphics are more of a management tool for decision-makers than for the public. At the same time, the public audiences do not favor visualizations that are too complex, so we need to find a balance” (Female, visual broker – department director/manager).
The two factors mentioned by Kennedy et al. that were not validated in this research were the following. 1. Source/media location. None of the interviewees discussed whether the source of the visualization influenced their trust in the visualization, and the same is true for the media location. Most interviewees, especially the elderly, believed that data visualizations had the same credibility as long as they originated from the government sector or official media. On the one hand, the advanced cognitive processing methods presented by data visualization indicate high-cost human and information resources, thereby enhancing its credibility. On the other hand, the credibility may stem from China’s widespread party control of the media system, as well as its strict control over fake news during the pandemic. 2. Time. Whether one is idle is supposed to be a factor that influences visualization engagement, but none of the interviewees mentioned or agreed with this. According to the official advice to lock down when the pandemic emerged, the vast majority of Chinese people displayed a high level of compliance of staying at home during the Chinese Spring Festival (Zhao, 2020), and they had plenty of time to engage in data visualization. Therefore, the time factor was not validated.
Moreover, based on the empirical material, two new indicators were observed that are believed to provide useful theoretical additions to the discussion of engagement in data visualization during a public emergency. 1. Frequency. When a subject matter and topic appear in public discourse too often, public users will be significantly less interested in it. In the presence of information overload, “COVID fatigue” (Michi et al., 2020) occurs, and data visualization is no exception.
IDI: “I don't really want to look at them anymore. They’ve been in my eyes too many times and I'm kind of tired of them” (Male, 32, PhD degree). 2. Relevance. A notable finding is that relevance (including psychological distance) affects the engagement of public users. The interviewees indicated that they tended to view data visualizations that were relevant to them, such as the number of new cases in the street, neighborhood, and city where they lived and the geographic locations of these cases (the micro- and meso-levels), rather than national or even global pandemic information (the macro-level). This factor was also mentioned in the visual brokers’ FGD. This finding corroborates the research by Peck et al. (2019), which posited that “data is personal.” The browsing of visualization is driven by personal experience, and thus people tended to focus on their current region or hometown, such as by checking where they are on a location map, or were interested in visualizations that were more relevant to information about their age, gender, occupation, etc. Visualizations that only presented an overall picture of the country had very low engagement levels.
IDI: “I know it makes sense to understand our country and the global pandemic situation, but I prefer to wonder what happened near my house” (Male, 50, bachelor’s degree).
Factors of Data Visualization Effectiveness
Via FGDs with visual brokers and IDIs with public users, we propose that data visualizations have different goals in different contexts, and therefore require updated definitions of effectiveness. When the design and narrative of a visualization tend to be neutral (e.g. a simple bar chart showing the change in GDP), namely as a basic visual representation of the raw data, it is referred to as “visualization as illustration,” which is applicable to everyday social contexts. When a visualization is designed and narrated with standpoints and conclusions, it is referred to as “visualization as persuasion.” Visualization as persuasion has the aim of increasing the possibility of citizen participation, advocacy, and campaigns, such as raising awareness of social issues or groups, and changing attitudes and actions.
Visualization effectiveness is defined as comprising four stages. “Visualization as illustration” requires only the first two stages as a framework for the definition of effectiveness. When the visualization is permitted to be persuasive to get people’s attention and to develop some degree of collective action, the definition of effectiveness must be measured in four stages. This definition was formed by drawing on McGuire’s (1968) six-step model of persuasion, which provides a basic and important reference for the design of various attributes of persuasive communication. We believe this adaptation will contribute to the discussion of data visualization effectiveness. 1. Awareness; “I would like to view.” In the first stage, the effectiveness of the visualization is based on the basic design-based work, that is, the topic, color, format, style, and other settings that determine the first impression of public users facing the visualization, that is, whether they are interested in viewing it or not. The influencing factors in this stage, such as the visualization being easily accessible and having an attractive design, were obtained via the empirical material. 2. Comprehension; “Understanding after viewing.” In the second stage, the effectiveness of visualizations lies in their content, that is, in making the information in the raw data more efficiently understandable. The empirical material indicates that public users prefer uncomplicated graphics or visualizations that provide annotations and explanations, which is the reason why location maps and health QR codes became more popular than traditional line charts and bar charts during the COVID-19 pandemic.
IDI: “Sometimes I'm not sure what the visualization wants to convey; maybe just writing the conclusion and telling me what happened would help me a lot” (Male, 44, high school education). 3. Conviction; “Being convinced after understanding.” In the third stage, whether data visualization can have a persuasive effect depends on changes in attitudes and emotions. The ability of information to elicit changes in attitudes and emotions has been the focus of numerous visualization studies (e.g. Joffe, 2008; Van Kleef et al., 2015; Kennedy and Hill, 2018), and is also considered in persuasion research. Our empirical material demonstrates that when viewing a data visualization that is recognized by a viewer, specific emotional responses are provoked (e.g. anger, surprise, sadness), and the viewer is willing to share it with others or to continue to learn about the data, which are considered as factors of the effectiveness of the visualization. 4. Action; “Taking action after being convinced.” In the last stage, the data visualization is considered to correspond to the effectiveness factor of a clear goal and direction of activism. For example, the health QR code dictates the range of options a person has corresponding to the different results. This visualization serves as an all-encompassing collective action via the perception of how to act, whether the action is in line with the policies, and whether everyone is acting.
The definitions and potential factors of data visualization effectiveness.
This framework helps facilitate the understanding of how visualization harkens back to a long tradition of media activism in its role as a technical facility and notion, and this is achieved based on a perspective that encourages users to understand the workings of data and reality. It can simultaneously provide activists with a toolkit to enhance the work of advocacy.
Conclusion and discussion
This article analyzed the rise of data visualization in social and political contexts. We demonstrated the potential use of data visualization for advocacy and social mobilization in a public emergency by using a case study of Shenzhen, China. Regarding the Chinese practice of tackling the COVID-19 pandemic, we claim that visualization mediates data and the real world. To answer RQ1, in times of public emergency, data visualization via the hybrid agency of visual brokers, policymakers, and public users aims to act as an infrastructuralized governance toolkit to inform residents of the reality of the situation and ultimately spur collective action. Regarding RQ2, we proposed a new framework to define the effectiveness of data visualization based on four stages, while adding frequency and relevance as two factors influencing visualization participation to understand the potential ability of visualization to mobilize subjects and facilitate their participation in proactive activism.
Visualization provides a technological practice to bring the public closer to the conclusions of data, and, in the case of China, lowers the threshold for understanding the results to a traffic light-like model, thereby allowing the mobilization of the entire population to act in response to a public health emergency. Over time, visualization has become an integral part of the population’s actions; once it breaks down, it could even collapse society. Which are all indicators of infrastructure. As a result, the three characteristics of co-action, spatio-temporality, and tri-partite power relations were proposed to describe China’s data visualization politics during the COVID-19 pandemic. Beyond the practice, we considered a new epistemology for visualization as an infrastructure of datafication. While it supports the perception, participation, proposal, and critique of data politics issues, it reduces the disclosure of data sources and processing details. While premeditation and bias are always present in big data (Boyd and Crawford, 2012), visualization may make it more difficult for those blinded by visual appeal to detect these problems.
The limitation of this article is that the investigation of the audience was limited by the inability to verify the influence of authoritarian regimes on visual brokers or public users, who may be influenced by social desirability or collectivist culture to maintain cooperation with all policies related to visualization. Although we can note a certain degree of proactive activism adopted in China during the COVID-19 pandemic, it is not a “bottom-up” initiative (Gabrys et al., 2016), and the government has remained dominant. The surveillance and exposure of personal information may be tacitly approved. The Chinese population represents an unspecified subject of rights that may fall short of what Ruppert and Isin (2020) expect regarding the conditions of “data citizens.” Future research will seek to expand the horizons beyond China to investigate how the concept of data visualization politics differs spatially and socially to determine if the findings of this study are solely context-dependent.
When visualizations influence politics, especially when they are deployed by ideologically relevant efforts, we are inevitably concerned with visualization design itself, as well as the associated transparency and accountability issues. Field (2020) called attention to visual brokerage issues related to COVID-19, including faulty data interpretation of classifications, color schemes, and scale symbols. Visualizations based on intentional or unintentional errors can become ineffective guides to action and even cause social panic (Pandey et al., 2015). Moreover, the encouragement of data activism resulting from the COVID-19 pandemic may trigger an outpouring of ethical questions about privacy, surveillance, and morality. Policymakers and visual brokers must make trade-offs between being too detailed and too vague without increasing tensions between privacy and public health (Boyd and Crawford, 2012). For example, many Shenzhen residents demanded the further collection and disclosure of the personal information of COVID-19 patients (Zheng and Wen, 2020). The request to disclose the patients’ names and the detaialed address of their residence caused controversy regarding data surveillance and security.
Therefore, we assert that this research is premised on the idea that persuasion related to data visualizations can present positions rhetorically, but should be in the public interest. We agree that there is a need for a technical due process to ensure the legitimacy of data input and visualization output for political influence. In addition, we recommend focusing on “visual governance” beyond “data governance” (Beer, 2013: 118–119; Correll, 2019; Hill et al., 2016; Kennedy et al., 2016a), that is, introducing mechanisms and strategies to ensure that visualization as infrastructure can support views and action claims in a way that aligns with the overall vision of society. This is, of course, a difficult problem and will serve as a direction for our future research.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by National Social Science Foundation of China (19BXW098).
