Abstract
Social network services (SNSs) provide the functions for both the management of social networks and information diffusion. We considered SNS networks as the networks that embrace both the strong links formed through the making-relationship function and the weak links formed through the sharing function. For identifying the structural role of each link, we constructed the SNS networks using Twitter data and analysed them focusing not on nodes but on links of the networks at a component level and at an ego-network level. More than 200,000 tweets were randomly sampled through the streaming API of Twitter. As a result, we found that weak links formed through the sharing function play a more important role in maintaining the range of information diffusion and provide more structural advantage in acquiring and controlling information than the strong links formed through the making-relationship function.
1. Introduction
Twitter and Facebook are social network services (SNSs) that enable individuals to establish and manage relationships online. A large number of people all over the world use these two SNSs. The very purpose of SNS is to support the management of social networks, but it also enables users to spread user-created information through their social networks.
When discussing fast and extensive information diffusion on SNSs as social media, we are generally referring to Twitter and Facebook [1, 2]. Research works that point out fast diffusion of information on Twitter refer to retweet, the sharing function of Twitter [3–5]. Twitter users can share tweets written by other users and their followers through the retweet function. Facebook also provides the share function that enables users to deliver specific posts to their friends.
However, SNS users do not only use the sharing function to distribute information. boyd et al. [6] found that people utilize the sharing function not only to spread information, but also to build and manage relationships with other users through confirming that they are listening to others, expressing their agreement in public or showing their friendship by accepting others’ requests for a retweet. Studies on retweet motivations for social relation management also show that a retweeter uses the retweet function to both distribute information and establish a relationship with the retweetee who posts the retweeted tweet [7, 8]. That is, the sharing function has an influence on not only information diffusion on SNSs, but also the formation and maintenance of social networks. Thus, the relationships formed through the sharing function have to be reflected in SNS networks.
The relationships formed through the sharing function cannot be the same as the relationships formed through the making-relationship function. The former is characterized by less durability and continuity compared with the latter because the relationship formed through the sharing function is based on the unit of content on SNS networks not the user. However, several kinds of links of different strengths could coexist together on social networks and have their respective merits [9]. Granovetter [9] explains that the strength of social tie (i.e. link) hinges on time spent on the tie, emotional intensity, intimacy and the reciprocal services characterizing the tie. According to the criteria, the relationships formed through the sharing function can be regarded as weak ties while those formed through the making-relationship function on SNS networks can be regarded as strong ties.
Nevertheless, studies on the structure of SNS networks have so far paid most attention to only the relationships formed through the making-relationship function – the follow function of Twitter and the friend function of Facebook [1, 10, 11]. Considering the lack of study on the retweet and the significant influence of SNS on social relationship formation and public opinion formation, for example, it is necessary to investigate the effect of the retweet as a sharing function on the formation of social relationships. Better understanding of such SNS by focusing on specific functions would provide more insightful understanding of the complex social-media-powered communication phenomenon.
Hence, this study proposes a novel approach to understand and analyse structural characteristics of SNS networks as networks that embrace both relationships formed through the making-relationship function and those formed through the sharing function. In particular, focusing on the weak links formed through the sharing function, we examined the effect of the sharing function on the structure of SNS networks.
This study focuses on two functions, the making-relationship function and the sharing function. The making-relationship function enables a user to receive automatically all contents that a specific user uploads on SNS, that is, the follow function of Twitter and the friend function of Facebook. The sharing function enables a user to deliver specific content to the users who are connected with him through the making-relationship function, that is, the retweet function of Twitter and the share function of Facebook.
This paper is organized as follows. Section 2 synthesizes previous studies related to the subject of this study. In Section 3, we describe the data collection procedure and the construction of the SNS networks. Section 4 introduces the methods used to analyse structural characteristics of the links in the networks and presents the results. Finally, we summarize our findings in Section 5.
2. Literature review
boyd and Ellison [12] pointed out that SNSs discriminate against the existing online communication services because the SNS networks are formed not by topics or organizations but by individuals represented by their profiles. However, integrating with Web 2.0, which creates key value through user participation based on open service structure [13], the social network of SNSs has gained the characteristics of an information network that distributes information created by the users. Under the paradigm of Web 2.0 summarized by openness, participation and sharing, the services for helping the production and circulation of information have adopted SNSs and so a wealth of information has been created and distributed within SNSs.
Twitter is an SNS service on which users can publish and read texts shorter than 140 characters. boyd et al. [6] mentioned the following four things as the conventions of Twitter: the @user syntax to refer to specific users; topic indication through the combination of a hashtag (#) and a keyword; the provision of links to outside content by including the URL in the tweets; and retweeting. In particular, retweeting is the specialized sharing function that exists in the forms of forwarding and RSS (Rich Site Summary or Really Simple Syndiation) in email and blog, respectively. The reason why SNSs starting with Twitter have received attention as social media is their capability to spread information through the sharing functions like retweet.
Therefore, many researchers have understood retweet as a key mechanism for information diffusion on Twitter and tried to investigate its characteristics. boyd et al. [6] stipulated that retweetting is both ‘a form of information diffusion’ and ‘a means of participating in a diffuse conversation’. Kwak et al. [1] identified that, if a tweet of a user is retweeted once, it is delivered to an average of 1000 users, regardless of how many followers the user has. They also pointed out that, unlike previous SNSs, the distinct characteristics of Twitter, which make it new communication media, are the fast information propagation through retweet as well as the directionality and the low reciprocity of relationships.
However, studies on retweet have largely concentrated on investigating the factors influencing the retweet frequency. Lee and Kim [14] focused on a retweeter’s cognitive process and showed that, if the message of a tweet is public and political or coincides with the retweeter's attitude on the issue, then the likelihood of retweeting the tweet is higher. There are also content analysis studies on the retweeted tweets. For example, Zhao and colleagues [15] developed a Twitter-LDA (latent Dirichlet allocation) model and carried out topic modelling on both the entire tweet data and the retweeted tweet data. The detected topics on the former and on the latter were different in frequency of categories and types.
However, users do not use the sharing function only to distribute information. boyd et al. [6] asked Twitter users why they used the retweet function, and they categorized the responses into 10 types. The responses showed that people utilize the sharing function not only to disseminate information but also to build and manage relationships with other users through listening an responding others, expressing their agreement in public or showing their friendship by accepting others’ requests for retweets. Recuero et al. [7] tried to apprehend retweet practices from the perspective of social capital through a qualitative analysis of questionnaires administered to Twitter users and four quantitative case studies of retweets. They discussed that Twitter users use the retweet function to express their agreement, provide support and create conversation, as well as diffuse information. Another study investigated Twitter users’ retweet motivations through an online survey and revealed six primary motivations: social interaction, sharing of value-added information, the role of the curator, altruism, timeliness and the sharing of public anger [8].
Previous studies on retweet motivation have indicated that the retweet function has an effect on not only information diffusion, but also relations among users on Twitter. A retweeter can retweet a specific tweet to form the relation with a retweetee. Therefore, social links could be created and maintained through the retweet function, the sharing function, the follow function and the making-relationship function.
To the best of our knowledge, no research has integrated the links formed through the retweet function into the Twitter network. Most of the earlier studies on the Twitter network assume that the Twitter network is built only on the links formed through the follow function. Regarding this, Kwak et al. [1] presumed that the Twitter network is constructed by the follow function, and Chang and Ghim [10] also traced followers and followees repeatedly based on Korean Twitter network data. On the other hand, Achananuparp et al. [16] considered that meaningful relationships are formed through the retweet function as well as through the follow function. However, their network includes only the links formed though the retweet function instead of integrating them, like our study, with the links formed through the follow function, and proposed a model that explains information production and promotion on Twitter based on that network.
Certainly, the relationships formed through the sharing function cannot be the same as the relationships formed through the making-relationship function. The former are less durable and less frequent than the latter. However, weak ties are the decisive factors in social structures, although many social scientists overlook them [17]. Granovetter [9] assumed that the strength of a tie (i.e. a relationship in social network) is ‘a (probably linear) combination of the amount of time, the emotional intensity, the intimacy (mutual confiding), and the reciprocal services which characterize the tie’, and he insisted that weak ties are very likely what connect the specific two nodes solely within a group or link two nodes belonging to different groups.
Gilbert and Karahalios [18] pointed out that related studies do not reflect that not all relationships are the same in SNSs. Few studies have classified the relationships on SNSs and social media into strong ties and weak ties and discussed changes of aspects by the strength of ties on information sharing and interaction [19], Question & Answer [20], or positive feedback on political opinion [21]. However, there is no case that has interpreted the relationships formed through the making-relationship function and the sharing function of SNSs as strong ties and weak ties, respectively, and integrated the two kinds of ties within a network. In this study, we considered the relationships formed through the follow function as strong ties and the relationships formed through the retweet function of Twitter as weak ties, established the networks embracing the two kinds of ties and carried out a structural analysis.
3. Methods
3.1. Data collection
3.1.1. Twitter data
This study was carried out on Twitter. Twitter is representative of SNSs because of its fast and extensive information diffusion, similar to Facebook. Twitter provides two API (application programming interfaces), Streaming API and REST API. Streaming API enables us to collect Twitter’s global stream of tweet data and access streams of the public data or of specific single/multiple users’ data [22]. However, Streaming API delivers only a small fraction of the total volume of tweets by random sampling [23]. Although the volume of Streaming API tweet data changes depending on the data traffic conditions, it is about 10% of the total tweet data [24]. REST API defines many elements of Twitter as resources and allows us to collect the properties of each resource [25].
3.1.2. Collection procedure
The data collection consisted of three steps: node data (user data) sampling, retweet relationships data collection and follow relationships data collection. Because the Streaming API provides only part of the streams of all public tweets by random sampling, we focused on users who posted tweets as the node data of the Twitter network. However, some users had several overlapping tweets. The details are shown in Table 1.
Number of collected tweets.
In the second step, we collected the retweeters’ user IDs and the retweetees’ user IDs of the retweeted tweets among the whole tweet data. We picked out only the retweet relationships between sampled nodes.
Lastly, we identified the follow relationships just in case the nodes have the retweet relationships. To ascertain all of the relationships and compose a network, we needed to collect the follow relationships between all collected nodes. We determined that it is enough to assess the network including only the nodes that have the retweet relationships and compare the links formed through the retweet function with the links formed through the follow function. Because we focused on the links formed through the retweet function, the sharing function, we considered that it is practically impossible to confirm the follow relationships between all the nodes owing to the rate limitation of Twitter API. Users who withdrew from Twitter because they were not willing to provide information about themselves or were suspended from Twitter during data collection were removed from the data.
The data collection was carried out on two separate occasions. Table 2 shows the summary of the first data collection conducted in October 2013 and the second data collection conducted in March 2014. The reason that the time consumed to gather similar volume of data differed in the first data collection and the second data collection is that Twitter adjusts the amount of tweets data delivered through Streaming API depending on traffic.
Summary of data collection.
The outlines of the network data collected through the three phases are shown in Table 3. Although there is no specified requirement for the ratio of retweeted tweets to the whole tweets in this kind of analysis, from a similar study [10] that had 13%, we can justify our data, which reached 27.9% on average (first, 21.7%; second, 34.3%). It is necessary in further studies to investigate whether such a difference comes from criteria of Twitter data collection or changes in Twitter use pattern owing to the time gap between the two follow-ups.
Outlines of network data.
3.2. Visualization and characteristics of Twitter networks
The visualizations of the two Twitter networks created from the first and the second datasets are shown in Figures 1 and 2, respectively. We used Pajek, a network analysis tool, and the Kamada–Kawai algorithm to construct the two networks. Arcs in Figures 1 and 2 start from followers and retweeters to head followees and retweetees. Bold lines mean that the relationships formed through the retweet function, and thin lines mean that the relationships formed through the follow function.

Twitter network from first dataset.

Twitter network from second dataset.
Figures 1 and 2 illustrate that most of the nodes in the two networks except for those included in the respective largest weak component have relationships with a small number of nodes. Because we collected the random sampled Tweet data without any restriction like language or country, it is difficult to expect that well-connected networks will be established. The fact that the densities of the two networks provided in Table 4 and the densities of the respective largest weak components provided in Table 6 are very low also can be explained by the same token.
Summaries of Twitter networks.
The distributions of the two networks follow a power law, which means that a few hubs having exceptionally numerous links coexist with a large number of nodes having few links. The indegree and outdegree exponents in the first Twitter network are 2.359 and 2.516, respectively. The former and latter in the second Twitter network are 1.955 and 1.984, respectively. The Twitter networks constructed based only on links formed through the making-relationship function are known as scale-free networks [1, 10, 16, 26]. The Twitter networks created based not only on links formed through the making-relationship function, but also links formed through the sharing function in this study are also scale-free networks.
4. Analysis on retweet links
This study focused on links rather than nodes because it aimed to investigate the specific function that helps develop the relationships within an SNS network. We conducted network analysis on the largest component that includes the links formed through the sharing function at both component level and ego-network level. The component also embraces the links formed through the making-relationship function to identify structural characteristic difference by kinds of links. We utilized two network analysis tools together, Pajek and UCINET.
At the component level, we carried out bi-component analysis and examined distributions of bridges by kinds of links. At ego-network level, we performed structural holes analysis to quantify structural advantage. In particular, we computed the redundancy of all links and showed difference in redundancy distribution by kind of link.
4.1. Component composition
Component is a subnetwork that consists of connected nodes. There are direct and/or indirect paths between every node in a component, which represents a unit in which information or resources are distributed [27, 28]. Strong and weak components can be defined in a network. The former considers directionality while the latter does not. Social network analysis generally considers weak components because social relationships are reciprocal and it is difficult to disentangle directionality [29]. Thus, we also assumed weak components.
The spreads of weak components in the Twitter networks constructed from the first and second datasets are shown in Table 5. The components (n = 294 in the first; n = 255 in the second) in the networks include 23% (first) and 13% (second) of the total nodes, respectively. Most components (90% in the first, 80% in the second) are made up of only two nodes. Thus, it is expected that homogeneity in information exchange between subgroups is not high. One reason for this could be that we collected the Twitter data without restricting languages or countries of users. Applying these restrictions in the process of collecting data would create more homogeneous network.
Weak components in Twitter network.
The outlines of the largest weak components in the first and second Twitter networks are shown in Table 6. The largest weak components in the two Twitter networks have greater density and average degree compared with the whole networks. However, the degree distributions follow a power law. The indegree and outdegree exponents in the largest weak component in the first Twitter network are 2.099 and 2.225 and in the second Twitter network are 1.860 and 1.886, respectively.
Summaries of largest weak components.
4.2. Analysis of links
4.2.1. Classification of links
Social networks are generally assumed to be undirected [29]. Network analysis techniques that developed based on social networks also suppose undirectionality in most cases. Bi-component analysis applied in this study presumes undirected networks. Although structural holes theory presumes directed networks, there is no difference between directed and undirected network in the case of the binary network because the theory understands that the links between two nodes are unconditionally reciprocal. Therefore, we conducted link analysis after converting the Twitter networks into undirected networks.
When directed networks are converted into undirected networks, the duplicate links could emerge because two arcs that had headed in different directions changed into the same links. In this instance, we considered only one link.
Additionally, duplicate links could form through the follow function and through the retweet function between the two nodes in the Twitter networks. In this case, we considered only the strong link formed through the follow function, as the weak link formed through the retweet function is likely to be established on the account of the strong link.
The classification of links of the largest weak component in the Twitter networks according to the above two criteria is shown Table 7. When the Twitter networks were considered as directed networks, the proportion of arcs in the largest weak component in the first Twitter network and in the second Twitter network was about 1:5. However, after converting the Twitter networks into undirected networks and removing the duplicated links, the proportion changed to about 1:4 because more duplicated links were eliminated from the second Twitter network, which had greater average degree compared with the first Twitter network.
Classification of links in the largest weak component.
4.2.2. Distribution of bridges
A component is a unit in which information or resources of a network are distributed. If removing a link in a network increased the number of components, then the link would be called a ‘bridge’. Bridges play major roles in maintaining the connectivity and diffusing information in a network [28, 30].
We grouped all links of the largest weak component in the first Twitter network and the second Twitter network into bridges and links that are not bridges. The results are shown in Table 8.
Distribution of bridges by link type.
Overall, 227 retweet links are in the largest weak component in the first Twitter network. Among them, 162 retweet links (71.4% of all retweet links) are bridges that are unique paths between the specific two nodes, whereas the proportion of bridges among 691 follow links is only 34.3%, less than half of retweet links.
We saw similar tendency of the largest weak component in the second Twitter network. Although the proportion of bridges among all the links decreased to 18.1% from 43.5% in the first Twitter network, the proportion of bridges among 399 retweet links was 44.6%. This figure is over twice 15.0%, the proportion of bridges among 2787 follow links.
A χ 2 test was conducted to examine the distribution of bridges by kinds of links. The result of the χ 2 test on all links of the largest weak component in the first Twitter network was significant (Pearson χ 2 = 95.542, p-value < 0.001), showing that the distribution of bridges depend on kinds of links. The result of the χ 2 test on all the links of the largest weak component in the second Twitter network was also significant (Pearson χ 2 = 209.855, p-value < 0.001). For each test, the degree of freedom was 1.
The distribution of bridges depends on kinds of links, meaning that the probability of maintaining the connectivity of a network after removing a link from the network fluctuates with the kind of the link. In other words, retweet links among which the proportion of bridges is higher than follow links contribute more to sustaining the connectivity of a network as well as to the knowledge and information diffusion from a perspective of the entire network.
4.2.3. Structural holes: distribution of redundancy
The structural hole, proposed by Ronald S. Burt, is a concept that can stipulate the positional status of each node in ego networks [31], and it has been applied widely in the measurement of structural position of a specific node in a network [32]. Structural holes mean the positions without which certain two nodes cannot connect. Those are the positions in which there is no duplicated connection. Nodes in structural holes have more opportunities for information access, timing, referrals and control [33, pp. 8–49].
Burt suggested two measures, redundancy and constraint, to quantify structural holes [33, pp. 50–81.]. Redundancy can be computed for links, and it shows the degree of indirect duplicated connections, excluding the direct links between the two nodes. Constraint can be computed on both links and nodes. The former is called dyadic constraint while the latter is called aggregate constraint. Dyadic constraint placed on a link suggests that two nodes connected by this link lost structural advantage because of the existence of the link in a network.
Redundancy can explain the positional status of a specific link in an ego-network that centres around a specific node while constraint is more suited to the perspective on a whole network. Therefore, we adopted redundancy between the two measures in this study.
Structural holes theory assumes reciprocal directed networks. Thus, two relationships, one per node, exist per link between the two nodes of undirected networks. When structural holes analysis is carried out on links of undirected networks, redundancy and dyadic constraint are calculated for relationships whose numbers are twice as many as links. We computed redundancy for all the links of the largest weak component in the first Twitter network and in the second Twitter network. Descriptive statistics of the results are shown in Table 9.
Descriptive statistics of redundancy.
The Mann–Whitney U-test was used to examine the differences in the distributions of redundancy by kinds of links. The Mann–Whitney U-test is a non-parametric statistical method for assessing whether two independent samples are drawn from identical distributions [34]. If it is unclear whether a population follows a normal distribution, then using a nonparametric statistical method to test the hypothesis increases the confidence in the results [35].
We classified all the links of the largest weak component in the first Twitter network and the second Twitter network into follow and retweet links. Subsequently, we obtained the distributions of redundancy by kinds of links. The distributions of redundancy by kinds of links are plotted in Figures 3 and 4. The results of the Mann–Whitney U-test are shown in Table 10.

Distributions of redundancy by kinds of links: first dataset.

Distributions of redundancy by kinds of links: second dataset.
Descriptive statistics of distributions of redundancy by kinds of links.
p-Value < 0.001.
The means of redundancies are higher for follow links (first 0.104, second 0.172) than for retweet links (first 0.008, second 0.086). The results of the Mann–Whitney U-test show that the redundancy distributions by kinds of links are statistically different (first p-value < 0.001, second p-value < 0.001).
The bigger the redundancy of a link, the more indirect duplicated connections exists. In our experiment, the distributions of redundancy of the links differ by kinds of links and the average redundancy of follow links is bigger than that of retweet links. Therefore, retweet links have fewer duplicated connections compared with follow links and provide more structural advantages in terms of acquiring information or controlling information flow to individual nodes.
5. Discussions and conclusion
We considered SNS networks as the networks, which embrace both the strong links formed through the making-relationship function and the weak links formed through the sharing function. By adopting the concepts of bridge and redundancy, we analysed structural characteristics of the two links at a component level and at ego-network level. As a result, two conclusions were drawn.
First, at a component level, we observed that the weak links formed through the sharing function play a more important role than the strong links formed through the making-relationship function in maintaining the connectivity of a component and the range of information diffusion. The distributions of bridges that become unique paths of information diffusion were statistically different by kinds of links in the Twitter networks, and the proportion of bridges among the links formed through the retweet function was more than twice as high as the proportion of bridges among the links formed through the follow function.
Second, at ego-network level, the weak links are less redundant compared with the strong links so that they provide a greater structural advantage to acquire new information and control information flow between other nodes. The distributions of redundancy of the links in the Twitter networks differed by kinds of links, and the average redundancy of the strong links formed through the follow function was bigger than the redundancy of the weak links formed through the retweet function.
The findings of this study suggest that weak ties formed through the sharing function can expand the range of information and knowledge distribution from the perspective of an entire network. At the same time, weak ties can give more information benefits to individuals in a network. Weak ties could be used to bridge social capital [36], influencing the connection with external resources and the spread of information in the network.
As has been proved with the evolution of SNSs into social media, it is difficult to separate social networks and information networks. This study demonstrated that the information-oriented function also can contribute to the domain and robustness of social networks. When organizing a social network in an organization or online, we should examine and introduce information-oriented services actively for formation and extension of the network.
This study is meaningful in that it integrated the links formed through the making-relationship function and through the sharing function to construct SNS network. In addition, it investigated the structural role of the weak ties formed through the sharing function compared with the strong ties formed through the making-relationship function, focusing not on nodes but on links of the network. Despite these implications, this study has limitations in that it used only a limited amount of Twitter data among various SNSs. Therefore, further studies should replicate our findings using more data from different SNSs. Additionally, from the perspective of network analysis, a smaller number of indexes can be applied to link. Thus, it would be valuable to examine various types of links that exist in social networks.
Footnotes
Acknowledgements
This work was supported by the BK21 Plus (Brain Korea 21 Plus) project funded by the National Research Foundation of Korea (2014-11-0023).
