Abstract
The value and role of the Like button in social media have gained increased attention/focus, yet we know little about how liking relations form between Likers and Likeds. We study this problem in an online healthcare context from a social network perspective. Taking into account the effects of both the network structures and the attributes of Likers and Likeds, we utilize a theory-grounded statistical modelling approach, Exponential Random Graph Models (ERGMs), to model the liking network in an online healthcare community. The results of ERGM analysis reveal that, while network degree exhibits a big effect in the liking process, individual attributes like the level of past involvement and degree of activity also positively influence members’ future liking behaviour and performance. The evaluation indicates that our model is an effective method to identify the formation of liking networks. The findings extend the understanding of online liking behaviour and provide insights into harnessing the power of liking.
1. Introduction
Health 2.0, the term that describes Web 2.0 technologies developed to support healthcare [1], creates an unprecedented opportunity for patients to obtain ‘social’ interactions [2] and support [3]. Among other Health 2.0 technologies, ‘Like’, a social button that allows individuals to express their positive affective reaction to online content, has been increasingly introduced into online healthcare communities. Liking is inherently dyadic, involving two parties: the Liker who shows his/her positive preferences with a Like, and the Liked whose online content is liked by others. Compared with traditional Web 2.0 websites that highlight the role of Like on the improvement of reputation [4, 5], given the importance of social support for patients, liking in online healthcare communities offers a unique channel to observe positive emotional support between Likers and Likeds. Through clicking the Like button, positive emotions of Likers, such as agreement, appreciation and compassion, are produced, leading Likeds to feel understood and that they matter to others. In order to increase the number of beneficiaries and contributors of liking behaviour, it is important to understand how Likes are generated between Likers and Likeds in online healthcare communities.
This study aims to address this problem from a social network perspective. When the liking behaviour occurs, a liking tie is created from a Liker to a Liked. A liking network refers to a set of individuals and a set of liking ties between them. Liking networks establish a bond between individuals and their positive preferences. In the new point of view, we can investigate the problem by considering how the effects of structural and individual characteristics operate together.
To achieve this objective, we consider the liking tie as a dyadic variable and propose a conditional probability, assessing whether a liking tie forms between a Liker and a Liked, given all the other ties. However, the dynamics of liking networks is a complex process involving multiple microprocesses. Members differ in network position and individual attributes, which influences the way liking ties generate. It is difficult to identify the effects of individual intrinsic properties from network dependencies. Another challenge is that dyads within a network are not independent observations [6]. For instance, the degree centrality of an actor is related to the degree centrality of her/his peers in networks. Thus statistical methods with independence assumptions violate the interdependent nature of network data. Further, we expect to go beyond most common statistical techniques (e.g. regression), which attempt to describe the existence of one or more associations in observed data and find a well-fitting model of liking networks that is able to model networks as a whole by considering difference dependences. A network modelling method that can include individual attributes and structural dependencies simultaneously and handle the interdependency of network data well is needed.
In this study, taking into account the effects of both network positions and attributes of Likers and Likeds, we construct a liking network based on the data from an online healthcare community and utilize a theory-grounded statistical modelling approach – Exponential Random Graph Models (ERGMs) – to model the liking network. We then conduct a goodness-of-fit test to examine how well our generated model fits the observed network data.
Identifying how liking networks form in online healthcare communities carries value. First, existing literature on how Likes are generated is limited. We attempt to fill this gap in an online healthcare setting. An effective approach is proposed to model liking networks. Second, this study can help create a relationship between liking behaviour and users’ attributes and network positions. Whether patients can gain needed support is critical to the sustainable development of Health 2.0 websites. A Health 2.0 website provider should be aware of the liking network and try to drive it to evolve in order to benefit more patients. It is critical to determine whether and to what extent certain network structural or individual characteristics facilitate liking.
The remainder of this paper is organized as follows. In Section 2, we briefly review recent related work and develop the hypotheses. Section 3 describes the data and ERGM analysis. We present the results of ERGM analysis in Section 4 and report the evaluation of our model in Section 5. Section 6 concludes the paper.
2. Related work and hypothesis development
2.1. Related work
Our work draws on three streams of related research: patient networks, Like and ERGMs. The objective of this section is not to display the whole picture of these areas, but to highlight the gap we seek to fill.
The role of peers in improving health condition has been identified by the existing literature. Prior studies emphasize that the decisions to change health behaviour are not made by persons in isolation, but rather reflect choices made by people within a short social distance in networks. How ties are formed among patients can influence whether health behaviour changes [7, 8]. Patient network formation research mainly examines how and why patients connect with each other. Multiple factors could influence the formation process of patient network ties. Most of the factors from prior research can be categorized into structural dependencies and actor attributes. For instance, structural dependencies have been found to drive the evolution of patient subscription networks [9]. Another study shows that individual health-related characteristics contribute to the formation of patient friendship networks [10]. Whereas extant studies focus on friendship and subscription relationships among patients, to our knowledge, there has not been a study on liking relationships in the online healthcare context.
Since Facebook’s introduction in 2009, the Like button has been widely adopted by various social media. The Like function offers a channel for people to evaluate online content and express emotions. Companies and politicians have increasingly realized the value of Like and utilized it for various purposes in recent years. Likes can be an important indicator of online reputation [11], which increases the reading probability. When the Like button is clicked, a score is added to record the number of likes from different Likers. Individual liking behaviour transforms into a single number that is countable and comparable [12]. The number of Likes indicates an evaluation by crowd efforts, which becomes WOM (word-of-mouth) and influences the decisions of other people [4]. Existing evidence has shown that Facebook Likes can drive additional product sales [5]. The liking behaviour provides a unique way to record and quantify users’ interests. A recent study shows that liking behaviours can be used to estimate private traits [13]. While previous studies mainly highlight the role and value of liking behaviour, less attention has been paid to how liking relations form between Likers and Likeds.
Various statistical techniques have been proposed for social network analysis. The assortativity coefficients, proposed by Newman [14], can be seen as an important tool to study assortative patterns in networks [15]. Blockmodelling allows the interactions within and between sets of nodes with particular attributes to be examined. However, neither assortativity coefficients nor Blockmodelling can examine multiple attributes simultaneously and they are unwieldy to assess the relative contribution of different effects. Logistical regression is able to predict network ties by attributes, but it cannot deal with the interdependencies in network data well. Other methods, such as Multiple Regression Quadratic Assignment Procedure (MRQAP), can be adopted to examine multiple relations at a time, but it cannot handle well skewness in the distribution [16]. Importantly, to better understand network dynamics, a well-fitting network model is of greater value than techniques that only examine whether certain associations exists [6]. ERGMs are a class of network models that examine the presence of a tie between a pair of individuals with a set of predictor variables, such as network structures, individual attributes and dyadic covariates [17]. This class of models provides a solution to both of the problems mentioned above. First, this approach allows us to capture the variability that is hard to model in detail, as well as the regularities in the process of tie formation [18]. Also, ERGMs are capable of incorporating individual characteristics and structural effects simultaneously and assess the relative contribution of each social processes by estimating the effect of each dependency given the others [6]. Most importantly, unlike many statistical methods with independence assumptions, ERGMs assume that the likelihood of that the presence of network ties are interdependent, which can deal with network data well [16]. In recent studies, Wimmer and Lewis utilized ERGMs to investigate racial homophily in a friendship network based on Facebook [19]. Another study found the presence of both direct and indirect reciprocity in network exchanges in online communities using ERGMs [20]. A ‘performance-based clustering’ phenomenon was observed within a large open-source community by examining strategic selection and homophily [21]. These studies suggest that ERGMs represent a promising class of model to deal with network data in the online social media context, which meets our need.
2.2. Hypothesis development
Following prior studies mentioned above, we consider the effects of both structural dependencies and individual characteristics and propose our hypotheses correspondingly as follows.
2.2.1. Preferential attachment
Among other structural dependencies, preferential attachment has been mostly studied in extant social networks studies. Actors in networks who have a large number of ties are often considered to be highly local, influential and popular. For directed networks in which the edges are directed arcs, the in-degree measures the number of incoming links an actor has. In a liking network, it shows the number of sources of Likes a user receives. The number of arrows from an actor is measured by out-degree. It indicates the number of Liked whose content has been liked by the actor. Degree centrality of an actor can be used as an indicator of the importance of its potential communication activity [22]. ‘Hub’ nodes with many more links than others are recognized as opinion leaders or influential members who have a big impact on peers’ decisions [23]. According to the theory of preferential attachment, hubs tend to get new links [24]. Existing studies tend to consider preferential attachment as an essential factor in the formation of links [25, 26]. It is natural to expect that users with many Likers have more incentive to contribute. Also, users’ high out-degree reflects their intrinsic attributes of reading and liking extensively. Thus we hypothesize that in-degree or out-degree of a patient will be positively correlated with liking relations.
Hypothesis 1.1: Patients with high out-degree are more likely to give more likes to different Likeds.
Hypothesis 1.2: Patients with high in-degree are more likely to receive more likes from different Likers.
2.2.2. Past involvement
Past involvement, in this study, refers to a user’s experience of liking or being liked. The level of past involvement positively influences later participation [27]. Individuals with liking experience are more likely to be familiar with the Like function and use it more often. Meanwhile, according to the game theory, the imperfect information environment will lead to ‘reputation effect’ [28]. Given uncertain environments, members with a good reputation may gain trustworthiness and more social contacts. In sociology, it is also termed as the ‘Matthew effect’ that people who receive recognition early tend to gain more credit than those who do similar work later [29]. People who receive many likes signal a good reputation, suggesting that their content is often helpful and attractive. Their content will continue to be worthy of trust in the future, which creates a higher chance to be liked. On the other hand, the decision to like tends to be influenced by the previous choices of others by the effects of herd behaviour. Therefore, we propose the hypotheses that past involvement will influence likes received and given.
Hypothesis 2.1: Patients who have given more likes are more likely to give more likes to different Likeds.
Hypothesis 2.2: Patients who have received more likes are more likely to receive more likes from different Likers.
2.2.3. Activeness
Activeness measures the level of participation shown by the users. Individual activeness in an online community is often extremely uneven. A very small portion of the members are often much more active (i.e. they create more posts) [30]. Active members have a much greater willingness to participate in activities in an online community. Current evidence shows that there is a strong correlation between different participation behaviours [31, 32]. Therefore, members who participate actively tend to engage in more liking activities. Meanwhile, their active participation increases their visibility in online communities, which creates a higher chance for them to be liked. We expect that there is a positive relationship between activeness and the formation of liking relations. We therefore hypothesize the following:
Hypothesis 3.1: Active patients are more likely to give more likes to different Likeds.
Hypothesis 3.2: Active patients are more likely to be liked by different Likers.
3. Methodology
3.1. Data
We obtained the dataset from the Diabetes Forum (http://www.diabetesforum.com/), an online healthcare community that offers web resources about diabetes and for diabetics. We targeted this setting because diabetes mellitus is a chronic metabolic disease. Diabetics often need long-term self-management of their health condition (e.g. regular exercise, good diet and utilization of insulin) as well as professional therapy. Thus they have more enthusiasm to seek support from online healthcare communities. Aiming to assist patients to better manage their health condition, the website provides members with content and a forum to discuss topics related to diabetes such as diabetes news, self-introduction and complications. The community also contains Like elements. It allows patients to ‘like’ posts or replies, providing a fitting context for our purpose. Further, the community allows users to directly observe a post’s first three Likers’ names and the total number of the rest. By clicking the number, we could further learn the rest of the Likers’ names.
We collected all of the posts generated from 1 November 2011 to 31 October 2013. After excluding the posts without a Like, we finally obtained a dataset of 3593 posts with at least one Like and 308 members who were Likers or Likeds. The website also displays the cumulative number of occurrences of liking and being liked for each member from their registered time. We collected this information as past involvement.
Then we constructed a directed network with 308 nodes and 2193 edges. The network structure graph is shown in Figure 1. Table 1 reports summary statistics for our variables. Individual activeness is measured by the average number of posts written by the patient per day.

The liking network graph.
Summary statistics for ERGM analysis.
3.2. ERGM analysis
As mentioned above, we applied ERGM analysis to validate the hypotheses. The general mathematical form of exponential random graph models is:
where the summation in the model is over all configurations, κ is a normalizing quantity to ensure proper probability distribution, η A indicates a vector of parameters corresponding to configuration A, and g A(y) is the network statistic. A configuration is a subset (generally small) of possible network ties (and/or actor attributes) in the network [33].
In ERGMs, we specify an ERGM by including parameters related to local social processes in the liking network. Each parameter implies a particular configuration corresponding to a hypothesis. For instance, the alternating k-instar parameter measures the tendency to form a skewed in-degree distribution. Then the model represents a distribution of random graphs based on these configurations. ERGMs assess the prevalence of the configurations above what would appear by chance alone [34]. If a configuration is more prevalent than by chance, the configuration has a significant probability of existing in the network. There are generally two ways to estimate parameters in ERGMs: maximum pseudolikelihood estimation and MCMC (Markov Chain Monte Carlo) maximum likelihood estimation [19]. In contrast to MCMC maximum likelihood, maximum pseudolikelihood estimation tends to be misleading, thus, MCMC maximum likelihood estimation is preferred in extant literature [6]. Thus, we estimated the parameters with MCMC maximum likelihood estimation methods, as previous studies suggested [35, 36]. In the MCMC maximum likelihood estimation, a distribution of random networks is produced and their statistics are compared with the observed network. The more similar the network statistics are, the better the ERGM estimations are. The parameter estimates are revised until they converge by simulations. We estimated the model for ERGM analysis in the PNet. 1 The estimated parameters in the final fitted model indicate whether one or some of the configurations are statistically frequently observed.
Next, we examined whether our fitted model can capture the characteristics of the observed network. We conducted a goodness-of-fit test to validate how well the ERGMs fit the observed network. We generated a large number of random graphs from the estimated model and evaluated whether these simulated networks resembled the original network in a series of network statistics.
4. Results
Table 2 reports the results of ERGM estimates. The estimates and standard errors can be seen in columns 4 and 5. The significance of an estimated parameter indicates the existence of a specific configuration in the observed network. A parameter can be considered significant when the estimate is at least twice the standard error [37, 38]. t-Statistics indicate the degree of convergence, which is preferably not greater than 0.1.
Estimation results of ERGM.
Notes: t-statistics = (observation − sample mean)/standard error.
Statistical significance.
The degree-related parameter estimates are positive and significant, indicating that degree has a positive influence on liking relation formation. Hypotheses 1.1 and 1.2 are supported. This tells us that either in-degree distribution or out-degree distribution in the liking network is skewed. Hubs are prevalent in the liking network. Some popular members are clearly preferred than others as Likeds. Meanwhile, some members like others more often. This finding provides new evidence to support the preferential attachment theory [24]. The effect of past involvement is identified. Hypotheses 2.1 and 2.2 are supported. The level of past involvement positively influences members’ future liking behaviour and performance. Experienced members are more likely to use the Like function again and more often. The result thus supports the widely accepted notion in previous research that past behaviour plays a critical role in predicting future behaviour. On the other hand, the more likes a patient has received, the more likes the patient will receive from different Likers. A member with more likes indicates that her/his generated content is often more valuable or meets the need for support and is more likely to be liked than others. Also, since the decision to like is often influenced by the previous liking of others, this leads to a herd effect, enabling a post to be liked more after being liked. The effects of activeness are also identified. Both Hypotheses 3.1 and 3.2 are supported. Active patients in online healthcare communities tend to enjoy liking activities and be liked more often. Active members generally have more energy and spend more time in engaging in all kinds of activities in the online community, including the Like function. On the other hand, members who frequently post are more visible and more likely to receive incoming liking ties from others.
To sum up, in the liking network, the number of past received likes and activeness can drive the Liked to gain more liking links, while out-degree, the number of past given likes and activeness can lead Likers to develop more liking relations with others. Among others, network degree exhibits the largest effect, suggesting that network structure plays a greater role than individual attributes in the formation of the liking network. Further, the effect of past involvement is less than that of activeness on the generation of the liking relations.
5. Evaluation
We conducted a goodness-of-fit test to examine how well our generated model fits the observed data. We generated 8,000,000 simulated networks and picked up 1000 samples. A series of network statistics were compared between the samples and the observed network. The lower the difference is, the better our estimated model is.
The goodness-of-fit testing results are shown in Table 3 for a variety of network statistics. These results were generated by comparing the observed network and simulated networks. The difference in the statistics of the observed network and generated samples indicates whether the estimated model fits the data well. Column 1 includes the parameters in our model. Column 2 shows the observed value of the network data. The mean and standard error from 1000 sample simulated networks can be seen in columns 3 and 4. The ERGM is a good fit for the observed network if the absolute value of a t-ratio is less than 0.1 [39]. The results show that all estimated parameters have t-values less than 0.1. The results indicate that our model is capable of reproducing the observed network well.
Results of goodness-of-fit test.
Notes: t-statistics = (observation − sample mean)/standard error.
6. Conclusion
Our study explains how liking networks form in the online healthcare context. We utilize a theory-grounded statistical modelling approach, Exponential Random Graph Models, to model the liking network in an online healthcare community. The results of ERGM analysis reveal that the joint effects of preferential attachment, past involvement of being a Liked or Liker and activeness are responsible for generating the liking network.
The research has important implications in theory and practice. First, to our knowledge, this study is the first to examine how liking relations are formed from a social network perspective. Our findings add to the growing literature of liking behaviour in social media and extend our understanding of its use in online communities. Second, similar network structures may be caused by different social processes. We thus propose an ERGM to model the liking network, which makes it possible to assess the effects of individual attributes and structural dependencies simultaneously. By explicitly investigating this, our study sheds light on the relative importance of different social processes. The results of goodness-of-fit indicate that our model can well fit the observed liking network data. The study also has important practical value for patients and Health 2.0 platform providers. As our research has shown, patients who seek like-type emotional support should actively participate in online healthcare communities. In the meantime, creating close relationships with others (e.g. making online friends) who have given many likes will increase individual opportunities to be liked. It also provides insight into how to better harness the power of liking. Our research results demonstrate that network degree, past involvement and activeness do play a role in the liking process, which helps Health 2.0 platform providers who are interested in pursuing strategies of enhancing liking behaviour to better understand how liking generates from a network perspective. When trying to evaluate the importance of each key member for liking behaviour on a website, their microstructural social environment, past involvement and activeness should be considered.
The major limitation of our study is that we focus only on a healthcare community. It is imperative to extend the analysis to other social media like Facebook. Secondly, owing to data limitation, we only examine the behavioural characteristics of patients in liking networks. Patients tend to gain information and emotional support from others with similar health-related characteristics, which may also contribute to the formation of liking networks. Further work will extend to these fields.
Footnotes
Acknowledgements
Many thanks to Joshana Shibchurn for her assistance in language revision. Also, thanks to the editor and all anonymous reviewers for the constructive comments.
Funding
This work was supported by the National Natural Science Foundation of China (grant nos 71328103 and 71171067).
