Abstract
Influence analysis, derived from Social Network Analysis (SNA), is extremely useful in academic literature analytic. Different Academic Social Network Sites (ASNS) have been widely examined for influence analysis in terms of co-authorship and co-citation networks. The impact of other network-based features, such as followers and followings, provided by ASNS such as ResearchGate (RG) and Academia is yet to be anatomised. As proven in ingrained social theories, the followers and followings have significant impact in influence prorogation. This research aims at examining the same in one of the widely adopted ASNS, RG. The rendering process is developed to render real-time RG information, which is modelled into graph. Standard centrality measures are implemented to identify influential users from the constructed RG graph. Each centrality measure gives a list of top-k influential RG users. The results are compared with RGScore and Total Research Interest (TRI) to discover the most effective centrality measure. Betweenness and closeness centrality measures have shown the outperforming results compared with others. A procedure is established to discover influential RG users that are commonly present in all top-k centrality results to identify dominant skills, affiliations, departments and locations from the rendered data.
Keywords
1. Introduction
The social media platforms provide means of social communication or personal information sharing [1,2]. Likewise, Academic Social Network Sites (ASNS) are profusely used for various research centric activities such as research sharing, academic collaborations, scientific communication, feedback collection and technical information dissemination [3]. ASNS provide a flexible medium to manage the explosive growth in research activities compared with traditional scientific outlets, such as journals and conferences [4]. Prevalent exploitation of ASNS, such as ResearchGate (RG), Google Scholar (GS), Academia.edu and Mendeley, has opened a novel paradigms of research in the field of scientific influence finding, inspired from social influence finding [5,6]. Finding scientific influence in terms of scholars [7,8], research topics [9], journals [10,11] and publications [12–14] using co-authorship and co-citation networks is a well-known concept, as these networks are available for many platforms such as Mendeley, Academia.edu, Scopus, GS, Digital Bibliography and Library Project (DBLP) and Association for Computing Machinery (ACM).
Apart from the co-authorship and co-citation connections, ASNS, such as RG and Academia, provide user-to-user connections in terms of followers and followings. It is observed that the perspective of follower and following connections is yet to be explored with respect to scientific influence identification. Recently conducted surveys demonstrate the higher utilisation of RG among other well-known ASNS in terms of number of registered profiles [6,15] and active participation [16–18]. This motivates our research focus mainly on RG. In this research, we explore RG in terms of connected graph with an aim of anatomising the impact of follower and following connections on influence analysis.
Information shared on RG is categorised into user profile content: research items, citations, reads, recommendations, department, position, affiliation, etc., and user-to-user links: followers, followings, co-authors and co-citations. A data rendering process is developed to render RG users’ profile content based on follower and following connections present among users. The rendered information is modelled into graph using graph modelling technique. The proposed methodology is an initial step to generate and utilise graph-based model of RG information to find influence considering user content and user-to-user followers and followings information present on RG. The visual graph layout helps in identifying how the influence propagates through follower and following connections.
Standard graph centrality measures – betweenness centrality (BC), closeness centrality (CC), eigenvector centrality (EvC) and PageRank centrality (PRC) – are implemented on constructed RG graph for influence identification. The concept is to leverage the follower and following connections to see how the influence of one RG user affects the influence score of other connected RG users through centrality measures. Each centrality measure identifies a list of top-k influential users from RG graph. The results are compared in terms of ranking correlations with influence assessment scores – RGScore and Total Research Interest (TRI) – given to each registered and active user by RG. RGScore is mentioned to be calculated based on the research in author’s RG profile and the interactions of other authors with it. TRI is mentioned as the sum of the research interest of each research item present in authors’ RG profile.
Also, a process is established to find the common influence, in which a list having distinct influential users that are present for at least once in centrality top-k lists is computed. The significant characteristics – skills, affiliations, departments and locations of identified common influential users – are discovered.
In section 2, the pioneer work in the domain of influence finding in ASNS is briefly reviewed. In section 3, the methodological framework is discussed. In section 4, the results are presented along with the detailed discussion and significant reasoning. Section 5 summarises this research with future expansion possibilities. The sample XML graph model of the prepared RG graph, the list of abbreviations and the list of notations are presented in Appendix 1.
2. ASNS and influence
The web-based ASNS are serving baseline mediums in achieving the competitive success in research sector [19]. In research community contextual and influence finding paradigms, ASNS are receiving intense attention. They offer broadly adaptable, efficaciously available and digitised medium to the prominent researchers of diverse disciplines and geologically scatter areas to share and access ad hoc, pre-print or full-fledged research and to collaborate, communicate and acquire assistance from scientific experts [1,3]. For such reasons, ASNS such as RG, Mendeley or Academia.edu claim to have several millions of recorded users [4,5]. Massive scientific data in relevant fields of research can provide pivotal information for sound scientific research and analysis; although, the present literature discloses a little interest in finding influential users/researchers based on their real-time research activities on any ASNS platform.
Each user in ASNS can possibly extend his scientific impact; however, a few users demonstrate their supremacy in scientific impact propagation [20,21]. Scientific influence finding is recognising these predominant (influential) users who can affect the assessments of different users by viral advancements or explicit activities. The recognition of such influential users is valuable for showcasing, promoting or spreading influence [22]. Identification of research domains, locations, department of work and so on of these influential users provides bright possibilities of communication and collaboration for emerging researchers [23]. The identified potential users with their specific domain of research interest can be invited for various research activities [24,25] such as article review, expert talk, conference speaker, multi-domain research collaboration and more [4,26].
2.1. Graph-based influence analysis in ASNS
Modelling social media information as a graph of connected data has its roots in many profound applications such as influence finding, topic analysis, clustering, expert finding, community detection and recommendation systems [22]. Graphs are demonstrated to be profoundly proficient for different social media data analytics [27–32]; among which, influence finding is preeminent.
Various ingrained social theories illustrate that the impact (influence) of one entity propagates to other connected entity and such influence can be effectively analysed using graph models [27]. For social media platforms such as Twitter and Facebook, the follower and following relations are proven to be substantial while analysing influence propagation and identifying influential entities. Furthermore, the technological advancements provide various competent mechanisms for graph storage and graph information retrieval [30]. Graph-based centrality measures [33] are widely employed for influence analysis in many application domains such as social media [34–37], web platforms [37–39], finance [40] and academic literature [41,42].
In academic literature domain, pioneer of work has been conducted in the field of finding influential researchers/scholars. Majority of such work utilises the available co-authorship [25,43–50] and co-citation networks [7,10–13,49,51,52] from graph data repositories. Apart from using the historic information of co-authorship and co-citation, identification of real-time influence by fetching the real-time linked data serves an establishment of developing potential research in academic literature analytic [53].
2.2. Related work
This research focuses mainly on analysing the impact of RG followers and followings on influence identification. The related work is discussed with respect to the existing studies that (1) compare RG with other well-known ASNS and (2) conduct influence analytics in RG.
2.2.1. Comparing RG and other ASNS
In literature, various analysis and surveys are conducted on well-known ASNS to identify their popularity in terms of total registered profiles, usage frequency and degree of activeness.
An analysis was conducted on the distribution of Spanish National Research Council researchers’ profiles on Academia.edu, GS and RG [15]. The study discovered the higher utilisation of RG in terms of number of registered profiles with the statistics of 4001, 2036 and 1156 profiles on RG, GS and Academia.edu, respectively. A study offering an overview of established and emerging ASNS – Scopus, WoS, PubMed, RG, GS, Academia.edu, Open Researcher and Contributor Identification (ORCID) and Publons – disclosed that RG is a widely utilised platform [18]. Profiles of 400 RG users were examined with respect to RG and h-index (GS and Scopus) [54]. RG was identified to be a leading ASNS in a way it employs researchers’ data.
A comparative overview of various ASNS, such as RG, GS, ORCID, Scinapse, Semantic Scholar, ScienceOpen and Kudos, disclosed RG to be superior among others due to its prominent status features [55]. A survey on Academia.edu, Mendeley, RG, MyScienceWork, Humanities Common, Social Science research Network, Profology and Trellis was carried out to learn their usage frequency [17]. The results showed substantially more frequent usage of RG among all with 85% of users using RG least monthly. Another survey conducted on Academia.edu, Mendeley and RG to measure the degree of activeness revealed RG receiving the greatest attention [16]. An assessment of researchers’ presence was performed with respect to four platforms: ORCID, ResearcherID, Academia.edu and RG [56]. The analysis disclosed that out of 1047 researchers studied, RG had by far the highest number of researchers (54.3%) present.
2.2.2. Influence analytics in RG
RG facilitates researchers to generate their profiles, share their research [26], discover the peers with expertise in relevant domains [20] and get insights into inclining research [23,24]. In existing studies, majorly the profile features present on RG are analysed statistically to identify the influence with different prospects. Few studies also explore the co-author networks of RG to identify influential RG users.
A study was conducted to explore the real-time RG publications from Canadian computer science researchers to analyse the correlations among diverse profile features of RG including views, publications, downloads, citations, questions, answers and followers [57]. The research investigations showed that the Canadian researchers are highly active in research associations and scientific information sharing while the statistics illustrated the moderate correlations among various RG profile features such as number of views, downloads and citations. The number of views was around 2 (1/2) times the number of downloads and the number of citations was around 2/3 times the number of downloads for the collected data. An analysis was carried out to inspect the effect of RG statistics on institute ranking [58], in which, the results unveiled moderate correlation.
RG network of Macedonian University was explored for specific research areas [59]. Centrality measures were used to find the influential researchers with multi-domain knowledge. A cluster analysis was performed on RG based on various RG profile features [60]. Three clusters – Active users, Representers and Lurkers – were created. The results disclosed that in the collected data, RG users’ distribution in respective clusters was 8%, 4% and 88%. An analysis was conducted on RG users from 61 US research universities at different research activity stages (categorised by the Carnegie Classification of Institutions of Higher Education) to inspect the impact of institutional differences on RG reputation metrics [4]. The results confirmed that RG is a research-oriented academic social networking site that thoroughly and convincingly mirrors the research activity level of institutions. RG helps in maximising the influence for emerging researchers by increasing RGScores, publications visibility, citations, profile views and followers.
Sample of 4800 RG profiles was analysed to identify the communication patterns and demographic attributes [61]. The influence of age, status and institutional rankings was measured against the researchers’ network activities. The outcome suggested that age or status had no influence on network activities of researchers, whereas the institutional rankings demonstrated the influence on network behaviour. A study to examine the interrelations of scientists’ network communications in RG and their real-life proficient accomplishments was performed in [62]. The results demonstrated that the researchers who were not actively involved in RG network activities had little privilege on the role of academic influential person.
A study was conducted using network science to create RG collaboration network of Iranian Scientific Institutions [63]. The results demonstrated that geographic location closeness and ethnic attributes have influential roles in academic collaboration network establishment. Also, the popular scientific centres in the capital city of Iran, Tehran, have a significant influence on the production flow of scientific activities. A sample of 77,902 RG users from 61 US research universities at different research activity levels was analysed [64]. The sample users were categorised into six groups based on their affiliated departmental disciplines as stated on their RG profiles. The results showed that user participation and user characteristics vary by discipline. In addition, users from higher research activity level universities are identified to be influential in RG metrics regardless of discipline. A study investigated the characteristic usage behaviours of Essential Science Indicator (ESI) highly cited researchers on RG [65]. The results showed that 45.21% of ESI researchers from diverse disciplines have registered with RG. The Spearman correlation denoted that information sharing behaviours have a strong correlation with researchers’ influence compared with social interaction behaviours.
3. Methodological framework
The recent surveys demonstrate RG as greatly utilised ASNS among others, but it is yet to be explored with respect to influence analysis in terms of followers and followings present on RG using real-time RG information.
The significant contributions of this research are as follows:
Render real-time linked information of RG based on follower and following connections
Formulate collected linked information into a graph model and implement graph-based centrality measures on the constructed RG graph
Measure the relevance of obtained results with respect to RGScore and TRI
Develop a method for collective influence finding and explore the properties of commonly present influential RG users in centrality-based top-k results
The proposed methodological framework for anatomising the impact of RG followers and followings constitutes five sub procedures: data rendering, graph modelling, graph construction, implementation of centrality measures and collective influence finding as demonstrated in Figure 1.

The proposed methodological framework.
3.1. Data rendering
The data rendering process contemplates following qualitative observations.
3.1.1. Observation 1
For the known influencing researcher R in RG having F number of followers and G number of followings, let the influence score of f in F and g in G be I
scf
and I
scg
, respectively. For
3.1.2. Observation 2
In a given time span tp, unlike social media networks, research networks tend to undergo a smaller number of changes in terms of content and links. Henceforth, this research does not include the concept of evolving network. It rather takes into account the real-time RG information at a specific time t.
While for many ASNS, the co-authorship and co-citation graphs are available, it is challenging to collect real-time linked RG information. Unlike social media platforms such as Twitter, RG does not provide any Application Programming Interface (API) but the structured organisation of information present on RG can be rendered using the ‘web rendering’ concept. A rendering process, specifically designed and accurately correlated with the structure of RG, is developed. It renders publicly accessible user content and user-to-user associations. Here, user content and user demographics, user-to-user links and associations words are used interchangeably.
The rendering process has two sub functions: ContentRender and LinkRender. To demonstrate the working of both the functions, consider a sample RG network as displayed in Figure 2. For any selected target user, U, U1, U2 and U3 are the first level of followings (displayed as FW1). Based on the Observation 1, for a selected target user U, only the followings are rendered in first phase. For the first level of followings of U, the followers (depicted as FL2) and followings (denoted as FW2) are rendered in second phase. The follower and following links are rendered using LinkRender, whereas for every user including U, the profile content is rendered using ContentRender.

Sample RG network.
The algorithmic steps of the developed data rendering process are demonstrated in Algorithm 1. Line number 3–10 demonstrate the working of ContentRender, while line number 11–18 depict the functioning of LinkRender. The rendered content is stored in file C and the rendered associations are stored in file R.
Data rendering.
FW: Following of Node N; FL: Follower of Node N.
3.2. Graph modelling
To transform the rendered RG information into graph structure, a graph model is prepared. Employing the prepared graph model presented in Algorithm 2, the collected information is transformed into nodes, relations and properties forming a connected RG graph.
Graph modelling.
As displayed in Algorithm 2, line number 1–4 depict that for each content row in input file C, the nodes are modelled. Each rendered property is attached with the respective node. Line number 5–7 denote that for each row in input file R, the relations are modelled among nodes.
3.3. Graph construction
The set of nodes, relations and properties are imported into Neo4j - the leading graph database - to construct RG graph. Cypher Query Language (CQL) is used to interact with the graph in Neo4j. The pseudo-code of the graph construction process is demonstrated in Algorithm 3. Line number 1 displays the initialisation of Neo4j, line number 2–4, respectively, show the import process of nodes, properties and relations in Neo4j.
Graph construction.
N: number of nodes; R: number of relations.
Each ‘node’ in the RG graph represents an RG user while their associations are modelled as ‘directed edges’ with label ‘Following’ and ‘Follower’. N nodes are created and R relations are formed among nodes based on rendered follower and following information.
3.4. Centrality measures
To grasp a comprehensive impression on rendered RG information, we conducted a study to determine which users are at the ‘centre’ of the whole graph; therefore, graph centrality measures are implemented.
The notion of finding influential nodes is subjective depending on how the ‘influence’ is defined. Different centrality measures use different metrics which define the importance of a node from different perspectives. More specifically in this research, four standard centrality measures – BC, CC, EvC and PRC – are adopted and implemented in iterative manner on the graph to check their performance in finding influential users. Each centrality measure computes the list of top-k influential users as mentioned in Algorithm 4. Line number 1–4, respectively, demonstrate the calculation of BC, CC, EvC and PRC for each node in the constructed RG graph. The obtained results are discussed in section ‘Experimental analysis and results’. The evaluation of implemented centrality measures is performed in terms of ranking correlations with respect to RGScore and TRI.
Centrality measures.
N: number of nodes; R: number of relations; TBc: top-k list derived from betweenness centrality; TCc: top-k list derived from closeness centrality; TEvc: top-k list derived from eigenvector centrality; TPRc: top-k list derived from PageRank centrality.
3.5. Collective influence finding
To expansively examine the landscapes of influential users given by each centrality measure, a new list Tcom having collective results is generated as displayed in Figure 3.

Process to find collective influence from centrality-based top-k lists.
Here, Tcom is a set of distinct influential nodes ni which are commonly appeared in TBc, TCc, TEvc and TPRc with occurrence threshold
Collective influence finding.
TBc: top-k list derived from betweenness centrality; TCc: top-k list derived from closeness centrality; TEvc: top-k list derived from eigenvector centrality; TPRc: top-k list derived from PageRank centrality.
The skills, affiliations, departments and locations possessed by the identified influential users in Tcom are further derived and analysed in this research. These results are presented in section ‘Experimental analysis and results’.
4. Experimental analysis and results
A summary comprising details of implementation setup is shown in Table 1.
Implementation setup.
BC: Betweenness Centrality; CC: Closeness Centrality; EvC: Eigenvector Centrality; PRC: PageRank Centrality.
The rendering process renders the profile content of 1544 RG users. For each user, 18 profile properties are rendered. In total, 27,792 (1544 × 18) properties and 1646 connections present among 1544 users are rendered. Table 2 shows the 18 rendered properties with the assigned IDs.
List of rendered properties.
The visual representation of the constructed RG graph is shown in Figure 4. Each RG user is depicted as a node. The rendered 18 properties of each user are stored with each respective node. The rendered follower and following connections are displayed as directed edges with labels ‘Follower’ or ‘Following’ among connected nodes.

Performance of BC for top-k RG influential users. (a) Performance of BC for top-10 RG influential users.(b) Performance of BC for top-20 RG influential users. (c) Performance of BC for top-30 RG influential users.

Visualisation of the constructed RG graph.
The performance of standard centrality measures is evaluated based on how efficiently these measures identify top-k influential users from the constructed RG graph. The experimentation is performed taking three values of k with identical intervals,that is, k = 10, 20 and 30. Four centrality measures are implemented on the graph in iterative manner, and eventually, four lists of top-k influential users are recorded. In Table 3, top-k (k = 10) influential RG users received by centrality measures are presented. Here, Rank 1 denotes the highest influence. The commonly found users in all four centrality top-k lists are highlighted in Table 3.
Identified top-k (k = 10) influential users in RG graph.
BC: betweenness centrality; CC: closeness centrality; EvC: eigenvector centrality; PRC: PageRank centrality.
To identify the relevance of RG followers and followings with respect to RGScore and TRI, top-k lists generated by centrality measures are compared with the top-k lists computed by RGScore and TRI.
Table 4 shows the sample ranking tuples generated for every centrality with RGScore and TRI for top-k results. The generated ranking tuples disclose the degree of relevance between two ranking lists.
Ranking tuples notations for top-k.
RG: ResearchGate; TRI: Total Research Interest; BC: betweenness centrality; CC: closeness centrality; EvC: eigenvector centrality; PRC: PageRank centrality.
The generated ranking tuples for

Performance of CC for top-k RG influential users. (a) Performance of CC for top-10 RG influential users. (b) Performance of CC for top-20 RG influential users. (c) Performance of CC for top-30 RG influential users.

Performance of EvC for top-k RG influential users. (a) Performance of EvC for top-10 RG influential users. (b) Performance of EvC for top-20 RG influential users. (c) Performance of EvC for top-30 RG influential users.

Performance of PRC for top-k RG influential users. (a) Performance of PRC for top-10 RG influential users. (b) Performance of PRC for top-20 RG influential users. (c) Performance of PRC for top-30 RG influential users.
To identify the correlations among the generated ranking tuples, rank correlation co-efficient is computed. The results are demonstrated in Figure 9.

Ranking correlation measures. (a) Ranking correlation for top-10. (b) Ranking correlation for top-20. (c) Ranking correlation for top-30.
It is evident from Figure 9 that for both RGScore and TRI, BC performs efficiently while EvC and PRC deliver inefficient results for the rendered data and constructed RG graph. Higher correlation in the case of BC suggests that high number of RG users depends on one another to make connections. Wider set of RG users in the constructed graph has many connections that are very crucial in the propagation. EvC and PRC achieve lower correlation which depicts overall less number of RG users in each user’s research network. It is observed that for k = 10 and k = 20, BC outperforms, while for k = 30, CC achieves higher correlation when compared with both RGScore and TRI. Overall, CC performs efficiently over EvC and PRC, which suggests in the constructed graph, a moderate set of RG users has ties to several active or influential users. Overall, as the k in top-k increases, the performance of EvC and PRC decreases. BC and CC perform stable with increasing k values, demonstrating BC as the most relevant to both RGScore and TRI for the constructed RG graph in influence identification and ranking yielding the average relevance of 67% with RGScore and 40% with TRI.
4.1. Collective influence list
Based on the collective influence finding method presented in Algorithm 5, for the influential users present in each top-k list of four centrality measures, the users with occurrence threshold of at least 1 are fetched. The properties of such derived 59 users are examined in detail to identify their skills, degree affiliations, departments and locations.
Figure 10(a) represents the influential skills of users present in Tcom list and having occurrence count

Property analysis of top-k RG users in Tcom list. (a) Skills. (b) Degree affiliation. (c) Department. (d) Location.
4.1.1. Findings
The skills of identified common influential RG users are examined in-depth to gain knowledge about in which research fields the influential researchers are working. It is recognised that areas, such as Macroeconomics, Development Economics, International Economics, Financial Economics and Econometrics are highly significant with occurrence count
In the computed Tcom list, majority influential users possess PhD in ‘Economics’ with
The influential university departments found in Tcom are Department of Economics with
The observed influential locations in Tcom list are İzmir and Nevşehir with
5. Summary and future work
The proposed methodology provides a novel approach to treat RG information as a connected graph of users in terms of followers and followings. The aim is to anatomise the impact of RG followers and followings on influence identification. As the already established graph of RG data is not available, research demands an efficient method to collect linked RG information to form RG graph. The methodology contributes to offer a data rendering process to fetch linked RG information. From the fetched information, RG graph is constructed. Standard centrality measures – betweenness, closeness, eigenvector and PageRank – are implemented on the constructed graph for influence identification. The performance is compared with RGScore and TRI. The results reveal that BC and CC perform efficiently on our constructed RG graph. This research is an initial attempt to analyse RG information as a complete and connected graph by implementing centrality measures with an aim to study the impact of follower and following connections on influence identification.
ASNS data undergo less frequent changes unlike social media data; hence, this research provides a static pattern that does not consider the change in centrality measures over a time at this stage. We aim to consider the effects of the topological changes on influence in our future work. The dynamic changing patterns in the graph will be helpful to study symmetry of the complete RG graph and grasp deeper insights into influence change over a time. In addition, the centrality measures applied on RG graph do not consider the statistical measures of RG such as number of research items, citations, reads, etc. In the future, centrality measures and RG statistics can be assembled to identify the influence more accurately.
Footnotes
Appendix 1
List of notations.
| U | ResearchGate User |
| N | Number of Nodes |
| N i | ith node |
| P | Number of properties |
| P i | ith property of node N |
| R | Number of relations |
| FL | Follower of node N |
| FW | Following of node N |
| I scn | Influence score of node n, where |
| TB c | Top-k list derived from betweenness centrality |
| TC c | Top-k list derived from closeness centrality |
| TEv c | Top-k list derived from eigenvector centrality |
| TPR c | Top-k list derived from PageRank centrality |
| T com | Collective top-k list derived from centrality results |
| t o | Occurrence threshold |
| C o | Occurrence count |
| S i | ith skill of node N |
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
