Abstract
By applying different clustering algorithms, the author strived to construct the best visual representation of scientific domains and disciplines in Poland. Journals and their disciplinary categories constituted a data set. A comparative analysis of maps was based on both qualitative and quantitative approaches. Complex patterns of eight maps were evaluated taking into account both the local proximity of disciplines and the whole structure of presented domains. Final clustering quality value was introduced and calculated in reference to the knowledge domains. The authors underlined the role of quantitative and qualitative methods in combination in the mapping evaluation. The best results were obtained with the T-distributed stochastic neighbour embedding (t-SNE) algorithm. This youngest technique may have the biggest potential for semantic information studies and in the scope of broadly understood semantic solutions.
1. Introduction to science mapping
Sophisticated visualisations play an important role in science. They create opportunities to discover new research trends; specifically, they elucidate transdisciplinary frontiers that are difficult to describe quantitatively. Such visual representations can show newly emerging scientific fields, including where and how they overlap to give rise to new areas of knowledge. They can also show where areas of knowledge disappear by revealing unpopular research disciplines. There are several terms referring to information visualisation used in and for science. Most often, the process is called ‘science mapping’ and results in appropriately, maps of science [1,2]. Synonyms for ‘map’, such as ‘landscape’, appear in the subject literature, aiming to outline special correlations and similarities between data as well as to draw nascent conceptualisations [3]. These maps can present the scientific community and their mutual relationships (through citations or co-authorships) or scientific fields and disciplines and their interactions. Aside from the many known contributions of the great bibliometrician Eugene Garfield to the current science of science (SciSci) studies, his concept of scientography – or the graphical overview of the history and state of a given body of research – should be noted as the first turn towards visually representing science for qualitative analysis [4,5]. As an analogue to scientography, the concept of sociography refers to social structures and can be successfully used in the visual analysis of the structure of an academic community. Two main streams of recent science-mapping research can be distinguished, which are as follows: (1) studies oriented towards scientist collaboration and (2) research considering science organisation. As a result, social structures (i.e. sociograms) or science domain structures (domain analysis) can be studied at different scales and levels [6]. Chaomei Chen coined the term ‘knowledge domains visualisation’ (KDViz) as a synonym for ‘mapping science’ [1].
Science maps are the product of the last stage of the visualisation process [7], and they simultaneously initiate the analytical cycle while they are explored, read and comprehended. Ultimately, these infographics are admired for delivering new knowledge about research areas, centres, institutions, and the science community, depending on the selected units of analysis. For analysis, metadata (journal titles, authors, keywords of articles, categories, and field descriptors) and citations are used. The relationships between them are obtained by counting co-occurrences. Then, multiple data sets can be mapped on a graph or output space, which requires dimension reduction. Nees Jan van Eck and Ludo Waltman identify these two approaches as ‘distance-based’ and ‘graph-based’ mapping [8]. Research areas or interests are formed on maps through local dense concentrations (or linkages). The visualisation’s patterns determine the conceptual structure of a research field or discipline and even a knowledge domain.
The process of science mapping starts from the conceptualisation of a research problem and a strategy for data gathering [7,9]. Where to obtain representative data for mapping is one of the basic challenges for scientists, who strive to investigate and measure phenomena arising in the world of digital information. For data retrieval and collection, it is acceptable to use global bibliographic databases, such as the Web of Science (WoS), Scopus or Google Scholar. However, for the visual representation of scientific trends and developments, researchers have recently been drawn to choose national, specialist or secondary indexes, such as CrossRef, Association for Computing Machinery (ACM) Digital Library or computer science bibliography database systems and logic programming (DBLP) [10,11]. In terms of completeness, they are not as representative as global indexes are, but they can provide comparative studies at a specific level.
In Poland, no complete, complementary bibliographic database that represents national scientific work exists. One aspiring to be such is the Polish Scientific Database (Polska Bibliografia Naukowa or PBN), which is in the development phase [12]. The PBN is updated by submissions from Polish scholarly units and also from the authors of publications. As the process is optional, the database’s complementarity is not uniform [13]. Another problem is that the citations database POL-index is also just being developed. It is part of a larger, national academic system POL-on still being elaborated by the Ministry of Higher Education. These issues and uncomplementarity of data sources makes it difficult to track and analyse citation patterns in Polish scientific literature at a level representative of the whole country for most research domains. Therefore, we are forced to look for alternative solutions.
Academic journals published for a selected community, such as those in a nation, can reflect the needs for specialised scientific knowledge. In particular, the accumulation of this knowledge within a basic domain space can be observed on science maps. As de Solla Price [14] conjectured, a journal database should contain the structure of science. Currently, researchers reveal this phenomenon in the subject categorisation of journals used [15]. Polish journals have increased quickly since the 1970s, achieving a critical number of journals per scientific unit [16]. As the total number of journals became stable (4500), they presented a valuable source of scientific knowledge for organisations and research.
The author used the professional database of Polish journals, organised into several scientific disciplines. This categorisation became the basis of the similarity metrics used in this study. Due to the lack of a unified bibliographic data, the resulting disciplinary mapping of journals presented here can be used to map the Polish knowledge landscape. This landscape illustrates the deliberate organisation of scientific knowledge by revealing the grouping of fields into main domains, current trends in research and technologies and the integration of similar areas and specialities. This landscape can then be the basis for visual analyses of present science structures and future changes.
2. Brief literature review of mapping algorithms
Recently, the concept of ‘Big data’ has been thoroughly established in science, forcing a methodological turn in scientific writing analyses and influencing scientometric frameworks. There are two main reasons for this phenomenon. First, there has been a dramatic nonlinear increase in authorship and a quick growth in publication output, creating substantially more papers [17]. Second, machine-learning algorithms and artificial intelligence applications used in big data solutions have seen rapid development in the last two decades. Tens of millions of scientific documents to be analysed can be processed and then clustered or classified. Clustering routines can partially or fully solve the main problem of a large data set visualisation and dimension reduction into two-dimensional (2D) or three-dimensional (3D) output spaces. This becomes the basis of science mapping that can identify research topics, scientists and collaboration paths. The scale of mapping studies depends on many factors, such as big data accessibility, familiarity of algorithms, experience and others [18].
The principle of layout construction is to place similar objects in N-dimensions (where N = 2, 3) so that they are close to each other and dissimilar objects are far apart [1]. So called distance-based maps are created by algorithms that minimise the distance between analysed units. A smaller distance means a stronger relationship between them (semantic similarity) and inversely, very distant objects have weak (or no) relationships. Mapping algorithms usually generate the structure of clusters that are overlapping, showing discovery patterns and trends in science.
The popular technique for dimensionality reduction for science domain analysis is multidimensional scaling (MDS), which is a way to rearrange objects in the configuration that best approximates the observed distances. It actually shifts objects in the space defined by a requested number of dimensions and controls how well the distances between the objects can be reproduced by a new configuration. MDS has been widely applied in constructing knowledge maps of authors, articles, journals and keywords [8,19]. Such visualisation maps can be used as information retrieval interfaces [11]. In the ranking of dimensionality reduction techniques, principal component analysis (PCA) or factor analysis is second in popularity after MDS in aiming to extract data in multiple categories and achieve a more qualitative typology [1,15].
In 2010, a new popular mapping technique, visualization of similarities (VOS) was introduced in the user-friendly application VOSViewer 2 (vosviewer.com) by Van Eck and Waltman [8, 22]. It was intended as an extended version of MDS, because VOS and MDS have a close mathematical relationship [20].
Graph-like maps of science (where data interconnections are indicated by links) rely on mapping techniques – the list of which is rapidly growing due to the spread of social network analysis (SNA) applications in the social sciences [21]. Chen [1] first introduced the pathfinder network scaling technique, which has been further tested in several projects [22,23]. Popular network layout algorithms described in the subject literature network include Kamada-Kawai and Fruchterman-Reingold [24,25]. Researchers can manipulate graphs by selecting different layout algorithms in network modelling software, such as Pajek, Gephi and Cytoscape [26]. SNA measures are especially useful in developing clustering algorithms for Big data. Waltman and van Eck [27] introduced modularity-based clustering techniques, which are used for topic clustering in large document sets that can reach tens of millions of documents [28].
As quickly as artificial neural network techniques develop, they are implemented for science visualisation. Self-organised maps (SOMs), or Kohonen networks, aim to produce a 2D representation of the input data using a neighbouring function. White wrote that these maps are ‘in effect the same map, differing mainly in matters of nuance’ [29]. SOMs are used less frequently than, for instance, MDS. Perhaps, this is because they have few instances of co-occurrence data, while Kohonen networks require a large input data vector to generate a qualitative map [30]. SOMs are often used in comparative studies of mapping algorithms [31], which typically conclude that MDS is distance preserving while SOM is a topology-preserving technique [32].
In the face of manifold mapping techniques, there is a need for evaluation of the results. One approach is the comparison of maps produced by different algorithms or based on different analysis units. In quantitative terms, this means a comparison of cluster accuracy that is one of the basic research problems in contemporary science mapping [28]. The scientometrician’s watchwords ‘same data, different results’ reflect the main shortcoming of science-mapping frameworks [18]. The metadata of bibliographic collections produce a wide spectrum of combinations for the examination of different analysis units. Direct citation maps can be compared with co-citation or bibliographic coupling links. The patterns generated by textual characteristics can be compared with co-citation layouts [28]. So, with different mapping techniques and parallel metadata, the visualisation maps lead to complex and multi-perspective study of scientific topics and communities.
There are many questions posed by science-mapping professionals. For example, how do distinct approaches affect the results [18]? How does one select the proper approach? In other words, which approach represents the ‘true’ scientific structure (and not artefacts)? How does one reproduce the same results in scientific studies [33]? These issues motivate the search for common evaluation methodology with regard to relatedness measures. One recently published study uses a quality function of clustering and its scaling [34], and another applies a matrix of the intersection of clusters [23]. It should be noted that big data in scientometric mapping generates unique problems, such as cluster labelling that requires different levels of natural language processing (NLP). Based on the above solutions and considering the specifics of the data, this research introduced a mapping quality measure in reference to domain analysis.
3. Experiment, methods and data
3.1. Data
There are common journals based in Poland where researchers, particularly in the humanities and social sciences, directly publish output generally written in the Polish language. Indeed, scientific research and documentation cannot be limited to a native language, otherwise it will impede knowledge dissemination and the flow of ideas. Polish historians and linguists often have difficulty illuminating local and national research problems for foreign audiences. Technology, science and medicine tend to deal with more global issues relevant to the world at large and all people. Predominantly, the archives of international journals, particularly those with higher impact factors (IFs) are the target of such authors.
The database of Polish scientific journals Arianta (http://www.arianta.pl/) is curated by professionals and is systematically updated (N = 4338 as of June 2017). This is the largest continuously updated Polish journal reference source [13]. The database interface allows filtering combinations that can present the relationships to global indexes. In comparison with the WoS, one can note that 146 (3.3%) Arianta journals are indexed in the Science Citation Index (SCI) Expanded and 10 (0.2%) are in the Social Science Citation Index (SSCI). Common records with Scopus and ERIH PLUS equal 469 (10%) and 220 (5%), respectively. Therefore, as to the basic limitations of Arianta, Polish journals that serve mainly the humanities and social scientists have uneven disciplinary coverage in global databases. In addition, one should take into account that among the scientific journals, this database also contains popular science and professional magazines for engineers, managers, and medical professionals.
The author conducted previous research that analysed a journal collection extracted from the bibliographic database Expertus at Nicolaus Copernicus University in Toruń, Poland, which is the most complementary database at a local level [35]. Within Expertus, the total list of journals was N = 5829 (as of December 2018), 1858 (30%) records were not covered by Arianta. This ratio is considerable but does not take into account the poor representation of international journals in the institutional bibliography (less than 5%) [34]. We can say Arianta was the most representative journal database for university scholars at that time. Arianta is not just a simple list of Polish science and popular science magazines; the creators, Aneta Drabek and Arkadiusz Pulikowski, equipped it with a professional classification scheme based on Science Classification in Poland (SCP) according to the 2010 version [36]. It consisted of 171 categories, mainly derived from the WoS subject categories. Since 2018, the number of disciplines was reduced to 47. For analysis purposes, categories were grouped into domains (this is analogous to the WoS research areas).
Thus, the main benefits of Arianta as a selection for investigating mapping and layout were: largest national and/or local complementary database in terms of sources, systematic updating, disciplinary categorisation, curation by professionals and web accessibility. The following assumptions were made: (1) the majority of Polish journals are recorded in Arianta; (2) this set is the most representative for Polish universities (excluding technical universities and medical academies) and provides a large spectrum of research fields and specialties; and (3) at the same time, international scientific journals are covered to a different degree at each Polish university.
The creators of Arianta maintain the complete journal descriptions, in other words, publisher, disciplinary categories, Polish ministerial scores in different years, publishing frequency, indexes, coding, URL and so on. Data about the ascribed categories were gathered directly from the web site and sorted accordingly (Figure 1).

Web interface with examples of the journal’s descriptions.
Then, the data were sorted and arranged (a pre-processing step was not required because of the high quality of the downloaded data from a professional database) in order to create the journals-disciplines matrix; the resulting 4338 × 171 matrix was not symmetrical. Co-occurrence of categories defined the thematic/disciplinary similarity of the analysed units, journals. Compared with co-words, co-citations, and co-authors, co-categories as a unit of analysis are less frequently used in science mapping [15,31] because of their lack in a basic metadata set. Concerning digital libraries, on one hand, automatic classification still does not generate satisfactory results, but on the other, manual classification for large data sets may be too costly. For evaluation purposes, the predominant discipline versus one or more additional disciplines were identified in the data set and weighted. The ratio between them was regulated by various weights that were applied in mapping tests. Co-category distribution was very important later in the analysis as it turned out, and their properties are presented in Table 1.
Data characteristics for co-occurrence.
3.2. Methods
For the tabular data representation of the journals categories, two different methods of spatial visualisation were possible. The first focused on calculating distances between categories (columns) and the final arrangement of the set of classes on the plane or 3D space. Then, the journals set was positioned within a semantic space, where semantics refers to preserving the domain/disciplinary proximity. We called it a top-down approach, when the final visualisation reflects the organisation of the original classification on a predefined level, and therefore, replicates its taxonomical disadvantages [31]. The alternative approach was down-top, when the lowest-level hierarchy items, in other words, journals (rows) are the subject of the dimensionality reduction algorithm. They are clustered into areas that can be identified as single or multiple categories that do not necessarily coincide with the primary categorisation. Such visual patterns help to uncover new knowledge about the data and their structure. The authors mapped the set of journals categories by both approaches using four algorithms: MDS, spectral clustering, isomap and T-distributed stochastic neighbour embedding (t-SNE). MDS as a basic technique in SciSci research was selected as a preferential one, while the others, the relatively newer methods, had never been tested in information science applications. Two of them represent nonlinear dimensionality reduction (NLDR) variations.
3.2.1 MDS
2D Euclidean space coordinates for each of 171 or 4338 nodes were calculated using a Minimisation of the cost function called stress. The stress coefficient is a measure of goodness-of-fit, how well (or poorly) a particular configuration reproduces the observed distance matrix. It was calculated as[37]
where dij stands for the reproduced distances, δij stands for the input data, i = 1,…, 171 and j = 1,…, 4338. In this formula, dij stands for the reproduced distances, given the respective number of dimensions, and δij (delta ij ) stands for the input data (i.e. observed distances). The degree of correspondence between the distances among points implied by the MDS map and the matrix input by the user is measured (inversely) by a stress function. Generally, MDS enjoys high popularity in information science applications and, as it turns out, with no time-consuming computation.
3.2.2. Spectral clustering
The spectral clustering algorithm makes use of the eigenvalues of the similarity matrix of the data to perform dimensionality reduction before clustering in fewer dimensions. The similarity matrix is provided as an input and consists of a quantitative assessment of the relative similarity of each pair of points in the data set. As it applies to image segmentation, spectral clustering is known as segmentation-based object categorisation.
The spectral clustering uses a standard clustering method on relevant eigenvectors of a Laplacian matrix of A, where Aij ≥ 0 represents a measure of the similarity between data points with indexes i and j Laplacian matrix is defined as[38]
where D is the diagonal matrix
An algorithm of spectral clustering must be performed in three main steps: (1) the creation of a similarity graph between data; (2) the computation of k eigenvectors of a Laplacian matrix to define feature vectors; and (3) the execution of k-means clustering [39]. Spectral clustering has become increasingly popular in exploratory data analysis due to its simple implementation and efficient solving by standard linear algebra methods [40].
3.2.3. Isomap
Isomap is an NLDR method that extends the MDS metrics by incorporating the geodesic distances imposed by a weighted graph. If MDS is connected with the pairwise distance between data points, which is measured using a straight line in Euclidean space, then isomap exploits the geodesic distance induced by a neighbourhood graph. Intrinsic geometry of the data is measured along a low-dimensional manifold analogous to a ‘Swiss roll’ [41]. A constructed neighbourhood graph allows an approximation to the true geodesic path to be computed efficiently as the shortest path, for example, computed using Dijkstra’s algorithm. Next, the lower-dimensional embedding (such as MDS) is computed.
As noted by the algorithm’s creators: Our approach is capable of discovering the nonlinear degrees of freedom that underlie complex natural observations, such as human handwriting or images of a face under different viewing conditions. In contrast to previous algorithms for nonlinear dimensionality reduction, ours efficiently computes a globally optimal solution, and, for an important class of data manifolds, is guaranteed to converge asymptotically to the true structure. [41, p. 2319]
3.2.4. t-SNE
T-SNE is a technique based on machine learning developed by van der Maaten and Hinton [42]. It is an NLDR technique well suited for embedding high-dimensional data for visualisation in a low-dimensional space of two or three dimensions. The technique is an optimised variation of stochastic neighbour embedding [43] in which the conditional probability that ‘pj|i would pick xj as its neighbour if neighbours were picked in proportion to their probability density under a Gaussian centred at xi’ [42, p. 2581]
where σi is the variance of the Gaussian that is centred on datapoint xi. Then, the similarity of map yj and map y
i
is modelled by the condition of the Gaussian variance equal to
The cost function is given by
This method is based on the symmetrisation of the cost function C and uses Student’s t-distribution rather than a Gaussian one to compute the similarity between two points in the low-dimensional space [42]. The advantage of the t-SNE visualisation relies on reducing the tendency to crowd points together in the centre of a map. It better reveals structure at many different scales. This algorithm has a wide range of applications, including computer security research, music analysis, cancer research, bioinformatics and biomedical signal processing.
In the current work, the dimensionality reduction algorithms on the journals-categories matrix in both configurations (rows and columns) were used. The experiment’s phases were repeated for weights between the primary discipline and additional one(s) in the ratio 0.6:0.4.
Calculations were performed with a series of Python scripts on the Jupyter platform (http://jupyter.org/). The programme’s execution time was essential output information that was taken into consideration. Finally, multiscale data were visualised using a scatter plot with an interactive preview in Plotly. Interactivity provided the possibility of estimating generated patterns in terms of disciplinary similarity. Qualitative analysis was the dominant approach in the evaluation of the visualisation maps.
3.2.5. Domain clustering quality measure
Visual analysis was used for the rough evaluation of the final maps. The specifics of data units (i.e. disciplinary categories) were grouped into overriding categories (domains), which were identified by colours to facilitate the visual estimation. In present applications relating to domain analysis, we can assume that qualitative clusters would have no colour overlap (i.e. strict disciplinary journals) at their centres, but this would occur at edges (inter- and multi-disciplinary journals). Thus, mapping layouts can be divided by domain areas and clusters’ correspondence with domain uniformity defining a qualitative measure.
Suppose the map reveals N clusters grouped into D domains. It should be noted that typically in most science classifications, D is in the range (5–8). Let sij denote the population of cluster j qualified to domain i (i = 1,…, D and j = 1,…, N). One cluster can appear in data from different domains; it is preferable when the clusters are uniform as much as possible. Then, we can evaluate the final cluster-domain relationships by introducing the following quantitative function Q
where the global diversity parameter α is a sum of local diversity values αi defined as
Local diversity depends on the number of apparent ‘foreign’ domains li in the cluster ascribed to domain i and clustering consistency parameter k, which is defined experimentally based on the data structure and varies in the range (1–2). For compact, well-defined clusters, k = 1 is used, and for fuzzy ones, k = 2. All intermediate values indicate the clearness degree of clusters within every domain. If a domain distributes on several clusters, its special continuity is evaluated. If a distribution reveals no clusters, then it can be considered as one multi-category component, and Q will be an extremely small value regardless of the primary domain selected for the calculation.
4. Mapping results
Figure 2 contains the visual patterns generated by the four algorithms for both the weighted and unweighted disciplines of journals (N = 171). The algorithm’s running time is also noted for easier identification of the best performance. As happens in dimensionality reduction algorithms, both axes represent arbitrary units. As most statistical and machine-learning techniques allow input data vectors to be constructed from the columns (categories) or rows (journals), these multi combinations can shed more light on hidden relationships between data in the context of SciSci studies.

Mapping results for unweighted (a) and weighted (b) disciplines made by the four algorithms.
As we expected, weighting slightly increased processing time by several hundredths of a second. The most time-consuming mapping algorithm was t-SNE. Journal clustering, when the disciplines were unweighted and weighted, accordingly, are shown in Figure 3.

Mapping journals classified according to unweighted (a) and weighted (b) disciplines made by the four algorithms (colour versions with better resolution are available at: http://wizualizacjainformacji.pl/journals_maps).
The same conclusion can be made when comparing the running time for both processing units; t-SNE was the most time-consuming technique, and weighted categories required a bit more time than unweighted.
5. Mapping evaluation and interpretation
The map of categories representing the final thematic grouping reduced the dimensions of the data and thereby, can help in redefining selected scientific fields. The effect should resemble PCA or factor analysis objectives, discovering new factors reflecting changes in science and scientific research. This is a top-down approach, because the analysis starts with the top-level categories. The down-top approach means that the clustering relates to the journals set. In cognitive terms, this can be more surprising because the clustering effects can diverge from an imposed hierarchy.
From Figure 2, we can see that isomap was not an appropriate method for data representation, probably because of the fewer degrees of freedom in the input vector. The output layout shows a one-dimensional distribution (on a curve) regardless of the weights and must be excluded from further consideration. Spectral clustering led to an excessive squeezing of the data points. MDS and t-SNE revealed uniform (no clustering) and loose distributions that eliminated them from consideration in further ranking.
In the case of journal mapping (Figure 3), the spectral clustering can be eliminated because of an unrepresentative output space that appeared to be data loss. Spatial differences between data points were negligible (they appeared at the 17th decimal place). The weights between categories were selected empirically. It was found that an optimal visualisation response of manipulation by weights must be uniform with no critical changes [31]. Therefore, the next evaluation criterion related to the stability of the graphical distribution at the transfer of the unweighted–weighted data. This property, stability in distribution, can be primarily estimated by visual inspection. It is possible to note that point configuration preserved a relatively stable structure on three maps: t-SNE, isomap and MDS. In the case of t-SNE, weighting made clustering more intensive; MDS weights reduced the number of concentric rings, and isomap decreased scattering. Both maps represent clearly designated clustering; therefore, in the final selection, we will focus on these two instances.
The author performed a visual analysis of the graphical distribution by scanning a predefined viewport 1 through which one could seek any neighbouring points that thematically did not match each other. This way, we could locally test the quality of the visualisation maps. At least one found case of clustering inaccuracy was enough to decide about the disqualification of the examined approach. Figure 4 shows an instance of bad local semantics and overall clustering for the MDS algorithm.

MDS mapping of journals classified according to unweighted (a) and weighted (b) disciplines. Examples of catching incorrect clusters: A –‘Science Philosophy’, B –‘Geopolitical Review’, C –‘Aeromechanics’ and D –‘Psychological Review’, E –‘Technology Engineering’ and F –‘Miscellanea Geographia’.
Figure 4 illustrates examples of disciplinary mismatch for just two MDS clusters. So, for the purposes of current discipline mapping, MDS was not a proper technique. These qualitative (not machine-based) tests showed that among the three algorithms, we could trust t-SNE mapping. At this time, it is justified in applying an evaluation framework with the clustering qualitative measure Q [7].
A new SCP [36] (since 2018) group disciplines into eight domain areas unambiguously: humanities, social sciences, medicine and health sciences, exact and natural sciences, engineering and technology, agricultural sciences, theology and art. Thus, operating by the primary discipline of the Arianta classification, it was possible to extract information about the knowledge area from the discipline–domain relationships. The maps were coloured according to the domain attribute; the results are presented in Figure 5 and with better resolution on the web portal (http://wizualizacjainformacji.pl/journals_maps). To calculate the Q cluster borders, they should be defined primarily. Due to irregular shapes, they were drawn directly on visual layouts; finally, they were divided into distinct slices indicating domains. As Figure 5 shows, in the t-SNE case, the process was maximally simplified. Table 2 presents comparisons of the three mapping algorithms with two combinations each.

t-SNE visualisation of journals classified according to unweighted disciplines.
Clustering quality measure regarding domain distribution for the three algorithms.
MDS: multidimensional scaling.
The best performer, t-SNE, as shown in Figure 5, provided the spatial continuity of domain areas and a qualitative designation of clusters. The clusters are labelled with reference to primary discipline, and as can be observed, a large coherence within clustering organisation was obtained. It is worth noting that unweighted vectors resulted in increasing quality in one case, that of t-SNE. The weighted version illustrated the distribution with clear clusters on the peripheries and a dispersed set in the centre. On the contrary, MDS and isomap, which revealed only a medicine cluster, kept it only for weighted data. It is probably that t-SNE favoured the structure of data with a distribution as presented in Table 1.
An interactive version of the t-SNE map would allow examination of how clustered journals were similar to each other. The colour distribution presents uncommon and logically proximal areas of knowledge that are usually difficult to realise because of input data requirements (large-scale and/or qualitative information from the one side and dimensional limitations of a plane from the other).
The positive evaluation qualifies the t-SNE layout for the Polish science map, which becomes a valuable source of specific, important conclusions on how it is functioning and how scholarly communication is organised at a national level. They are presented below:
Contemporary medicine draws on cooperation with science and engineering, which is reflected in the literature.
Computer science issues are diffused among exact sciences that are associated with computing and engineering (e.g. industry and technological processes).
Problematics of natural sciences are close to those of medicine and engineering and technology.
Natural sciences are divided into two groups: biological-based, associated with medicine and social sciences, and geological ones, applying mostly to engineering and technology.
Single cluster assemblies are forestry, agricultural and veterinary, which are not marginal but centred within engineering and natural and exact sciences.
Art is characterised by a similar isolation, but within the humanities and social problematics, which means that up until now, there has been little to no Polish writing about modern art using ICT.
Social sciences are the majority of Polish scientific journals.
The theology cluster arranges at the periphery and simultaneously within the humanities field. Its appearance in the newest Polish classification is groundless.
Despite LIS being connected with social sciences (since 2018) and its roots in the humanities, it is nearer to social issues; it also has close ties with culture studies.
6. Summary and discussion
The qualitative phase included a rough selection based on a visual analysis of the data distribution. There were two perspectives of examination due to data vector projections to both categories and journals. The categories mapping (top-down) using the proposed algorithms did not show any regularities that one could focus on. However, the journals clustering, the down-top approach, brought added value to insightful science study. Manipulating by weights between primary and additional disciplines ascribed to journals allowed us to observe how different algorithms behaved while the data vector was slightly changed. The weighting caused the mapping patterns to change in different directions. MDS, the first of the three selected algorithms, reduced the concentric rings into two, and similarly, t-SNE reduced the number of clusters, while isomap generated one central cluster. The colouring of maps according to journal domain illustrated clustering well and allowed for inferences about the clustering results.
Evidence of this quantitative analysis ranking was added, based on equation (7), which consisted of diversity value (8), reflecting the cluster–domain relationships. To perform the calculations, the clusters having irregular shapes were drawn manually using graphics software. Herein, the visual analysis, graphical processing, Q parameter calculation and qualitative insight (i.e. expert interpretation of final visualisation) constituted the methodological framework of the bipartite evaluation.
This SciSci research, despite the uncertainty, derives benefits from domain visualisation in imaging the current state and possible development of science. Mapping well in terms of knowledge domain visualisation (KDViz) means that scientific fields and disciplines are organised into more or less outlined clusters that reflect semantic similarity between them [1,31,43]. However, the visual representation of scientific knowledge cannot be even and linear because it is dynamic and inherently nonlinear in nature. Distribution therefore should not be so dense so that separate categories can be distinguished and be characterised by some degree of lacunarity [44] in order to leave enough space outside of the cluster for potential growth of research areas. Therefore, the utility-based recommendations for mapping consist of uneven distribution, clearly drawn clusters, and no local crowding.
When visualisation represents whole knowledge domains using a 2D area, the neighbouring, in particular at the aggregation level, defines the logic of the domain organisation. Aside from local conditions, there is overall structure. Thus, an area related to many others (i.e. multi-disciplinary fields) can be expected near the centre of the layout with more independent fields at the edges. Authors noted that both conditions are met in the t-SNE clustering, whose result is shown in Figure 5. Consequently, one can formulate important features of a 2D map: (1) grouping into clusters and (2) semantic proximity of neighbours. Additional dimensions on 3D maps are a distinct, poorly discussed research and engineering problem; there are still too few projects relating to 3D science visualisations in the form of, for instance, 3D graphs [31].
It is worth mentioning that the journal set was investigated under the assumption that it was representative of Polish science. However, if we take into consideration that the analysed journals did not entirely cover science, medicine and engineering because of mostly international publishing, the data set should be extended by a proper selection of what is planned in further works. The question of whether resulting patterns reflect the real organisation of Polish science becomes a separate research problem, which can be referenced to similar works [45].
In 2018, the science classification table in Poland was updated [36], and consequently, the change of domain grouping was introduced on the map (Figure 5) and confirmed by a final assessment with experts. Due to the visual representation, we can evaluate current modifications of organised categories. The given layout can serve as a basis for development work on classification structure.
7. Conclusion
Through science mapping using different algorithms, the authors sought the best scientific domain representation in relation to SCP. The article depicts the comparison of visualisation patterns taking into account both the local proximity of disciplines and the overall, rational structure of the domain. Qualitative analysis was enriched with quantitative evaluation based on calculating the clustering quality Q in reference to knowledge domains. The mapping of science always contains uncertainty in formulating laws and dynamics [46]. However, exposed visualisation maps led to conclusions on how scholarly communication at a national level functioned and was organised.
The best results were obtained with the t-SNE algorithm. This newer technique is well known from bioinformatics and biomedical applications for visualising Big data. Hopefully, the presented approach will initiate the use of the promising t-SNE method in information science within the scope of broadly understood semantic solutions.
Contemporary SciSci provides ‘insights into the conditions underlying creativity and the genesis of scientific discovery, with the ultimate goal of developing tools and policies that have the potential to accelerate science’ [46]. The author underlines the role of quantitative and qualitative methods in combination for mapping evaluation. They both are essential in multi-contextual and multi-disciplinary analyses, which can be useful for information scientists, classifiers, science analysts and policy makers, scientometricians and all that are inquisitive about the overall development of scientific knowledge.
In the face of the popularisation of science mapping, we should pay attention to the reality that researchers are becoming the direct users of graphical layouts, which are increasingly interactive. Analogous to user experience (UX) as applied in web design, similar evaluation practices applied to visualisation should be introduced. Such utility-based recommendations for mapping, resulting from this study, would consist of uneven distributions, clearly drawn clusters and no local crowding. Naturally, we cannot forget about science mapping having the scientific discipline’s roots in information visualisation, which besides usability embraces philosophical and sociological contexts.
Footnotes
Acknowledgements
The author wants to thank the anonymous reviewers for their careful reading of the manuscript and many insightful comments and suggestions.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship and/or publication of this article.
