Abstract
We present semantics-based mechanisms that aim to promote reflection on cultural heritage by means of dates (historical events or annual commemorations), owing to their connections to a collection of items and to the visitors’ interests. We argue that links to specific dates can trigger curiosity, increase retention and guide visitors around the venue following new appealing narratives in subsequent visits. The proposal has been evaluated in a pilot study on the collection of the Archaeological Museum of Tripoli (Greece), for which a team of humanities experts wrote a set of diverse narratives about the exhibits. A year-round calendar was crafted so that certain narratives would be more or less relevant on any given day. Expanding on this calendar, personalised recommendations can be made by sorting out those relevant narratives according to personal events and interests recorded in the profiles of the target users. Evaluation of the associations by experts and potential museum visitors shows that the proposed approach can discover meaningful connections, while many others that are more incidental can still contribute to the intended cognitive phenomena.
Keywords
1. Introduction
Information and communication technologies are progressively transforming the way that the public can appraise cultural heritage, both virtually and in situ. The inception of the World Wide Web first made vast repositories of information available online [1]; then, the social revolution of the Web 2.0 allowed people to share cultural experiences [2]. The development of the Semantic Web (a.k.a. ‘Web 3.0’) enabled computers to process growing amounts of metadata, for example, to automatically assemble personalised groups of items that can be presented to the target users as multimedia galleries in an app, or tailor-made itineraries to follow in a museum [3]. Nowadays, many research efforts on the use of Web 2.0/3.0 technologies in the area of cultural heritage are focused on the storytelling aspects, aiming to ensure that the items and the pieces of information presented to the users make sense together and serve to deliver messages that they may reflect on and retain easily [4–7].
The use of narratives is already widespread in museums in various forms, aiming ‘to present inclusive and nuanced history, to make big ideas less overwhelming and abstract, and to create frames of experience that encourage deep and satisfying engagement for visitors and online users’ [8]. Notable developments have allowed location-specific narratives to be presented to visitors, through systems that exploit spatial information, maps and synchronised content [9,10]. The narratives are commonly developed by venue curators, who can use specific templates and formats [11] or semantics-based aids for emplotment, that is, for the selection of significant events in a story and the identification of pertinent relations among them [12]. Several experiences have been reported on uses of social media to collect narratives from the museum visitors, who could contribute their personal stories and experiences [13–15]. Studies on group-based narratives have shown that not only do they increase cultural understanding but they also significantly enhance visitor cooperation [16,17]. Overall, it has been found that, whichever the medium of delivery (text, audio guide, video, even augmented reality), narratives can be a very effective way to increase the level of immersion and, hence, the quality of the cultural experiences [18].
In this article, we address the problem that arises when there are several (or many) narratives available for the museum visitors to choose one from. Expecting visitors to read textual descriptions could be an option when there are a few narratives. With sizable sets, in contrast, it is necessary (at least) to list the narratives in decreasing order of relevance, taking into account the context and interests of each visitor. There is abundant evidence that humans retain more easily information that is presented to them in relation to anecdotes, unexpected paraphernalia or situational context [19–21]. Accordingly, our approach seeks to identify connections between the narratives and the current date (either due to historical events or annual commemorations) in order to find subsets of narratives that are more relevant to each given day. Thereupon, we use word embeddings [22–24] to compute personalised recommendations, by sorting out those subsets in relation to any personal events and interests recorded in the profiles of the museum visitors.
Our proposal takes place in the context of the European project CrossCult, 1 which aims to foster reflection on cultural heritage and history through interconnections among cultural digital resources, physical venues and citizen viewpoints. The technological platform created within the project is briefly presented in Section 2, followed by the details of the components that support and implement the management of calendar-based associations in Sections 3–5. Section 6 presents the pilot experiment conducted in the Archaeological Museum of Tripoli (AMT; Greece), which aims to harness major crosscutting topics (e.g. freedom, health or the role of women in society) to deliver multiple narratives through the museum items. A discussion driven by the findings of the experiment is given in Section 7, and the article concludes with Section 8.
2. The CrossCult platform
The CrossCult platform is a complex ensemble of software aimed to provide services to different types of stakeholders, including museum curators and experts, data scientists, cultural app developers and system administrators (through different Web-based frontends) as well as current and future museum visitors (through Android or iOS apps). The operation of the Web-based and mobile frontends is supported by a backend that provides infrastructure and instrumentation for hosting the service components.
At the core of the platform, the CrossCult knowledge base (CCKB) is a repository for storage, management and retrieval of semantic information. It implements a semantic layer of common cultural heritage concepts and their relationships, building on standard Semantic Web technologies to facilitate interoperability and linking with linked data resources. Linked to the CCKB, a number of software modules implemented as microservices [25] provide high-level, application-oriented services covering six major functional areas, namely ‘Association Discovery’, ‘User Profiling’, ‘Recommendation’, ‘Context Awareness’, ‘Social Networking’ and ‘User Experience’. Each functional area is covered by one or more technological modules, which offer distinct services (e.g. chatting and microblogging) or address different facets of a single issue in a complementary fashion (e.g. carousel-based profiling vs interaction-based profiling, item recommendation vs path recommendation).
In the following sections, we present the details of the platform components that implement our proposal for the management of calendar-based associations, seeking to identify the most relevant narratives for museum visitors. These are the following:
The CCKB as a whole, with particular attention to the semantics used for the reflective narratives and the special dates (Section 3).
An ‘Association Discovery’ microservice that can provide connections between the heritage items of any venue registered in the CrossCult platform and the special dates modelled in the CCKB (Section 4).
A ‘Recommendation’ microservice that sorts out a set of narratives according to personal events and interests recorded in the profile of any given visitor (Section 5).
3. Knowledge base structure
The CCKB [26] is a comprehensive, standard-based structure of semantic definitions and formalisms, developed for facilitating interoperable connections among cultural heritage data. Its architecture is shown in Figure 1, with different sections carrying different semantics and every box built/expressed using the definitions of the ones beneath it:
The bottom section includes four ontological schemas that constitute the foundation of the architecture, with Conceptual Reference Model of the International Council of Museums (ICOM)’s International Committee for Documentation (CIDOC-CRM) [27] being the most prominent. This is an international standard (ISO 21127:2006) for modelling cultural heritage information, providing an extensible semantic framework that any cultural heritage information can be mapped to. The framework is complemented with the semantics of the Simple Knowledge Organization System (SKOS 2 ), which is a W3C recommendation designed for representing thesauri, classification schemes, taxonomies, subject-heading systems, or any other type of structured controlled vocabulary. The Dublin Core 3 schema is adopted as a standard vocabulary for describing web resources, and the FOAF (Friend of a Friend 4 ) ontology is used for describing user related entities and their interests.
The middle layer of the CCKB accommodates the semantics of the upper-level ontology, which captures common concepts and relationships across a diverse range of cultural heritage data. It is driven by a core-subset of CIDOC-CRM semantics, complemented with a set of project-specific definitions to model reflection aspects. The upper-level ontology enables augmentation, semantic linking, semantic-based reasoning and retrieval across disparate data resources, and its instances are enriched with links to DBpedia 5 concepts, which provides additional interoperable properties and connections to a large body of general knowledge.
The side section accommodates the CrossCult Classification Scheme (CCCS), a faceted vocabulary structure that aggregates terminology from standard thesauri resources, such as the Arts and Architecture Thesaurus of Getty (AAT), the EU’s multilingual thesaurus (EuroVoc), the UNESCO Thesaurus and the Library of Congress Subject Authorities (LC) vocabulary.
Finally, the top section of the architecture contains the venue and the user ontologies. The venue ontology is a fully CIDOC-CRM compliant structure, which aims to model the spatial arrangements of the different venues that participate in the project. The user ontology, in turn, is aimed at supporting user modelling requirements with respect to interests, visiting preferences, personality and cognitive traits, background and other ethnographic information. The ontology combines elements from the FOAF and CIDOC-CRM models, while introducing new properties to describe particular user characteristics, such as fatigue, prior knowledge and behaviour.

The CCKB stack of ontological layers.
3.1. The semantics of reflective topics
Further to the elements of CIDOC-CRM, the upper-level ontology contains a project-specific class (Reflective Topic) and a set of adjunct properties that allow creating a network of points of view, aiding reflection and prospective interpretation over a topic. Examples of reflective topics in CrossCult are ‘Daily life’, ‘Migration and industrial revolution in Europe’, ‘Mortality and immortality’ and ‘Religion and pilgrimage’.
The definition of a reflective topic requires accommodating a range of semantics relevant to a theme. As illustrated in Figure 2, the Reflective Topic class connects to other upper-level ontology classes via a set of well-defined semantics, some of which constitute project-specific extensions of standard CIDOC-CRM properties. In detail, the class can be understood as extension of the E89.Propositional Object 6 class, extended by the project-specific property ‘reflects’. This property sets a reflective topic instance as the primary subject of reflection of a physical or conceptual source. For example, the Eiffel tower can be used to drive a reflection on engineering and industrial revolution: hence, the physical object ‘Eiffel tower’ reflects the reflective topic ‘Engineering marvels of Europe’.

Semantics of the Reflective Topic class, with the example of rt_0053 Education Apollo and Muses.
A reflective topic is characterised by a Title (E35.Title) and is further contextualised with links to CCCS terms (skosConcepts). Besides, a broad reflective topic can be composed of more specific (narrower) ones. The property P148.has component allows for recursive composition, which can be experienced sequentially via the semantics of the has first, has next and has last properties. Finally, multimedia elements, modelled as E73.Information Object, further describe a reflective topic by accommodating text and audiovisual materials. Such elements fall into three categories which are distinguished via property:
P67_2.has media is the most generic and is assigned to elements that simply complement the topic.
P67_3.has intro is assigned to media that introduce the topic or act as a trigger for engaging with it.
P67_4.has narrative is assigned to the elements that drive reflection through a narrative.
Reflective narratives are short stories authored by Humanities experts, aimed at contextualising a reflective topic with inspiring viewpoints and historical/social facts, complemented with links to digital resources. As such, the narratives revolve around a particular museum item or a broader collection of exhibits, aiding reflection and reinterpretation by storytelling. Stories, images, hypertext resources and audiovisual elements can be interwoven into rich compositions within the semantic environment of the CCKB, via (a) author-based assignment of CCCS terms to physical items and reflective topics and (b) automated processes of named entity recognition and resolution. As an example of the latter, the reflective narrative cckb:E73/rn_0053 in Figure 2 was enriched with links to a number of DBpedia resources (e.g. db:Kithara, db:Sappho, db:Muse and db:Apollo) using DBpedia Spotlight [28], which automatically recognises named entities in natural language text.
3.2. The semantics of special dates
The core motivation for this article is to model special dates (historical events and annual commemorations) in the CCKB in order to trigger calendar-based associations with cultural heritage items, which act as entry points for delivering potentially interesting narratives to users. The ontology classes and properties used to model special dates enable connections with other classes of the ontology such as those used to model reflective topics and physical items.
The special date entries carry descriptions about developments that occur annually or occurred in the past on particular dates, which somehow affect or have affected some states or behaviours. In this respect, they can be formally understood as events under the definition of the CIDOC-CRM class E5.Event, which ‘comprises changes of states in cultural, social or physical systems, regardless of scale, brought about by a series or group of coherent physical, cultural, technological or legal phenomena’ [27]. Figure 3 illustrates the semantics, classes and properties that are employed to formally describe special dates. At the core of the definition is the E5.Event class, which holds together the various elements of a special date. The E52.Time-Span class defines the actual date of the event, which is expressed as an instance of time in the form of the xsd:dateTime 7 datatype. The E50.Date class complements the temporal definition of a special date instance by providing a date in the form of an appellation. The actual description of a special date is accommodated by an E73.Information Object which also carries (P67.refers to) links to CCCS and DBpedia concepts, just like reflective narratives (see Section 3.1).

Semantics of the special date of Franz Grillparzer’s ‘Sappho’ premiered on 21 April 1818.
Figure 3 illustrates the example of the special day cckb:E5/04d1e42250 that occurred on 21 April 1818, which refers to Austrian writer Franz Grillparzer’s ‘Sappho’ premiere. The description is enriched with links to the DBpedia concepts ‘db:Vienna’, ‘db:Sappho’ and ‘db:Franz_Grillparzer’, provided by DBpedia Spotlight from a textual description of the event given in www.onthisday.com.
4. Association Discovery
As explained in Section 3, in the CCKB, each cultural heritage item relates to one or more reflective topics, and through them to reflective narratives, which are enriched with links to CCCS and DBpedia concepts. Special date entries are similarly linked to CCCS and DBpedia concepts, enabling techniques for cross-searching and association discovery via a common layer of semantics. Using SPARQL queries, we can identify associations between museum items and special dates by generating subject-based matches via a common layer of concepts applicable to both.
The SPARQL query below is one sample from the catalogue of association discovery queries used in the CrossCult microservices. It exploits DBpedia enrichments
Figure 4 illustrates an example of a discovered association between a museum item and a special date. The museum item MT0034, which belongs to the AMT (Greece), is a marble plaque depicting an assembly of the nine Muses with Apollo Pythios in a rocky landscape. The item is used to drive reflection on the topic of Education, hence it is connected to (reflects) the reflective topic rt_0053, which is furnished by the narrative rn_0053. The narrative tells the story of Apollo and Muses and how music played an important role in the education of Ancient Greeks, particularly of women. It then moves into highlighting the role of the female poet Sappho in the music education of women in ancient Greece. The narrative is linked to several DBpedia concepts, one being db:Sappho, which is also related to the special date 04d1e42250, the day Franz Grillparzer’s play ‘Sappho’ premiered in Vienna, on 21 April 1818. Through this entry, an ancient artefact, depicting Apollo and the Muses, can be related to a 19th-century tragedy inspired by the life of an ancient Greek female poet. Both ends support and stimulate a discussion about education of women, originated by a reflective topic in the CCKB.

Example of association discovery between a museum item and a special date via a DBpedia concept.
5. Personalising associations
Through the process described in Section 4, the number of candidate associations can be overwhelming once the CCKB contains annotations for more than a few hundred special dates. Depending on the terms provided by DBpedia Spotlight, associations can be found between one date (often based on the same historical event) and most of the items. In order to both limit the volume of associations presented to a visitor, and to provide only associations appropriate to each individual, a number of steps are taken as described below and visualised in Figure 5.

Flowchart of the association discovery process, linking museum items (yellow) and special dates (orange) indirectly via a DBpedia concept, and the selection of appropriate associations based on personal events and interests (green).
5.1. Important personal events
User profiles in CrossCult can contain data gathered about each visitor in different ways (explicit or implicit [29]) and accumulate the annotations resulting from the use of different apps or Web-based questionnaires. One of the profiling features of the apps allows users to provide the dates (day, month and year) for personal events that have importance to them, which are modelled by the user ontology of the CCKB (see Section 3). Examples of such events are provided in Section 6.2, including dates of birth, marriage or graduation. The users can specify the meaning of each event (which can be exploited in the personalisation processes), but they are not required to do so. Each personal event is used to filter associations based on different combinations of year, month and day information. For the sake of clarity, we assume an example of personal event on 14 February 1976. Initially, associations are sought related to the exact day, month and year (e.g. a historical event on 14 February 1976, such as a US nuclear test at the Nevada Test Site, or the establishment of Fondation Vasarely museum in Aix-en-Provence), then based on the exact day and month (e.g. all associations on 14 February, such as the annual event of St Valentine’s day), then based on the year (e.g. all historical events of 1976), then based on the month (e.g. all annual and historical events on February) and finally based on the day (e.g. all events on any month’s 14th day). The order of the filtering plays a role, as more exact matches are preferred for showing to the user than broader matches. The broadest match would be any date with the same day, but could still be presented in an appealing way to the user based on the proximity of the upcoming personal event: for example, on 14 December, the application could show an association with the introductory message: ‘Your personal event is exactly two months away! Today, …’.
5.2. Personal interests
The user profiles can also capture personal interests, which are currently chosen from lists of reflective topics or keywords. These may be provided explicitly by the users, or learned (without user intervention) by the profilers of the CrossCult platform by observing their actions in any apps (e.g. by keeping track of the topics they choose to read about when offered several choices). As explained in Section 3, reflective topics and keywords are curated by Humanities experts, and they can be used to further filter associations based on their similarity to the CCCS or DBpedia concepts that bring about the associations.
While there is a variety of ways to find which concepts are relevant to any personal interests provided by the user, we have resorted to Word2Vec [30] as the most general and scalable approach. Word2Vec uses artificial neural networks to reconstruct linguistic contexts of words. The model used in this work relies on a pre-trained Google News corpus (
Using Word2Vec, we compute similarity scores for each pair of association-related concepts and a user’s personal interests. For each association, we consider the closest personal interest to be the one for which the corresponding pair of association-related concept and the user’s interest has the maximum similarity score. We then take two steps to filter associations based on this information:
Remove trivial similarities. If the maximum similarity score of the association is below a specific threshold, then the similarity is considered trivial and the association is removed from the list. This threshold is needed in part because Word2Vec returns non-zero similarity scores between most words in its corpus. The choice of threshold affects the volume of associations removed, and for this article, it is chosen arbitrarily based on the use cases of Section 6. Indicatively, the similarity score of ‘family’ with ‘families’ is 0.55; with ‘child rearing’ it is 0.26; with ‘nuns’ it is 0.16. Based on these value ranges and following some experimentation regarding the volume of associations removed with different thresholds, this article omits all associations where the maximum similarity score between concept and personal interest vector is
Find best matching event per date. After all trivial associations are removed, one association is chosen for each day and month combination, so that the user is not overwhelmed daily by numerous associations. Instead, on any given day they may receive either one association (the one closest to their personal interests) or none (because no association exists on this date at all, or because this date is not relevant to their personal events, or because the association is only trivially connected to their personal interests). In this vein, associations found on the same day and month are sorted based on the maximum similarity score between their keyword and the user’s personal interest vector. For each day and month, the association with the highest similarity score is displayed to the user.
6. Case study: the AMT
We have tested our methods in the context of one of the CrossCult pilot experiments, titled ‘One venue, non-typical transversal connections’, that takes place in the AMT. The starting point for this museum is representative of the current situation of thousands of small and medium-sized cultural venues around Europe, which suffer from very little traffic and whose treasures are unknown to the vast majority of citizens. The museum owns a small collection of heritage items, arranged into different rooms according to chronology and accompanied by shallow, unconnected information panels that merely indicate the type of a statue or the transcription of the text carved on a tombstone, but nothing (or very little) regarding its meaning and context. In such conditions, the museum often failed to deliver even one of the many stories that it could tell.
Our hypothesis was that, if those stories were developed and annotated properly – including the management of calendar-based associations we advocate in this article – then it would be possible to deliver interesting content linked to events that are meaningful to each visitor. In this line, a team of Humanities experts from the CrossCult consortium developed a set of 75 reflective narratives about life in antiquity involving the heritage items of the museum, using a controlled vocabulary about appearance, mortality, religion, rituals, goddesses, humans, amazons, nudity, social status, education, daily life, weaving, dowries, food, names, wild animals and healing practices. These narratives were associated (via human curation, assisted by CrossCult tools for experts [31,32]) to 17 archaeological items displayed physically at the AMT. These narratives provide a wealth of opportunities to automatically identify the most interesting stories to offer to any visitor, enabling synthesised views on reflective topics that could hardly be conveyed before.
For this experiment, the CCKB was populated with 60,525 special dates. Most of them were created from the online resource www.onthisday.com, passing the short textual descriptions of each historical event through DBpedia Spotlight, with default settings. In addition, we compiled a list of annual commemorations observed by the United Nations and UNESCO, which convey global significance; we also processed the National Days and Flag Days listed in Wikipedia, which are useful for filtering associations based on the users’ nationality (in general, people are more interested in historical facts involving their own country than in others). These commemorations were annotated manually with AAT and EuroVoc concepts, and any textual descriptions (or even only the title of the commemoration) were passed through DBpedia Spotlight as well.
In the following sections, we first present the output of the discovery of special date associations for the museum. Then, we analyse the results provided by the personalisation mechanisms for three synthetic profiles. Finally, we present the results of an evaluation poll conducted with a team of experts and a set of potential museum visitors to appraise the wisdom, interest and value of the associations in relation to the intended phenomena of curiosity, reflection and retention.
6.1. General calendar-based associations for the AMT
Performing association discovery, as described in Section 4, between the semantic annotations of the reflective narratives created for the AMT and the two types of special dates results in an extensive set of matches. In total, 3856 associations are found, out of which the majority (3544) is with historical events and 312 are with annual commemorations. This imbalance is hardly surprising, as the volume of historical events is massive (with more than 160 events daily). Similarly, the common concepts found between the narratives and the events are different based on repository: historical events are linked to the AMT reflective narratives through 57 DBpedia entities (such as ‘db:Ancient_Greece’, ‘db:London’ and ‘db:Track_and_field’) while annual events are linked predominantly through AAT concepts (16 of them, including ‘fertility’, ‘health’ or ‘women’), which are usually broader. This does not mean that there is no overlap, however: gay pride, theater, slavery, Greece and Cyprus are found as links in both sources. 8 As expected, some links appear far more often than others: in associated historical events, the most popular links are Greece, London, Athens, Ancient Greece and slavery, which account for 79% of all associations with historical events. Associations with annual commemorations are mostly found with the following concepts: men (30%), health (20%), family (16%) and breast (11%). All the concepts present in associations are shown in Figure 6.

All keywords used to find the associations between special dates and the reflective narratives from the AMT (size and colour both denote prominence).
Considering other aspects of the associations discovered, Figure 7 shows how associations are distributed in terms of the concepts used as links, the distinct dates (i.e. day–month combinations) and the reflective narratives and museum items that are linked. Evidently, much of the information comes from historical events, although annual events have associations with most narratives and museum items. Notably, there are two narratives where an association is found only with annual events, but not with historical events: one is associated with four different dates on the topics of fertility, family (twice) and women; the other is associated with men, which links it with every day of November (due to the ‘Movember’ month, dedicated to men’s health). Due to this association, November dominates (at 53%) the associations with annual events; however, every other month except for April and July is also represented with at least one annual event in the associations found. By comparison, almost every day has at least one association due to a historical event. Considering both historical events and annual commemorations, 360 distinct dates within the year are represented. This finding gives significant leeway in terms of finding events on a specific date important to one person, but also comes with a daunting task of filtering the most relevant personal associations from a vast pool of almost 4000 candidates.

Distribution of associations found between reflective topics of the AMT and different types of events.
6.2. Filtered associations for three personae
In order to evaluate how the numerous associations can be tailored based on a specific user’s personal events and interests, we use three distinct personae (synthetic profiles) representing potential visitors to the AMT. Each subsection comes with its own analysis of the findings.
6.2.1. Persona: Mata
Mata is a girl from Tripoli, Greece, in her early 20s. She has lived in Tripoli her whole life, and since her parents got divorced, she has joined the goth subculture and the preference towards mysticism and the morbid. In a CrossCult app, Mata has provided her birthday (31 August 1999), the date her father left the house (13 February 2013) and the day she finished her final school year (16 June 2017). In terms of interests, Mata chooses ‘Veils’, ‘Cloaks’, ‘Talismans’, ‘Mortality’, ‘Funerary sculpture’ and ‘Cemeteries’ from the choices contained in the CCKB.
Mata provided three personal events, which are used to first filter only associations related to those dates. For her birthday, 4 associations are found with 31 August; 14 with the year she was born (1999); 446 with the month of August and 55 with the 31st of any month. Using all three dates in a similar way, 1280 associations on 119 distinct dates throughout the year are found without taking into account her interests. Unsurprisingly, most of these dates are in August (26%), February (24%) and June (24%). Given the large number of associations, it is important to filter them further taking into account Mata’s interests. Among them, ‘Cloaks’ and ‘Talismans’ are not found in the Word2Vec database and are thus ignored in our approach. Using a similarity threshold of 0.2 between Mata’s four remaining interests and the concepts linking dates and reflective narratives, 418 associations remain.
As a final step, associations are filtered by similarity and the most similar to Mata’s interests (based on the closest word among those interests) for each date is chosen. The result is 58 associations, all on distinct dates. The distribution of these associations is shown in Figure 8. Most of them are derived due to Mata’s interest in veils (67%) which, interestingly enough, appears most often associated with the concept of slavery (the Word2Vec similarity between ‘veils’ and ‘slavery’ is 0.31). This explains why slavery is often prominent among the associations (45%). It is also interesting to note that Mata’s interest for cemeteries is used to filter associations based on the concepts ‘village’, ‘archaeology’ and ‘health’. Mata’s interest in funerary sculpture, however, is only exploited to find one relevant date: 24 August (7 days before her birthday), the day that on 79 AD ‘Mt Vesuvius erupts, buries Roman Pompeii and Herculaneum, 15,000 die’ (text from www.onthisday.com). This day is associated with a plaque depicting Apollo and the nine Muses, and is linked due to the term ‘Pompeii’ (in the narrative, the muse Sappho is also shown in an image of a painting from Pompeii). Unsurprisingly, again, of the 58 dates relevant to Mata, most are on the 3 months of her special dates. Finally, it should be noted that Mata’s associated dates refer to 15 annual events, while very few of the historical events are in the 20th century.

Distribution of Mata’s final associations (up to one per date).
6.2.2. Personae: Üter and Irmgard
Üter and Irmgard are a naturist German couple in their 50s, stopping in Tripoli on their way through clubs and beaches from Cephalonia down to Kalamata and then to Corinth. While searching for an afternoon activity on his phone, Üter found a link to the AMT and filled in an online form under the title ‘Let us personalize your visit’. When asked for three relevant dates, he provided his birthday (27 March 1960), Irmgard’s birthday (18 July 1963) and the date they got married (24 December 1980), and as topics of interest, he chose ‘Nudity’, ‘Marriage’ and ‘Mythology’.
As in the previous example, the couple provided three dates which are used to first filter only associations related to those dates. Based on these dates, a total of 1319 associations on 132 distinct dates throughout the year are found before taking into account Üter’s and Irmgard’s interests. Most of these dates are in March (32%), the month Üter was born, and December (20%); interestingly, July and August are almost equally represented (despite the former being Irmgard’s birthday), at 13% and 12%, respectively. Since the couple has a narrow set of interests, it is expected that with a similarity threshold of 0.2 most of the trivial associations will be removed. Indeed, 354 associations are close enough to the couple’s interests.
Finally, associations are filtered by similarity and the most similar to the couple’s interests for each date is chosen. The result is 66 associations, on distinct dates, distributed as shown in Figure 9. It should be expected that an Archaeological Museum has more links to the couple’s interest in mythology (58%) than nudity (11%). Interestingly, the concepts used do not refer to ‘Ancient Greece’, ‘Greece’ or ‘archaeology’ (only two associations are linked to Greece), despite the fact that these topics were prominent in the pool of associations. ‘Marriage’ is often used to choose associations, but all those associations are based on ‘slavery’ (the two terms have a Word2Vec similarity of 0.35). However, ‘nudity’ is used to choose associations based on the concepts ‘bikini’, ‘men’, ‘nun’ and ‘water’. Most of the 66 dates are on a month of someone’s birthday (March and July) and to a lesser extent on the month of their marriage (December). Unlike Mata’s case, the couple’s associations are not often with annual events; moreover, there are more events referring to the 1900s and 2000s. The most recent one, for example, is on the 31 March (4 days after Irmgard’s birthday) of 2012, when ‘Fiji floods kill 2 people and force thousands to be evacuated’ (text from www.onthisday.com). This event is linked via the DBpedia concept ‘db:force’ with a tondo depicting Heracles and Auge from the third century AD, and its narrative on how he forced himself on her, the daughter of his host king Aleus of Tegea.

Distribution of Üter and Irmgard’s final associations (up to one per date).
6.2.3. Persona: JD
A user (identified as JD for John/Jane Doe) has used the anonymous login option for a CrossCult app. Concerned about privacy, JD chose not to give any of his or her personal events as information to the profiling mechanisms. However, he or she chose to put down some of his or her interests from the list provided: ‘heroes’, ‘athletes’ and ‘gay pride celebrations’.
Unlike the previous examples, the lack of a set of dates means that all possible days of the year will be considered for JD. With no primary filter for personalisation besides his or her interests, the total number of associations is 646. As expected, JD has far more associations than for Mata (418) and Üter and Irmgard (354), since the associations were not pre-filtered based on personal events. It is important to note that with more (or different) interests the number of associations could be much higher: for instance adding ‘patriarchy’ to the current interests would increase the total number of associations to 2713.
Choosing up to one association per date of the year (based on the highest similarity to interests) results in 178 associations. These associations are spread more uniformly across the months of the year than use cases which include personal events, as shown in Figure 10. However, there is still some imbalance between months, for example, the same association is made on every day of November due to the commemoration of men’s health (‘Movember’) which is linked to a marble tombstone with a representation of a woman and a young athlete, and a narrative on their clothing (or lack thereof, in the case of the athlete). Most associations are made due to the ‘athletes’ keyword, which is associated with ‘Athens’, ‘Track and field’ and many other terms. ‘Heroes’, however, is only associated with the term ‘force’, and 86% of those associations are with the narrative of Heracles and Auge discussed in Section 6.2.2. Finally, the interest in ‘gay pride celebrations’ results in two associations: (a) with 17 August, when three members of Russian punk band ‘Pussy Riot’ were put into jail for 2 years in 2012 and (b) with 25 June, when the rainbow flag was first used, in 1978. Both dates are found associated with a headless statuette of a young girl and its narrative on education in Ancient Greek society (specifically, that young girls stayed at home while boys over 7 years old left the house to receive education).

Distribution of JD’s final associations (up to one per date).
6.3. Evaluation by Humanities experts and potential museum visitors
In order to appraise the ability of associations with dates to foster reflection, retention, curiosity and other cognitive phenomena, we asked four experts in Humanities and 81 other users (of a broad age range and potentially interested in visiting the AMT) to tag as many associations as they could from among the sets computed for the three personae of Section 6.2. The associations were automatically formulated in a way that presents the date, personal context, museum item, reflective narrative and associated event (see Table 1 for an example). The evaluation was conducted in two rounds between March and July 2018, recruiting non-expert users from among students of diverse degrees in the University of Vigo in Spain and the Arab Academy in Egypt. Feedback from expert users was solicited via direct contacts from within the University of Vigo.
Sample association description computed for the case of Üter and Irmgard.
To begin with, the participants were asked to assign any of the following tags (possibly none, possibly several) to the associations and the linked narratives:
Informative: the association/narrative provides new knowledge.
Thought-provoking: the association/narrative makes me reflect on the association itself.
Memorable: the association/narrative is probably to be remembered.
Curious: the association triggers curiosity to make the narrative attractive.
Personal: the association/narrative is connected to the user’s interests or dates.
Funny: the association/narrative can be perceived in a humorous way.
Clearly, the criteria for success would be to get many associations tagged as ‘Thought-provoking’ (as it relates to reflection), ‘Memorable’ (retention), ‘Curious’ (curiosity). The number of ‘Informative’ tags provides a measure of interest in the associations, whereas ‘Personal’ aimed to preliminarily assess the value of sorting associations according to personal dates and interests. The number of associations tagged as ‘Funny’ was a secondary aspect.
In addition, the participants had to choose one of the following mutually exclusive tags for each association:
Notable: the association is close to museum items or the their narrative.
Indirect: the association has some sort of connection, but this connection has several degrees of separation.
Irrelevant: the association is either purely circumstantial, uninteresting, misleading (because of incorrect interpretation of the meaning of a term) or unclear.
Finally, the four Humanities experts were asked to indicate whether the associations that they had found to be ‘notable’ or ‘indirect’ could also be described as:
Valuable: the association is worth showing to the museum visitors.
Useful: the association can increase the visibility and/or the understanding of the museum items.
Potentially offensive: the association involves terms that could be offensive to some potential visitors, and should therefore be filtered.
Table 1 shows a sample of the association descriptions that were provided for review, along with the persona descriptions. The associations were mixed and distributed randomly, with experts receiving a total of 136 associations and potential visitors receiving 323 associations. The final ratio of tags to presented associations are shown in Figure 11.

Ratio of associations tagged by participants, per type of tag and type of user.
Based on the total tags provided by Humanities experts and potential visitors, several conclusions about the quality and usefulness of our approach can be gleaned. Considering first the mutually exclusive tags, we observe that of the 136 and 323 associations rated by experts and potential visitors, respectively, 26% are deemed irrelevant by experts and 17% by potential visitors. While this is a promising finding, participants also predominantly considered the associations made to be indirect (56% for experts and 62% for potential visitors). Even indirect associations are deemed meaningful, however, since many participants tagged such associations as ‘informative’, ‘memorable’, ‘thought-provoking’ and ‘curious’. Notably, potential visitors were more prone to use such tags than experts, as they are probably less knowledgeable of the topics discussed in the narratives of the museum. Since the tools are intended to attract potential visitors, this is a positive finding. Even though it is not overwhelming, the presence of the ‘memorable’, ‘thought-provoking’ and ‘curious’ tags reinforces the intended value of our approach in terms of raising curiosity to deliver more information about cultural heritage in a way that increases retention, reflection and, in the end, understanding. The analysis of co-occurrence of tags, correlations with user data and other measurements such as inter-rater agreement ratio, is left for future work, since the size of the current sample is not sufficient and the associations were randomly distributed among participants. The predominance of ‘indirect’ associations (compared with ‘notable’ ones) suggests that connecting cultural heritage to dates could also be exploited to promote serendipity, in the sense of learning about valuable or agreeable things not initially sought by the museum visitors.
It is worthwhile to investigate results regarding the tag ‘personal’ in particular, which had a total of 53 occurrences among 369 ‘notable’ or ‘indirect’ associations. We believe that this low number of ‘personal’ tags is partially caused by the limited amount of personal information in the personal profiles, with a few relevant dates and interests from a closed vocabulary. The design of the experiment could also have an influence as this tag was optional and, ultimately, it is questionable that one might be able to properly evaluate whether something is ‘personal’ when it was selected for another person (in this case, a synthetic persona). Nevertheless, based on users’ feedback (through tagging and later discussions) we can conclude that many historical events would typically be outside the interests of most visitors, for example, being only relevant to certain nationalities. For example, events such as ‘Great Storm of 1987: hurricane force winds hit the South of England killing 23 people’ and ‘Lady Godiva rides naked on horseback through Coventry, to force her husband to lower taxes’ could be relevant to someone from Great Britain, whereas a Spaniard would probably be more interested in a less tragic event happening in his or her country, or in characters from local history. All in all, this suggests that additional personal information (such as country of origin) may be an appropriate additional filter to find matchings with the set of dates from the CCKB.
Experts additionally tagged the 100 associations that they considered not to be ‘irrelevant’ as both ‘useful’ (in 60% of cases) and ‘valuable’ (41%). Their verbal feedback about the associations they found ‘valuable’ and/or ‘useful’ also revealed that, even in the cases of ‘notable’ associations, it would be necessary to add one or two sentences to bring all aspects together, so that the associations could be properly understood in the end. In other words, experts could be put in the loop to reinforce the link between dates and narratives. For example, the historical event from 1959, ‘1st known radar contact is made with Venus’ can be related to mythology, but only after explaining why planets were named after gods. This can be considered in terms of explaining the associations to the user, which was not the main focus of this work.
In relation to the ‘potentially offensive’ tag, experts indicated that some discovered associations touch on issues that might be regarded as sensitive and controversial. Visitors, depending on their cultural background and beliefs, might feel less comfortable with associations exploring certain historical and social matters. Overall, 17 associations were identified as such, typically involving headlines about military confrontations (e.g. ‘Fighting breaks out between Turks and Greeks over dispute islands in Cyprus and 16 are killed’) or sexual orientation (e.g. ‘Gay pride events are banned for a century in Moscow’). While controversy has been known to act as a powerful trigger of reflection [33,34], the experts pointed out that museums should probably avoid unnecessary controversy, especially when the association of certain events to the cultural heritage of the place is indirect. Notwithstanding, they argued that the attempt to recommend (push) one or several narratives to visitors is most probably to have a beneficial effect.
Finally, the experts also noticed that the association discovery algorithms can provide them with clues for developing new reflective narratives to enrich a venue’s contents. This is a promising direction for future work which deviates from the original goal of showing content directly to users, and incorporates association discovery as part of a curator’s workflow of providing appropriate links with popular, commemorative, or indirectly related events and concepts. The additional step of a curator-driven enrichment can alleviate automatically discovered (and sometimes irrelevant or potentially offensive) associations, by removing or rephrasing such associations, and then applying the personalisation algorithms to select among the curated associations.
7. Discussion
This article presents how calendar-based associations between cultural heritage items (and collections thereof) and historical events or annual commemorations can be discovered automatically, as a means to bring specific attention to some of the many narratives that may be linked to a venue’s collection. We also propose a way in which those associations/narratives can be prioritised according to information captured in the profiles of venue visitors, plus an important feature of context: the day it is today. The proposed framework aims to make it easier for visitors to grasp the stories that the venue can tell, so that they can promptly decide which itinerary to follow. To the best of our knowledge, there have been no aids for such a decision in state-of-the-art projects about storytelling applied to cultural heritage experiences [8,35].
Our approach relies on mainstream standards for the semantic modelling of cultural heritage information, which we enhanced with a selection of additional resources plus new classes and properties in order to capture the relevant special dates, and to place reflective topics and narratives as the key mediating element in the association discovery and personalisation processes. It is worth noting that the construction of a thorough compendium of historical events is an open research problem [36,37], with notable contributions nonetheless in previous works [38–43]. The management of periodic commemorations and their meaning, in contrast, has not been systematised before. In turn, the reasoning features of our system combine an ontological approach and a word vector model which can handle any type of user interests (including, for example, free-form text). Recommender systems in the cultural heritage area (such as those of [44–46]) have not previously investigated such a word vector model.
Special dates can reveal an unprecedented range of possibilities for personalised and context-aware cultural heritage experiences. Namely, they allow exhibits to be rearranged in countless ways and can promote all the narratives written for a given venue throughout the year, thanks to links to universal topics and intra-venue to cross-border associations. Our experiment with the AMT shows that, in general, there may be a plethora of possibilities to promote reflection on a collection of cultural heritage items. Both the Humanities experts and the potential museum visitors confirmed that the associations and the linked narratives can deliver new information, helping to retain cultural knowledge and prompting further reflection. To a lesser extent, we have noticed that the approach can trigger curiosity, even though some associations are weak, many are indirect, and there is much research to be done on disambiguation, misleading words and relevance thresholds. Based on the evaluation results and the feedback from experts, we can conclude with some confidence that the recommended narratives are useful for promoting reflection and retention over significant cultural and historical topics.
This work will continue during the next months to assess the potential of the calendar-based associations, to fully investigate whether the connections to contextual and personal information contribute to long-lasting learning about cultural heritage. For this purpose, we will conduct new experiments with the AMT and other venues participating in CrossCult, following a more thorough experiment design and performing a deeper analysis regarding, for example, co-occurrence of tags, inter-rater agreement and other features highlighted in Section 7. A more longitudinal study could also be interesting, for example, by interviewing the participants in past experiments after a few months, with very specific questions aiming to assess how much they retain from the associations. In any case, from the point of view of scientific research and technical development, we have introduced the initial steps of semantic association discovery in the cultural heritage domain whereas many more steps remain to be explored. Among the lines of work that we want to explore in the future, we can highlight the following three: (a) adding to the CCKB a new knowledge layer with information regarding birth and death dates of notable people, another one of sports-related events, and the largely anecdotal commemorations listed in websites such as www.daysoftheyear.com; (b) replacing/supplementing the pre-trained Google News corpus of Word2Vec with a new corpus trained on a collection of documents from the cultural heritage domain, most probably Europeana (www.europeana.eu); and (c) improving the way in which the associations are presented to the end-users, following the advice of experts described in Section 7.
8. Conclusion
We have presented a method for associating historical events and annual commemorations with items and narratives specific to a museum collection for the purposes of context personalisation driven by user interests and personal dates. Taking advantage of a broad range of techniques for semantic modelling, named entity recognition and linking, online data repositories and word vector models, we managed to find associations, most of which were deemed accurate (directly or indirectly) by potential visitors. Evaluation results from a fairly free-form experiment involving domain experts and users suggest that calendar-based connections can reveal useful and valuable associations, which can be used to tailor user experiences and engagement with cultural heritage content. Further exploration of the proposed personalisation method is required to improve the accuracy and relevance of associations and capitalise on the most engaging aspects of the proposed framework. We see this work as a strategic point of action in the CrossCult project, in alignment with its ultimate goal of interconnecting cultural digital resources, physical venues and citizen viewpoints in order to foster reflection on cultural heritage and history.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work has been funded by CrossCult: ‘Empowering reuse of digital cultural heritage in context-aware crosscuts of European history’, funded by the European Union’s Horizon 2020 research and innovation program under grant agreement no. 693150. The authors from the University of Vigo were also supported by the European Regional Development Fund (ERDF) and the Galician Regional Government under agreement for funding the AtlantTIC Research Center for Information and Communication Technologies, as well as the Ministerio de Educación y Ciencia (Gobierno de España) research project TIN2017-87604-R.
