Abstract
This research investigated the strategic development of a large-scale transdisciplinary area, named intelligent infrastructure for human-centred communities, at Virginia Tech. Within such development, this study explored the future vision and anticipated scenarios of data infrastructure and digital libraries for smart community development. It draws upon the mixed-methods approach combining ethnographic participant observation, document analysis and semi-structured interviews. Grounded in socio-technical framework and rooted in empirical methods, this research produces results that augment design thinking and visioning practice for digital data libraries beyond traditional boundaries. The findings reveal the emerging scenarios around complex adaptive systems, intelligent data infrastructure and future digital libraries all in the context of building infrastructure for human-centred communities. Situated in this advancing reality, the results further discuss the next-generation data and information user experience, smart infrastructure data environment and future library capabilities. The article concludes that a smart library system, whether in its conceptual form of a ‘digital octopus’ or a ‘smart village data hub’ or an ‘intelligent virtual assistant’, will provide intelligence in data gathering, processing, summarising, communication and recommendations. By delivering unified and personalised data solutions, it will offer an end-to-end seamless experience for users throughout their journey of knowledge pursuit.
Keywords
1. Introduction
Today, there is the increasing use of ‘drones, smart city sensor networks, the Internet of Things (IoT), or proprietary device networks’ [1] for gathering valuable big data. This massive number of ‘smart’ objects, devices and sensors will ‘cooperate with each other, have their own metadata and continuously produce new data’ [2] in different forms. Data management will be a major and even bigger challenge. This will require the following: new techniques to harvest, filter and store relevant data; efficient processing, extraction and representation approaches for data objects to increase their usability and applicability; new protocols for interactively communicating decisions among objects; and smooth integration of heterogeneous data streams.
Powered by computing, sensing and communication technologies and featured by the interconnection of IoT with cloud computing, big data and cyber-physical systems, a smart infrastructure is ‘capable of collecting and analysing vast quantities of data to automate processes, improve service quality, provide signal feedback to users, and make better decisions’ [3]. It thus presents great promise for the building of thriving and resilient communities.
As the smart infrastructure agenda proliferates, higher education institutions are undergoing important transformations to converge intellectual domains and seek transdisciplinary solutions in this highly competitive area of global future. Academic libraries are at the nexus of big data, smart technologies and data-driven governance and have an important role to play in accelerating and coordinating the development of smart infrastructure. As the data landscape evolves rapidly and transforms disruptively, it is critical for libraries to begin re-imagining their future and re-visioning the architecture for the age of smartness. This provides opportunities to reframe, enrich and contextualise data infrastructure and digital library architecture in a manner that mirrors smart design and construction.
No longer a speculative scenario or science fiction, smart infrastructure and autonomous systems are an advancing reality today and so are the supporting data infrastructure and smart digital libraries. On many fronts, Virginia Tech (VT) is capitalising on its existing strengths and already available technologies to advance smart devices and digital networking, from driverless transportation to cartridge housing to new forms of energy [4]. Building smart communities involves a much broader vision of data infrastructure and novel concepts of digital libraries that need to be figured into the system design and community engineering process.
With such goals in mind, this work raises and addresses the following timely and critical research questions:
What are the implications of intelligent infrastructure, ubiquitous mobility and human-centred communities for digital library systems?
What are the associated data management, documentation and preservation challenges in smart infrastructure research and development?
What are the future concepts and anticipated scenarios of smart digital libraries?
Altogether, this study presents a schematic view of how the intelligent infrastructure evolution shapes and organises new library infrastructure and how a future-facing smart library system can mobilise the intelligent infrastructure advancement.
2. The evolution of intelligent infrastructure landscape
A confluence of technological advances, such as ubiquitous mobility, connectivity and surveillance, networked sensors and artificial intelligence (AI), marks the advent of a new era. More objects and devices are connected to digital networks than ever before, with the ability to sense their environments and both generate and communicate information about what is happening [5]. ‘The separation between cyber, physical, and social systems is blurring’ [6]. As a result, data volume is due to growth at an unprecedented pace, much of it from embedded devices. Smart systems are expected to grow, fed by millions of data points from multitudes of human and physical sources. Networks are becoming ubiquitous, offering information on physical things, human activities and cyber systems.
Collectively, these developments lead to the emergence of a new field that is transforming the academic enterprise of higher education. It is the field of smart infrastructure where networking, human and physical realms meet [7]. The emergence of this field is powered by interconnected sensing networks that collect and transmit real-time digital information from almost any object – for instance, roads, food crates, utility lines and water pipes [8]. It is enabled by ubiquitous digital information and communication technologies, such as clever software for analytics and visualisation, computer firepower and cloud computing, that help interpret the huge flow of information and transform raw data to useful knowledge, so as to monitor and optimise complex systems [8, 9].
On the whole, a smart infrastructure is an integrated cyber-physical and human system equipped with intelligent and self-configuring capabilities [3] to achieve real-time monitoring, efficient decision-making and enhanced service delivery [10–12]. It is a complex adaptive system ‘with dynamic and autonomous adaptation capabilities to deal with changes in their working environments and within themselves’ [13].
Essential to its performance are real-time data acquisition and processing, timely and comprehensive presentation of data [12] and reliable transformation and accurate communication of actionable information to decision makers or to system operations, in order to improve performance or adapt to changing conditions [9, 14]. It is an infrastructure that functions on the principles of data acquisition, analysis, feedback loop and adaptability [14]. In this sense, data hold the key to the innovation of infrastructure and is at the heart of all smart solutions. Modern libraries are central data stewards and life cycle managers. In this regard, data infrastructure and digital library systems are the fundamental provisions to support the creation, acquisition, curation, distribution and use of data, whether in the context of explanation, prediction or modelling [15].
With the world getting increasingly instrumented and interconnected along with the opportunities afforded by the abundance of data, it is expected that the smart innovation of infrastructure will drive efficiencies in water distribution networks, electricity grids, communication webs, transportation systems, food distribution and natural resources allocation, as well as other complex systems [8]. The accelerated progress in the smart industry will require integrative research and systematic design by highly coordinated teams [16] and will depend on the adoption of transdisciplinary perspectives and approaches to deliver a unified vision of a digitally enabled future [15].
VT is riding this wave of evolution by leveraging and connecting its leading programmes in smart design and construction, energy, robotics and autonomous vehicle systems, ubiquitous mobility and materials through an integrated area of development, named intelligent infrastructure for human-centred communities (IIHCC) [17]. Through converging areas of strength and in an integrative manner, IIHCC seeks ‘transdisciplinary solutions that exist along the interdependencies between humans, community, and infrastructure’ with the ultimate goal of improving quality of life around the globe [17].
Within this large-scale community uptake, it is critical to understand this complex next wave of technology development and in what ways the profound changes in the ‘smart’ realm really do matter for libraries’ future. It represents opportunities for the breaking of new frontiers in library development by becoming an integral component of this smart infrastructure and being able to interact with autonomous systems. In this integrated system, libraries will facilitate human–machine and human-building communication, support people’s interaction with technologies and optimise the individual and societal implications of such interaction for improved human condition and community well-being [18]. By this vision, digital libraries should be complex adaptive systems that respond to the needs of communities with escalating requirements and innovative ideas [19].
Undoubtedly, the effect of big data in intelligent infrastructure research and development is immense and their collection, processing, representation, aggregation, analysis and interpretation will be the main hurdles [16]. It thus raises urgent questions about data infrastructure, particularly at a time when it is imperative to raise the data foundation of the institution to be able to cultivate it as an asset. With such understanding, we should craft and create future-facing ‘people-focused’ digital data libraries in alignment with the IIHCC community’s perspectives, goals and aspirations [20].
To this end, this current research presents the future vision and projective scenarios of data infrastructure and digital libraries for smart community development. There is no previous work that has particularly studied and projected the future of digital libraries within the intelligent infrastructure framework. This research thus contributes new knowledge to this area of study. Framed in the IIHCC initiative by faculty stakeholders and thought leaders, all the scenarios described below reflect a transformative new paradigm and disruptive new direction for smart library development. These represent the needs to plan innovative solutions, define new functions and re-position us to conceptualise and operationalise digital libraries.
3. Methodologies
3.1. Context
The critical challenges society faces, such as water scarcity, access to education and rising cost of healthcare, are complex wicked problems with many dimensions [21]. Developing solutions to today’s multi-dimensional challenges require the synthesis of humanistic, scientific and technological perspectives [22]. Poised to become ‘an international destination for talent, partnerships, and transformative knowledge’, VT launched the ground-breaking, boundary-transcending Destination Areas (DAs) initiative in the spring of 2016. By way of iterative design, these areas combine VT’s existing strengths with novel transdisciplinary teams, tools and processes to tackle the world’s most pressing problems [23].
IIHCC is one of the five DAs identified to date. Having commitment from the College of Architecture and Urban Studies, College of Engineering and College of Science, together with multiple leading institutes and centres, a stakeholder committee was organised in September 2016 to provide collaborative leadership for achieving the IIHCC deliverables. Driven by deep faculty engagement and collaborative leadership model, the committee includes deans, institute directors and faculty members with demonstrated strengths in the key components of this DA. It is charged to orchestrate and guide the faculty cluster hiring, joint facility planning, curriculum development, research strategy and external partnership engagement during this important phase of transformation and innovation [22]. Libraries can in this process become a community mobiliser, data agent and information driver.
3.2. Methods
The current mixed-methods approach includes ethnographic participant observation, semi-structured interviews and document analysis. This study began in October 2016. The first several months of the DA initiation and IIHCC ethnography provided a framework to design this study. The document analysis covers relevant progress reports and stage updates as well as press releases and media coverage from the spring of 2016 when the IIHCC initiative was launched and throughout 2017 when this research was performed. During this period, the Principal Investigator (PI) participated in various DAs-related town hall meetings, information sessions, open forums, workshops and symposia while observing the interaction, taking ethnographic field notes and joining discussions. Through interactions at the various meetings and working sessions, the PI identified similarities between the data challenges and infrastructure needs discussed at the meetings and those identified by the stakeholders and also across different DAs.
During the spring and summer of 2017, the PI contacted the stakeholder committee of IIHCC and a few representatives in its closely related Data and Decisions Destination Area [24]. A total of nine representatives agreed to participate in a semi-structured interview. It provides a reasonable sample size for grounded analysis and inductive reasoning and meets the qualitative research design and sampling recommendations of Creswell [25]. These participants have broad stakeholder perspectives representing areas from computer science, mechanical engineering, biological systems engineering, building construction, automated vehicles and human factors engineering, transportation safety, industrial and systems engineering, social and decision analytics to statistics. Each interview lasted from 45 min to 1 h and occurred in person at the participant’s working site or otherwise through telephone or Skype.
This study employed the qualitative methodological recommendations of Creswell [25] and Yin [26] and applied the critical incident technique of Flanagan [27], with emphasis on typical scenarios, practical examples and the significant experience of human subjects. The interview questions centred around the IIHCC initiative, planning and development, and in such context, the supporting data infrastructure and digital library systems. Through focused interviews and in-depth dialogue with individual stakeholders, the PI was able to engage the participants in a design and visioning exercise for the libraries, and by doing so, figure libraries’ role into their thinking, reasoning and mapping process for the DA architecture and network design. Through this process, these stakeholders recognised, enriched and even added values to libraries, and as a result may potentially advocate for libraries in the IIHCC movement forward.
3.3. Significance
The research results draw upon the semi-structured interviews with key stakeholders, the PI’s involvement in the community as a participant observer during meetings and events and the analysis of communication documents including IIHCC vision statement and white papers, publications, reports and emails between members. Through the mixed-methods approach and triangulation of evidence, this research produces results that augment design thinking and visioning practice for data infrastructures and digital libraries beyond traditional boundaries. With the advancing landscape of intelligent infrastructure, this study provides pathways for integrating libraries into the complex, interdependent and adaptive cyber-physical–human integrated system. Framed around the mutual shaping of human–technology interactions, this study also contributes knowledge to libraries’ work towards building a complex adaptive data system for community learning and knowledge networking. As smart infrastructure development increasingly aligns with data mobilisation, information exchange and knowledge transfer initiatives that are at the heart of library innovation, this study provides timely conceptual framework, theoretical understanding and practical implications for any institutions to also transform themselves around infrastructure development and library transformation. The following sections present emerging scenarios around complex adaptive systems, intelligent data infrastructure and future digital libraries all in the context of infrastructure building for smart, connected communities.
4. Results and discussion
4.1. Human-centred communities
When asked how they would envision these different domains of knowledge to converge and work as an integrated whole, the interviewees resonated across the board and articulated the focal point and core value of human-centred community building and problem-solving. Essentially, they echoed and foresaw ‘the emergence of smart work sites where physical work systems (e.g., infrastructure, materials, tools, buildings, and vehicles) and computer-based systems (e.g., digital project information, IoT data, and augmented reality displays or 3D visualisation) intersect at human perception and cognition’ to solve societal challenges [28]: It converges into the community part. And that’s the key part it … to improve quality of life, to improve equality, and to ensure that we have a healthier community … It’s Human-Centered … It just has not been so explicit before. It’s about bringing together all these specialties and giving people core competencies that allow them to move in and out, understanding the human scale socially, demographically, technically for all the things that are getting ready to happen … it’s an important focal point.
Another interviewee articulated the essential relationship between intelligent infrastructures and human-centric design systems that are interdependent and interconnected. Illustrating by examples below, the participant described a system view of infrastructure with ‘physical interdependencies’ coupled with ‘human-mediated interactions’: I think the broad area of intelligent infrastructures and human-centric design systems is really important in the coming days … Future generation infrastructure systems should not be studied in isolation. They are best studied by looking at the infrastructure, but also looking at the citizens and human individuals who use it. Let me give you an example … now with smart grid, people can play with the meters … can decide whether appliances come on and off. So humans are actually in some ways adapting their behaviors to the infrastructure and the infrastructure has to adapt its behavior to the humans. This massive interacting system view allows us to represent both human components and the physical components, which is what we call ‘human mediated interactions’. We are interconnected, so if the power goes off it shuts down the lights on the road, which affects the traffic. It might affect the healthcare delivery system. So a problem in the power system actually cascades into other infrastructures. This is called ‘physical interdependencies’. … There’s one more interdependency that is typically not taken into account. We call it ‘human mediated interdependencies’. It means that dependencies between the power network and the transport network … also go to the humans who use them … We would like to study the infrastructures in this form, to study interdependencies, not just physical interdependencies, but also human mediated interdependencies.
The above view addresses mechanistic and human aspects of infrastructure along with interdependencies among them. Humans and their surrounding environments are increasingly connected through rapidly changing intelligent technologies. Realised through solid scientific and engineering foundations and concrete technological advancement, the goal of intelligent infrastructure development is to enable new levels of economic opportunity and growth, health and wellness, safety and security and overall quality of life. It is centred around empowering and equipping communities with the structure and ability to engage in meaningful and efficient ways with their environments.
At the same time, it poses significant challenges at the complex intersection of technology and society and at the intricate interface of humanities and environments. It requires integrative research that simultaneously addresses the technological, environmental and social dimensions of communities. Such communities ‘synergistically integrate intelligent technologies with the natural and built environments, including infrastructure, to improve the social, economic, and environmental well-being of those who live, work, or travel within it’ [29].
4.2. Complex adaptive systems
The study of cyber-physical–human systems and how they interact, combine and change should be anything but static. In such systems, how the different components converge and adapt to solve human-centred problems will be based on dynamic data capturing and seamless information flow. Data mobilisation will be at the centre of it all. From a ubiquitous mobility standpoint, one interviewee particularly emphasised data gathering capabilities of infrastructure: We need a lot of ways of capturing information. The ubiquitous mobility side is basically … obtaining information about people, about what people are doing in buildings. That’s the smart design and construction … For example, if you have an elderly person … maybe in the future buildings understand that that person is coming … The building might have data gathering capabilities … That same information can be provided to an autonomous vehicle. So when the building knows that that person has a doctor’s appointment … [it is] able to schedule an autonomous vehicle that knows that person has limited physical capabilities … and [thus] is able to get a wheelchair inside.
When asked how they envision data infrastructure and digital library systems in the IIHCC framework, the participants projected libraries as complex leaning and adaptive systems, embedded in ‘scientific living habitats’ in the digital world: First thing, hard infrastructures need to be done, [such as] a smart grid facility or a smart road facility, places where you can drive autonomous vehicles … But [as to] data infrastructure, we have embarked on the concept of scientific living habitats. What it means is that we are building live representation of cities and infrastructures in the computer. I think that has all the elements of a digital library, [with] tools managing the data, storing the data, [and] everything that you can think of. We’re doing this for the human-centered communities, it needs to be accessible … and adaptable, and something that learns from itself. That’s where the algorithms might come in … It is able to take information and see how people are using it, and then repackage it. Learn how people are trying to access information in the database and then say, ‘well, it’s obvious that this is an important set of information and maybe now we need to identify this area for more collection of information’.
Taken together, the perspectives above indicate how the data produced under IIHCC can support improved understanding and prediction of the interactions between smart physical infrastructures and the populations they serve, as well as how these data can contribute to better understanding of demand for infrastructure-based services. To achieve these goals, enabling data infrastructure and library system needs to be effectively in place. In particular, the digital data library should provide intelligence in data gathering and processing, information seeking and knowledge discovering. Such scenario increasingly requires the adoption of ‘smart’ concept in library architecture that is capable of learning and analysing human information behaviours, providing assistance and making suggestions or recommendations for new data collections or information gathering [18].
4.3. Intelligent data infrastructure
IIHCC needs to capture mobile location-based information, data streams from sensors, social media recommendations and various IoT inputs for different events. The event data can be from various sources such as buildings, traffic systems, autonomous vehicles, weather forecast and pollution monitoring. These massive-scale data are only beneficial when application-specific, actionable knowledge can be extracted from them. As a connecting bridge between data capture and knowledge extraction, data processing and handling in between thus becomes really critical and faces new challenges.
In particular, the dynamic and diverse nature of infrastructure data requires different data management techniques and advanced processing methods. For example, we need new approaches to search and crawling of massive infrastructure data from various sensors and sources as well as new methods for storing, indexing and query processing such data. As a transdisciplinary area of innovation, IIHCC also requires novel conceptual organisation and thematic clustering of data collections for sense making and intellectual mapping. Furthermore, new interface design and performance capabilities are needed to support interactive data analytics and on-the-fly visualisation for time-sensitive explorations and investigative work.
Under new data research scenarios, there are a whole set of emerging topics surrounding infrastructure data acquisition, storage, management and processing that are critical for the research community.
4.3.1. New data research scenarios
According to the participants, new IIHCC research scenarios especially involve live streaming of traffic, energy, building or other infrastructure data. They are often characterised by on-the-fly data processing, visual interaction, computational manipulation and analytical capabilities: [Not just] the amount of data or the availability of data, it’s the types of data that are available now, not only available, but available in real time. That really opens up areas of research more in the real-time analysis of data, as opposed to just looking at something more static, whose validity I guess, kind of dies with time, because the farther away you go from it, things change … That can be in the traffic area, modeling simulations, understanding traffic patterns, understanding what disruptions to deal with if there’s a crash somewhere. Now we have probes everywhere that can provide you that information real time as the events unfold. … [To] actually use the data to its maximum potential means things like being able to implement very quickly machine learning algorithms … [So it involves] how to put them into the live stream online data, then ways of visualizing the data in a much more efficient way … Some kind of live streaming of the data that would allow maybe some first level processing of the data that would allow it to be plotted for visual interaction quicker. Then how to generate these algorithms, and cross referencing and matching things, and being able to create a lot of computational on the fly capabilities on the data, that would be great.
These featured scenarios are time-sensitive and require rapid throughput and quick turnaround of data workflows. Typical of a self-adaptive system, the associated data infrastructure thus needs to perform context-aware and agile management of content, from automatically harvesting real-time data, to algorithm-powered first-level processing, to smart recommendations for further investigation, all of which constitute an intelligent data life cycle. As a key data steward, libraries need to engage in the new data scenarios and integrate into the intelligent data life cycle.
4.3.2. Smart data accessing, processing and visualising
On top of automatic data harvesting and ubiquitous information transmission, new infrastructure requires ‘smart data’ searching, mining and drilling from big data. Although recently ‘unprecedentedly large amounts of sensory data can be collected with the advancement of the Cyber-Physical-Social systems, the key is to explore how big data can become smart data and offer intelligence’ [30]. Smart data aim to filter out the noise and produce valuable content, which can then be effectively used by human agents. In this aspect, frontier research has already started exploring brain-inspired or nature-inspired computing for harvesting smart data. The current participants recognised the needs to incorporate such advanced techniques and AI automation into the IIHCC data processing and visualisation framework: I think the next level of dealing with this data is high performance online, online as in live computations, and some kind of magnified cognitive capability, so [that offers] the ability to visualize the data, to hear the data, to interact with it using more than one venue. For example, we were talking about being able to walk around with Google Glasses and see the data, not only see the raw data, [but also] see the processed data. We study building movement, so I would like to be able to interactively see that data processed live. How we do that, at this level or magnitude of data, there has to be some kind of improved cognitive capabilities on how we interact with the data. I don’t think it can just be 2D window of a computer. We need some kind of smart data, smart visualization of the data, some next level interaction.
The above anticipation suggests a highly streamlined and cognitively magnified data platform that enables the simplification and dissemination of highly complex data sets. Such a system is equipped with high-performance processing and multi-spectral visualising, sensing and interacting capabilities. It can reduce content complexity to functional simplicity [31] and help users discover insights, gather competitive intelligence and explore visual highlights. To realise such a system requires multi-functional capabilities, the most prominent of which are data processing solutions, such as rules for relationship discovery, data formatting, visual generation, report calculation and so forth.
Notably within this platform, digital library solutions are indispensable for discovering the underlying structure from retrieved data in order to acquire smart content. Particularly, solutions as to how data attributes and characteristics are presented and how taxonomies are designed as well as how navigation is orchestrated all have downstream impacts to the user experience. Undoubtedly, data classification and taxonomy are an essential component of data science and its foundations, and together with query and indexing technologies, can support big data processing and analytics. To develop such an all-around versatile data platform, increasingly we need to collaborate across major functional responsibilities to create an optimal digital experience.
4.3.3. External data linking, harnessing and presenting
As noted above, the IIHCC work should be community-focused and purpose-driven towards addressing public needs and societal challenges. Thus, external sources such as public census and government administrative data naturally become valuable components of its data framework. As shown in the examples given below, the IIHCC stakeholders here are critically engaged in curating diverse data sources and scientific workflows to provide and sustain public services or industry applications and to generate values for society through data sharing and reuse: There’s a really important source of data – the vast amount of administrative data that exists at the local, state, and federal levels. How I like to characterize it is that having all of this data and building rigorous theory and methods to use the data – very disciplined ways to use the data, what it’s providing for us is like a new lens for observing the social condition. I like to draw the analogy to when the Hubble Telescope was released, it looked deeper into the universe than we could ever see before. Public, the census data, all the different data sets that are public and run by the federal government are very important. They’re critical for us to be able to understand housing markets and what’s happening demographically … There are groups that are able to harness that and present it to people in the housing industry in a way that they can use it, that’s something we do a lot.
By such accounts, the IIHCC data infrastructure needs to enable access to and spur use of important and valuable external data assets where relevant, whether it is about building construction and built environment, or sustainable planning and environmental innovation. This defines the need to intentionally and proactively seek out existing sources and discover values in data, internal or external. It also requires disciplined approaches and rigorous methods to ensure scientific rigour and empirical soundness of using varying collections and fusing multi-source content. In this regard, knowledge bases such as Wikidata and YAGO are becoming very popular and could provide valuable IIHCC-related data if carefully curated and rigorously vetted.
4.3.4. Software and system preservation
In the data curation realm, there have been significant advances in recent years on managing, preserving and making available data [32]. However, such progress is lacking on the fronts of other intellectual products, such as code detailing the research workflow, software for analysing and visualising the data and environment for running the software. In the case of environment for running the software, it further includes the operating system, function libraries and software dependencies. These products are highly interconnected and interdependent and thus need to be preserved and presented together to function as an organic whole. Their preservation challenges are reflected in the scientists’ continuing workflow struggles as described below: I’m more interested in preserving the tools and the capabilities. How do you perpetuate that kind of capability? So, [finding] a better way to keep the software around and the whole environment in which the software runs is a challenge. Because of bit rot. Because stuff gets old and won’t run anymore … [So how to preserve] the context, the stuff around the tool, the operating system in which it functions, the libraries that it relies on, and all those software dependencies? … data people, they want to publish data findings. We here want to publish demos. Here’s a tool that you can use to view data stuff with. We want to publish the tool itself, [so that] other people out there can try it. They might even be able to plant [it] in their own data. That would be amazing. We don’t have a good way to do that now. The way we end up doing now is recording a video of somebody using the visualization. And then we publish the video and stick that online. But that doesn’t help people actually use the tool. It just preserves the look of the tool, so that somebody could rebuild it later. It’s pretty low fidelity preservation.
While performing visualisation, the participants further highlighted the importance of conducting different types of data pulls and the interactive nature of data explorations. Interactive data performance is highly contingent on the software and its operational environment. In intelligent infrastructure, it will surely become even more important to preserve and publish software and its running environment to forge coherent data analysis and dynamic visual presentation. Such development will also enable viewers themselves to exercise customisable visualisations or improvise dynamic simulations while empowering them to deploy available tools for new explorations. It will provide an effective means and appealing intermediary to engage the public in scientific experimentation, visual demonstration and idea exploration.
4.3.5. Dynamic data documentation and robust provenance tracking
In scientific community, it has been widely recognised that provenance plays an important role in many scientific applications and use cases and enables reproducible research [33]. Likewise, the interviewee below affirmed the importance of thoroughly and systematically tracking, describing and preserving the lineage and processing history of data, including analytical tools, intermediate results and knowledge artefacts: … having provenance on the data is important … a whole history for every piece of data, of where it came from and what’s been done to it. So you want to keep all those intermediate results and the process that produced all those intermediate results. There’s a whole provenance set of information to preserve. Not just the data, but also the tools … a tool takes some input and produces some output, so you gotta keep that whole chain of data and tools preserved. And anytime somebody touches or modifies something, you want to add that to the record and have that whole history listed out.
Another interviewee further characterised the dynamic nature of data gathering in smart building environment where instrumentation and calibration are constantly adjusted and modified. As such, documentation practice in correspondence needs to be flexible and reliable, speaking authentically to the data in real time and reflecting faithfully to changes on the fly, all by a living document of track records: Our documentation needs to be very good. It needs to be a living document that changes a lot over time. For example, the way I gather data, things change drastically every day. Sensors go bad, there’s new instrumentation [and] new calibration. If data is gonna remain reliable, then the infrastructure of the data gathering needs to be reliable [and] the documentation of it needs to be reliable. When a sensor goes bad, you have to switch it … and you need to make sure that the whole documentation acknowledges that switch. Then that data remains reliable and comparable to past data.
Above all, dynamic data documentation and robust provenance tracking will be imperative to ensure reliable and comparable data flows and thus will assume a central place in an agile and sustainable data environment.
4.3.6. An agile and scalable data library system
With all the emerging scenarios, today we face an explosion of complexity, particularly with respect to data in myriad forms pulled from many sources and complex dependency webs. Scholars thereby emphasise the significant value of workflow efficiencies in handling data. To solve the complexity of data, analysis and software issues for IIHCC, a high-performing, living and adaptable aggregation system is needed in the name of increased simplicity and utility.
In this emerging enterprise, the IIHCC data library needs to be scalable and resilient in and of itself. It needs orchestration software that can integrate diverse set of infrastructure-enabled data assets into cohesive process, so that human agents can focus on running the business of discovery and innovation, not manually managing data bits [34].
Looking ahead, key use cases and practical applications from IIHCC will certainly drive a spike in the deployment of data library solutions. These include collecting, documenting, preserving and publishing infrastructure data sets, protocols, experiments and computer code (such as algorithms and heuristics for creating simulated or synthetic data), along with related tools used to generate and analyse these data.
4.4. Novel concepts and future scenarios of smart digital libraries
Then, what are the future scenarios of smart digital libraries? There are three main concepts emerging from the interviewees’ responses, including the digital octopus concept, smart village concept and virtual assistant concept, which are detailed in the following narratives.
4.4.1. Digital octopus concept
In this first concept, library acts like an ‘octopus with a central head and multiple tentacles’ reaching out; pulling all the data, expertise and knowledge back in; and bringing them together. Here, library is a distributed system embedded within the cyber-physical-integrated environment to collect data and combine information, push these back and forth and communicate them across: … embedding [library] with, say, a facility like a smart road facility, where the challenge, in my opinion, is to say, ‘What does library have to do with placing a unit along with VTTI [Virginia Tech Transportation Institute] facility in the smart road? What role can it play right there?’ It can play a role there in terms of taking all the data that they’re producing and curating it and adding it to the digital library. That’s a beautiful part of the library, because library now is not sitting in that place … but now it’s extended. It’s almost like an octopus where there’s a central head and all the tentacles are going out and touching all the different cyber-physical learning facilities and pulling the data, pulling the expertise, pulling the knowledge back in, and bringing it together … It’s all spread out. It has gotten a distributed view … It should be embedded with the cyber-physical systems. [It is] something which touches everything at the same time but combines the information also to some extent. It’s pushing information back and forth, but they’re talking to each other as well across.
The participant further elaborated the digital octopus analogy and explained the central and expanded role of library in the smart landscape. Particularly, ‘library in its conceptual way of being the place where data and information is generated, stored, curated, accessed’ is virtually embedded everywhere in research spaces. With its ‘tentacles’ reaching far out and its functions embedded deep in wherever data are created, all the scientists are virtually participants in this library, building out data, curating it and adding to the large system: To me library is not the traditional one. Library without walls is the right concept. The concept of library as a way to store, curate, manage information is central to all the Destination Areas … which means those functions of the library should be deeply embedded within these areas. That’s why I gave the analogy of an octopus, because aspects of the library that ever touch with data curation, collection management in each of the spaces that are doing research are actually part of the library system … they are actually the places where things are being collected … So now the library information is being curated by everybody. All the scientists are building out information, curating it, adding to this large system, and I think that’s how library should be viewed as. Library is not an object in a building right there, that’s one part of it, but is embedded deeply within institutions and spaces to provide, curate, store, [and] access data and information.
In this view, library-enabled processes primarily occur in on-premises research facilities and spaces. Here, a library can function and act locally based on data it collects and take advantage of on-site data ingestion, curation and storage. This is particularly meaningful as ‘more enterprises will push processing and analysis of data to the edge of the network in order to cut data ingestion costs and reduce network latency’ [34]. We will see increasing deployment of research processes requiring local data handling – from provision, curation, storage to access and utilisation – close to the connected devices that enable these processes. This requires us to rethink data challenges and unified user experience no longer as a data management issue but instead as a data and system integration solution.
4.4.2. Smart village concept
In the following smart village view, library will be the ‘heart’ of all the data, bridging gaps between users situated on the different spectrum and dimensionality of data workflows. It will function as a central hub for data gathering, processing, sharing and communication: For me, the library ends up being the heart of all this data. I built a very funny cartoon off of what I thought the new smart village would look like … In that village, we’re gonna have data and smart construction and the city of the future … For me, the central hub of that park was all the data … a hub where all the data, all the processing, all the sharing happened. That was drawn in the center of this village. The computer science and electrical & computer engineering people are more interested in either the algorithms or the hardware. I’m interested in the use of the data. No one is in the middle providing the infrastructure to make all that happen. That’s where I see the libraries [should] be. It could be a naïve way of looking at it, but I don’t see anybody else being able to fill in that gap, to create that infrastructure and the communication of that infrastructure for all of us.
By referencing the cartoon drawing, the interviewee further interpreted what the smart village would look like. Here, library provides a data hub and infrastructure solution that spans communication, development and computation in an integrative manner and supports creative synergies across the many components of the village: This is a village … all these buildings have data information. There’re smart buildings, smart cars, smart intersections, smart motorcycles, smart manufacturing, smart power plants. But all that data needs to go somewhere, so all that data goes here [a center space designated as ‘state of the art data facilities’ right beside a ‘creative synergistic space’ in the village] … here’s high powered computing, data archiving, storage, processing, cloud and all that stuff … The infrastructure and the hub that brings this all together in the era of big data, I think it’s the libraries … The library has to go anything from cost of design to actual implementation, so they play the spectrum. I think of [libraries] as a center hub of communication, development, [and] computation’.
Placing it at the centre of a smart village in this scenario, the interviewee vividly framed the unified and magnified role of library in a data-oriented perspective that encompasses different layers: from the technologies enabling sensor data collection, to the communication architectures needed to support data transfer, to the reasoning techniques applied to data in order to generate the information upon which services and functionalities are built [35]. Here, ‘playing the spectrum’ from conceptual design to actual implementation and serving as ‘a central hub of communication, development, [and] computation’, the concept of library is re-defined, its image is re-drawn and its existence is fully embedded in research data facilities. The reality is, while the data are there in disparate systems, the insights and connections to users facing applications are not [36]. In this consolidated view, library as a data and information broker works to bring connections, deliver unified and personalised omni-channel experience and inspire creative synergies.
4.4.3. Smart recommender and virtual assistant concept
‘In an age of both information overload and siloed systems, aggregating and accurately reporting on research output’ are challenging [37]. Struggling to function under conditions of electronically mediated data and information overload, researchers look to smart library and information solutions. Such solutions can spontaneously self-assemble constituent pieces of data into a knowledge package with a particular property and then bring it to the individual with relevant interest. Inherently, this concept reverberates the autonomous system and ubiquitous mobility themes of IIHCC. In the scenario below, the researcher expressed the desire to have specialised library navigation and suggestion systems that can proactively ‘push’ knowledge to users and provide personalised references: Looking for references right now is a passive task. What I would like to have, as I’m typing a report, is a feature that is able to understand what I’m writing … processing that information, and say, ‘I’ve found five references. Two of them agree with your statement, and the other three are different. Would you like me to look for more?’ Or further, ‘Would you like me to bring those articles to you?’ … Or say, ‘Hey! Careful with that statement’. [So, it’s something like] IBM Watson… the Watson version for different domains … And it would cut down the time for me to go to millions of other libraries to try to figure out what other people have put there. So, that is intelligent infrastructure for human centered communities … to provide services … to keep people informed, to gather information from the environment and make better life for elderly, for young people. We need to make data information more proactive, because right now it’s more reactive. The people are pulling the data. We want to push information to people instead of waiting for them to pull it. We’re trying to create a sense of community. The library should be everywhere. You need that virtual component … overlooking at what you are doing and helping you, saying, ‘Hey, you should consider this new article. It just came out, and it’s very relevant to what you are doing’. Libraries are about people gathering information.
The above view resonates with the growing adoption of AI-powered virtual assistant or chatbot technology in many industries. Such technology can provide users with personalised information review and content summary from multiple sources. No longer a distant future, the new AI landscape is already exploring how smart search engine, machine learning, text mining and computational simulations with the latest optimisation strategies can identify knowledge gaps in a series of scholarly publications. Coupling AI modelling with personal interests, the virtual assistant can create a more systematic way of discovering, assembling and recommending relevant data and new information that exhibit specific desired properties and research attributes.
With improved text summarisation, such breakthroughs will help reduce information overload and improve productivity by automatically reviewing a history of research themes and providing a global sense of the documents. For example, based on personalised data sets and AI modelling, the technology will be able to summarise numerous technical reports, literature reviews or multiple emails and documents needed to be reviewed and subsequently recommend research gaps, needed data, or suggest new hypotheses [38].
Adding to the many possibilities, this smart recommender and virtual assistant concept can very well be integrated and applied to learning systems, forging a unique and valuable living-learning experience for students. Bringing information to users, this is what future digital libraries should be able to accomplish. We can achieve such promise by focusing on the interplay of pedagogy, technology and their fusion towards the advancement of smart learning environments: … provide more information to the students … You’re collecting data on the students. You have which classes they’re taking, what type of homework and projects they are doing … and are able to do clusters of information … Then send a communication tailored to that person to say, ‘These are the most recent articles in areas that you might be interested … related to your project, or your statistics homework’. All the students have a student card … that card is walking with them all the time … gives them access to the library and to food, and they might have a calendar that says what they’re planning to do. [With all these information, the library system can] say, ‘Hey, you have one hour, and I just checked, the three people that you’re working with on this group have time. Would you like me to get you together at the cafeteria? And by the way, I already pulled all the information from the library that you need for that project’. That’s the human centered community part, [it] is having the information be there for people, assisting people. They don’t have to go and gather the information. It’s push, not a pull.
By and large, the ongoing digital revolution ‘has accelerated recently with the tremendous increase in electronic data, the ubiquity of mobile interfaces, and the growing power of artificial intelligence’ [39]. AI is shaping the future of everything from medicine to transportation to education and thus must be figured into future information network and library data architecture. This includes pushing notifications and delivering content to users who are in search of certain information. In other words, content distribution and delivery workflows may need to reform with the changing research process, with the evolving information access and knowledge distribution dynamics and with the advancing discovery experience and expectations. New academic environments require us to proactively and seamlessly figure library’s omnipresent roles into the research life cycle of scholars.
5. Conclusion
5.1. Next-generation data and information user experience
Within the IIHCC framework, an increasing number of academic programmes and areas will converge under newer, broader and more dynamic alignments: a cyber-physical–human-integrated ecosystem. Libraries should be part of a continuum of researcher experience and technology innovations and have a place in this integrated infrastructure for teaching, learning and research.
The flurry of recent strategic moves reflects the growing importance of user-centricity and the appreciation and expectation of people for a more seamless user experience. Particularly, with the DA’s prime focus on human-centred communities, the library and information world of this IIHCC ecosystem should be a highly user-centric model. Ideally, within this model, users can enjoy an end-to-end experience for a wide range of data and information products, services and applications through a single-access gateway, without leaving the ecosystem.
In this ecosystem, users will experience live streaming of traffic, energy, building or other infrastructure data and perform on-the-fly data processing, visual interaction, computational manipulation and analytical exploration. With relevant data and desired information automatically and systematically pushed to them by AI-powered virtual assistants, users receive personalised information summary and customised recommendations on research gaps, needed data or new hypotheses. In this smart, connected community, scientists are virtually participants in the digital library by conducting on-premises data ingestion, curation and storage and continuously contributing knowledge and expertise back to the integrated system.
5.2. Smart infrastructure data challenges and solutions
Developing IIHCC entails creating vast collections of infrastructure, environment, human activities and their interaction data, where the totality of an individual’s experiences, captured multi-modally through digital sensors, is stored systematically or even permanently. These data, for example, could be the description of the semantic locations and physical activities of human in a building. The diverse and dynamic nature of such data requires a higher-performing, living and adaptable collection, documentation and aggregation system in the name of increased utility and workflow efficiency.
On top of the daunting storage and processing task, one additional challenge is infrastructure data retrieval and summarisation. For example, users may need to analyse smart building data according to specific queries; or perform analysis on certain entities, entity properties or attributes of the data set; and make summarisation of them according to specific requirements. To address all these demanding tasks, there is the need for systems that can automatically analyse this huge amount of data in order to categorise, summarise and also query them to retrieve the information that users need.
5.3. AI capabilities in future digital libraries
In the changing landscape, AI is still in many regards a nascent technology but sure is transformative to empower better decisions and change the entire enterprise [40]. In this transition, libraries need to embrace AI to stay ahead of the game and become a driving force in smart infrastructure development and faculty engagement. There are two prime reasons behind this idea.
First, as data preservation and curation requirements grow complicated and the data itself increase in size and complexity, AI will bring the power of better decision and efficient execution in managing, processing, extracting, summarising and discovering data. Such examples could be dynamic classification of variable data using evidence accumulation approach or automated detection and monitoring of attributes, classes and errors in data set based on incremental learning techniques.
Second, as the DA envisions the development of innovative discovery channels, creative learning spaces and embedded cyber-physical–human systems, AI will surely play significantly into the fabric of teaching and learning experience of faculty and students. Whether in intelligent built environments, autonomous vehicles and robotics systems or smart living and transportation networks, AI can provide methods for processing data from autonomous sources in distributed environments and offer ‘tools for working out consistent knowledge states, resolving conflicts, and making decisions’ [41].
5.4. Future direction
All in all, in the emerging smart infrastructure landscape, researchers are increasingly turning to technologies and tools to harvest and harness data and information in their quest to stimulate scientific discovery. With all the activities involved and anywhere connectivity realised, data and information tie everything together and need to flow seamlessly between systems. Future digital library, whether in its conceptual form of a digital octopus, or a smart village data hub, or an intelligent virtual assistant, is an integral part of this infrastructure. It will provide intelligence in data gathering, processing, summarising and recommendations, and by delivering unified and personalised data solutions, it will offer an end-to-end seamless experience for users and learners throughout their journey of knowledge pursuit.
By providing novel concepts and analytical insights, this research advances the theoretical conceptualisation and empirical understanding of scholarly data practice and information management in a distinctive new context. It contributes to a new modelling of digital library systems in the emerging horizon of disruptive technologies and smart environments. Situated at the intersection of data and information technologies, socio-technical infrastructures and smart and connected communities, this research poses a new direction of research, development and innovation for the library and information field.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship and/or publication of this article.
