Abstract
This article proposes and formulates the visual urban vocabulary for tacit, intuitive, experiential, but none-the-less fast, plausible, generative, informative, sketch-like composition and visualization of urban stories. Through visual and socially ‘inherited’ clues, the authors explain the complexities of urban spaces, their elements, interrelations and cause–effect phenomena to expert and non-expert public alike. The rules, syntax and overall advantages of such a vocabulary are grounded in the existing linguistic, cognitive, psychological theories, visual sociology and theories of urban design, combined and supported by the authors’ own research into visualizations and tools for evaluating, understanding and presenting urban spaces. With many illustrations, the article demonstrates the use for – and the use of – generic urban stories in discussions about urbanity, urban environments, livable places, etc. and positions them into educational, research and participatory planning and commercial contexts.
Keywords
1. Introduction
Places with multiple meanings might be mental constructs but, nevertheless, their understanding is primarily based on visual stimulus and experience. From complex sceneries to plain elements, they help observers to interpret what they are seeing, extract meaning and anticipate its implications. We begin our story with the basic presumption that, by carefully looking at the image, the observer would not only see the scene per se, but could, through tacit knowledge and experience, implicitly discern how the place works, what it offers, how it functions, where it is (in general terms), how it is organized, who the users are, what they do, how they use it, how it is meant to be used, what infrastructure is present, what kind of place it is, what the centrality of it is, the general character, and so on. The reference to the work of Cullen (1971) and his propositions about experiential perception of urban spaces through serial vision and its emergent townscape is mandatory at this point in terms of inspiration, but also for the creation of foundations upon which we build our discourse.
It so happens that the photograph in Figure 1 is the actual representation of a specific moment of Ljubljana’s old town district. Yet to an unfamiliar observer, if the location is not explicitly stated, the photograph could be representative of any urban scene with similar flair and layout in Amsterdam, Copenhagen, Hamburg, Bremen, or even Prague, to name just a few. For the unfamiliar observer, the photograph subsequently represents a more or less anonymous generic scene, in a generic place. Rather than argue whether it is in fact generic to such an extent, we can focus on the questions of why it could be generic and why that would be beneficial to us as planners and users; what we can learn from analysing it through our tacit knowledge, experience and unfamiliarity with the place; and if and to what extent we could persuasively recreate it by following a set of rules and parameters.

The inspiration for the visual urban vocabulary came from the photograph: what can be a plausible, generative and informative representation of such an (and in fact any) urban place that can be widely used, quickly generated, but also widely understood and tacitly perceived by non-experts in a more or less unified way.
Overall, the image represents an attractive, vibrant and pleasant space, yet there is more to be gathered from the photograph: physical layout; composition and organization; character of the buildings; their volume and shape; general age and type; differentiations of streets and levels, to name just the most obvious physical characteristics. However, there are several more clues about the functionality, use and access: shop signs of differing type indicate different services such as gastronomy and retail, large umbrellas indicate seating areas where food and drinks are usually served, and the users can be seen sitting, walking and reading. Some of the observed activities are even more specific and indicative: a mother pushing a trolley; an observer capturing the spectacle via photography. A lot can be discerned from inanimate objects: bicycles parked in the distance; water slowly flowing underneath; trees casting shadows. All of these combined generate a lively and vibrant place where people like to go to and stay, which is represented by the numerous people sitting, standing, mingling and enjoying themselves. We can conclude by the sheer number of people that the place is popular, liveable and much liked. Much of that is purely our speculation and tacit deduction, yet an informed one, regardless of whether we have visited the place or not. As we have been able to read the story about the real but unfamiliar urban place represented through visual representations following the tacit clues, we could also read the story of a hypothetical proposal represented in similar ways and by similar means. We could compare the two places, or even evaluate the place in terms of its liveability, access, amenities, etc. and this is our basic assumption. Generic representation and interpretation also have their limits and explaining the part of the story relating to a place’s unique identity is one of the most obvious ones.
The interpretation of photographs which ‘tell’ a story, and the use of photography in social research, has been thoroughly advocated before. Visual sociology argues that images can be read as texts in a variety of ways. Becker (1974) pointed out the connection between photography and sociology but also delved into the scope of photography being used as scientific apparatus. In regard to composition, Becker searched for analogies with languages and music, coding conventions and interpreting – reading – the photographs; thus he reaffirmed the notion of visual language metaphor and brought to the fore the difference between ‘reading’ by experts compared to non-experts, which are very much of interest for our research.
Before him, Cullen (1971) had described similar observations, specifically for urban environments and for the emerging urban design. He initiated the term art of the relationship and through photographs and sketches showed urban scenes and interrelations of elements through urban situations. His aim was not only finding out the facts about the place but also learning from it for the purposes of design and future interventions. Cullen operated with notions of higher abstraction that require a higher level of comprehension and expertise (i.e. the notion of entanglement, kinetic unity, animism, etc.). He was, of course, speaking to an expert audience, with such expertise required in order to follow the author’s argument. We, on the other hand, want to explain relations and stories to both experts and non-experts. In this case, it is useful to operate with a vocabulary that can span the expert–non-expert divide in terms of explanations starting from simple, tacit and more direct connections that do not require much expertise, leading to more complex, abstract constructs of higher order such as Cullen’s publicity. Such a vocabulary can explain and mean something to someone even though he or she never reaches the highest order of meaning, but rather stays somewhere in the middle on the understanding scale. The scene may well ‘say’ Cullen’s kinetic unity to somebody, while to someone else it might read as this place visually continues to another one, which creates something dynamic or lower on the comprehension scale: this place opens up to another street or even more mundane: oh, there is another street.
Cullen’s art of the relationship between urban elements represents a certain quality of the urban places. He even touches on anonymous urban objects, such as advertisements, but presents them mostly as objects that form an urban, primarily visual drama. As it will be later argued, we differ from Cullen on this notion, and want to show that such objects also inform us about relationships and indicate their use. In other words, in representations, urban elements are not (or should not be) used as décor and paraphernalia of drama only, but primarily as content messengers.
A crucial distinction has to be made at this point between our research and intentions in relation to Alexander et al. (1977), Cullen (1971), Boeminghaus (1978) and others who used urban scenes, patterns, symbols and representations, and classified them into vocabularies of design notions, principles and scenes worth recreating. Our proposed vocabulary is meant for the representation of the urban environments and places, to show how they work, their qualities, the interconnection of their parts and elements, their potential, their comparison and the related discussion with non-experts. With its specific purpose it differs greatly from the attempts of forming glossaries for successful urban design. This is why it needs to be emphasized again that this vocabulary is not in any way an urban design recipe. Here we are dealing with the understanding of (usually already designed) urban spaces and urban relationships, not with the design of them per se or the design of their individual urban elements. We are merely coding or writing the spaces in a manner so that they can be understood by non-experts or compared with other similarly coded places, trying to equalize the forms and elements in them to see them through their use, not through the explicit form of their elements or socio-historical identity.
Returning to the Ljubljana photograph, it is easy to imagine that, by not being limited to the given photograph, using other representative ways, such as drawing or sketching, we could show the observer even more information in a purposeful manner, giving more visual clues within the same scene and by using different kinds of representation techniques that achieve greater visual legibility and generate less clutter.
At the end of the article we go further and propose that we can recreate the scene, given the correct input. By having a descriptive parametric model of spatial relationships, we can generate – a notion intentionally used – the user simulants and with the help of an automated model for generating visual representations produce the image with a similar urban flair. For now though, let us just focus on how to tell the complex and comprehensive urban story to non-expert observers through the use of tacit visual language, their existing knowledge and experience, by providing them with sufficient visual clues or indicators and explaining what this elementary urban visual vocabulary should include.
2. Plausible Scenery, Presentations and Levels of Details
Flight simulators and different, equally immersive games are using the notion of plausible scenery, relying on artificially created and partially auto-generated environments, the characteristics and methods of which could be modified, adapted and put into practice by the urban design and planning community. Such approaches can be of practical use when it comes to the visual communication between experts and non-experts, generating fast, meaningful, experiential and tacitly understandable images of the urban settings in a wide variety of scenarios, such as explaining before and after situations, educationally explaining cause and effect relationships to non-experts, showing comparison of different proposal solutions, etc.
The aforementioned plausible sceneries use specific viewpoints, and a wide but finite range of elements (objects) deemed essential to recreate the feeling needed for the game/simulation experience and immersion while also being adapted to the movement, scale and speed of the observer. They use generic buildings, elements, greenery, urban furniture objects, etc. where possible and rely on the actual underlay or distribution logic to place them on sites. This quickly and effortlessly recreates almost infinite numbers of urban, rural and landscape situations. At the same time, they are relying on the tacit knowledge and experience of the user, who is used to interpreting visual elements and in-world surroundings, to come to the conclusion about function, action possibilities and use through visual and sometimes very clichéd hints and clues.
They are also using the principle of just enough objects and level of details essentially needed to recreate the urban image and not (always) trying to recreate the photorealistic environment with every possible detail. In this way they are optimizing the performance to common technological minimums but also optimizing the time and efforts needed for the preparation of such scenery and consequently reducing the cognitive load on the reader. As Becker (1974) explains, every part of the image carries some information that contributes to its total statement, and a proper ‘reading’ by the viewer is to see and respond to it consciously.
There is a trade-off, of course, when it comes to the identity of the place that is usually defined with particularities. However, there are also advantages when it comes to questions about processor loads and times for creation of the scenery and their fast reproductions.
Using the analogy of aforementioned flight simulators and games, it has to be noted that the digital city itself (e.g. buildings) in computer games is not always generically created but the urban elements are repetitive and can be seen as generic. While the games try to recreate plausible scenery for the sake of better immersion of the user, we need and expect more from elements that we use in representations of urban environment. On the other hand, the urban spaces we want to recreate are not there just for the scenery and for the support of the main character of the story or for our plane to fly by but are the main purpose of the representation and there to tell the story about themselves and represented urban environments, purely through visual means. The movement through space itself can be regarded as a semiotic resource (McMurtrie, 2013) and the foreseen paths in simulations or lack thereof can greatly influence the story narration and experience.
Research (e.g. Mullins et al., 2002; Strothotte and Schlechtweg, 2002; Ucelli et al., 1999) has shown that representations of elements and urban landscape should be kept simple and with just enough detail – sketch-like or line drawing representations are thus natural candidates. Peterson and Kim (2001) state that early in the course of perceptual processing, certain regions in the visual field are assigned figural and others ground status. They have different properties. Line drawings and sketch-like silhouettes are, in our case, essentially figures with definite shapes which, according to Peterson and Kim, if familiar, can be recognized and appear as things or objects. Singh et al. (1999) also describe the complexity of the object recognition and characterize it as a computationally demanding process that uses cues such as shape, colour, texture, motion and context. But the ease with which we recognize an object without any cues, only shape, leads the authors to conclude that shape is a key aspect of our recognition. Polonsky et al. (2005: 1) state: ‘There are many possible 2D views of a given 3D object and most people would agree that some views are more aesthetic and/or more “informative” than others’ and continue by stating that in a computer vision community, good views are presumed to be the ones that make an object more readily recognizable by humans.
In design practices involving both experts and non-experts, sketch as a medium has certain connotations that are beneficial to the reception of ideas. According to Buxton (2007), following sketch characteristics has clear advantages: they are disposable; include minimal detail; have personal touch; have clear vocabulary; suggest and explore rather than confirm; and are ambiguous. The representations are also better accepted by non-experts (Mullins et al., 2002; Verovsek and Juvancic, 2009) if the visual language used is unified and does not mix too many different representation techniques (photorealistic with line drawings, shaded with unshaded models, etc.). However, the level of details needs modification depending on the speed of the movement through the scenery by the observer and the distance between the two (more in Luebke et al., 2003). Plausible reality for the flight simulator’s point of view and speed means something different to a pedestrian in an urban environment – the scope of elements/objects and their details differ greatly.
3. The Visual Clues or Elementary Urban Visual Vocabulary
In the visual representation of urban spaces, we usually use the elements and objects such as urban furniture, signposts, pavements, greenery, people, as either part of the design or more often as decoration (we can also use the term beautification). In the representation of existing environments, however, we try to omit them either because they are in the way, deemed not relevant for the same reasons as above, or we want the place to look better and they are not helping. But these elements themselves – not necessarily invented and designed specifically each time, but represented as acceptable, contemporary and sensible archetypes – can tacitly tell much more than we give them credit for (see Figure 2). It is this anonymous and, in today’s globalized world, very generic landscape of elements – from shopping bags, trashcans, lampposts, sign posts to advertising boards, etc. that gives us visual clues about contextual background, space use and urban flair, regardless of whether we are discussing urban places in Asia or Europe. They can be regarded as visual and readable, as well as meaningful language, a sort of Esperanto of our urban landscape around which experts and non-experts can build their discussions using almost no words. Therefore we propose building visual urban stories around them or, more precisely, with their help.

The richness of simple and basic visual vocabulary can be enormous, hinting at the age, social position, activities, accessibilities, affordances, function, context, mood, form, etc.
For the reason we attach and want to transfer not only the form, but also tacit meaning, implications, connotations, relations, etc. along with the visual representation of an urban element, our suggested elements are to be purposefully chosen and combined. They should be formed and tailored around the user perspective and their previous experiences, alluring to previous reactions to urban spaces and almost genetically inherited, at least to a modern dweller, phrasing of our urban vistas and of their interconnected parts (see Figure 3).

Showing the urban space in use by user simulants in combination with simple attachments and accessories can intuitively reveal to non-expert observers much about the urban contexts they are depicted in.
Using the metaphor of a visual urban language we can go as far as to describe that the elements form the basic urban visual alphabet – letters. Combining them, we form words, multiple words form sentences and through the sentences an urban story with its complex and multi-layered meanings, connections and interrelations is told visually.
The language metaphor should not be understood too literally but rather as a principle, a set of basic rules, which allow for order, structure, composition and hierarchy, but also for the creativity in creation of messages. Kress and Van Leeuwen (2006) have established a similar connection between grammar of language and visual ‘grammar’ where depicted elements combine into visual ‘statements’ of greater or lesser complexity as it happens in language where words combine in clauses, sentences and texts. Instead of letters, words, sentences, paragraphs, etc. which are compositional elements of written language (Halliday and Matthiessen, 2004), we could also use the principle of rank which operates with similar hierarchical elements of morphemes, words, phrases and clauses. The language analogy mainly interests us because language and grammar as established constructs are easily comprehensible by a non-expert and can be conveyed with fewer explanations.
To clarify the notions, we should look at the syntax through a couple of examples and describe each level of the composition later. Let’s assume the bus stop sign, the bench, the illuminated advertising panel and the small roof cover are our basic elements – letters, which form a composed element: bus stop – a word (see Figure 4).

The visual urban alphabet – an example. Having many meaningful but nonetheless mundane and generic objects in representation of urban places has an effect of telling the story, explaining how the space functions and offering an enormous amount of multi-layered, meaningful information solely by relying on the tacit knowledge of the observer and his experience.
This explains that public transport is a part of a given urban environment; the place is connected with other city parts through this node. Adding the numbers painted on the bus stop, letters again, will denote bus routes and extend the word bus stop to a sentence: the place is well connected through several bus routes – public transport – with other city parts. With a combination of several such visual sentences that address connection, we can tell part of the story about this urban place’s accessibility (see Figure 5).

The ‘syntax’ of visual vocabulary – from ‘letters’ to ‘words’ and onwards to the ‘sentence’.
Not just by using a single element and combining several into elements of higher order, but also by the iteration of elements at each level of order, we can complement the intended message and form meaning (see Figure 6). As mentioned before, the more the numbers on the bus stop, the better connected the place is with public transport. Likewise, the more frequent occurrence of the bus stop shelters in a row, the more central and significant in the city the represented urban place seems to be.

Redundancy of messages – through two channels: urban elements and users – increases the possibility for the transmission of meaning.
The omission of certain elements in the narration or more specifically, showing the omission of certain elements through cause–effect results in space is also an important instrument of storytelling (see Figure 7).

Multiple ways of achieving the same message by combining the urban elements with users or several users and accessories.
The redundancy of alphabet elements, words and sentences – meaning the same words or sentences can explain the same issues, notions and phenomena – ensures that the message does not get lost even though the ‘reader’ misses some clues (see Figure 8).

Individual sentences form groups or ‘paragraphs’ that address the higher level of understanding of the urban space (e.g. access, ‘life between buildings’, centrality, services, etc.).
The synonyms act in a similar way, reinforcing the message, sometimes through the objects themselves and sometimes through their use (see Figure 9).

Visual synonyms share a meaning, but come in different forms that also imply further hints of intended use.
The objects that could carry several meanings and cannot be fully understood alone are defined by their context, by combining them with other elements or by showing them in use. The multifunctional urban element, basically a wall of different heights in Figure 10, serves different purposes: as a point of reference in space and a meeting point, or, in some other setting as a barrier and division. Isolated, it does not convey much meaning except that it is some kind of wall.

Objects that could carry several meanings and cannot be understood alone are defined by their context, by combining them with other elements or by showing them in use.
As in language, some links between the signifier and signified are stronger and more direct than others on a higher level of the message syntax: the connection between a proposed alphabet and words can be understood more clearly and directly than between words and sentences. The story that results, though complex in meaning, can also be prone to different interpretations, which share meaning to a large extent but do not translate directly or in one exclusive way to all readers (observers). The communication noise increases with the complexity of the syntax.
4. Advantages of the Proposed Urban Visual Vocabulary
The attraction of the proposed vocabulary is in its openness at the lower end, at the selection of alphabetical elements and their alternatives, their numerous combinations and composition – in phrasing to use the language metaphor again – as well as openness at the highest levels of urban stories told in as much complexity as intended.
The vocabulary and syntax rely on the observer’s already established and culturally inherited understanding of the relation between an actual object and its representation – between signified and signifier. When Becker (1974: 6) discusses the conventions in photography, he emphasizes ‘we are not ordinarily aware of the grammar and syntax of these conventions, though we use them, just as we may not know the grammar and syntax of our verbal language though we speak and understand it.’ By stepping away from coded symbols used in expert fields and relying on the tacit knowledge and experience of the observer, we are closing the expertise gap that usually hinders expert–non-expert discussions.
Limiting the syntax, fixing the number of elements and defining the connections too strongly would be counterproductive when trying to transmit tacit meaning and appealing to tacit knowledge of the observer. Instead we rely on experts’ and non-experts’ intuition. Intentionally chosen elements by the composer (an expert) are chosen on the basis of his or her experience, knowledge and expertise. The objective intuition, however paradoxically sounding, used to choose among possibilities to prepare the narration, has firmer ground in his or her expertise and needs only to be brought to daylight in a form that addresses tacit knowledge of a non-expert. The observer’s non-expert intuition originates in his or her tacit knowledge and experience. This fragile consensus between the one who prepares the narration and the one who receives it constitutes their mutual frame of reference. The narrator considers the narration as suitable and the representation as adequate. He or she is also anticipating the span of an observer’s tacit knowledge. The observer (the reader), on the other hand, accepts the narration based upon his or her actual tacit knowledge and interprets it within those limits. Neither side is entirely confident about the other’s abilities nor seeks any proof they exist; however, both are ready to ignore those uncertainties as long as the messages seem to be understood. The syntax aspires to clear messages and clear connections between elements and their meaning, but, because of this consensus, functions also when connections are looser and cannot be guaranteed due to abstractness or complexity of messages. The further appeal of a proposed vocabulary lies in the possibilities to steer the narration – whether showing the qualities of the place, problematic aspects, cause–effect relationships – through the aforementioned intentional choice of elements. They appear and are shown only if they carry a certain meaning and are meaningfully contributing to the narration (i.e. never for decoration purposes only).
Bringing all the advantages together, we could argue that the main attraction of the proposed vocabulary lies in its simplicity, openness, fluidity, reliance on intuition (expert and non-expert) and short (or almost no) learning curve for all parties involved.
The fluidity of the alphabet, vocabulary, connections and the emergent complexity of the narration are not random occurrences but rather stem from the order and structure based upon simple rules, characteristics of elementary units, which will be explained in the following section. We also point toward the higher hierarchical order of categories or notions which we aspire to or want to relate to at some point in our urban stories.
5. The Urban Visual Alphabet: Fundamental Elements and Rules
Given the abundant, complex conglomerate of an urban scene, what elements are we actually able to discern, track and perceive? Furthermore, what elements are useful when conveying meaning to non-experts, less skilled in visual analyses and more prone to intuitive reading of visual material?
The resulting alphabet is not a set of standard letters but rather a fluid selection of basic units satisfying the following characteristics:
The element is visually traceable in the observed environment;
The element as a carrier of meaning can be represented as perceived – in an experiential way – the relationship between signified and signifier can be deducted from the representation of the element itself, or is inherited from tacit knowledge and previous experience (i.e. we could also use symbols for representation but they would require a mind leap, which we strive to eliminate at the alphabet level);
The element is considered a basic unit when it can be perceived in space, and is recognized as the carrier of the smallest possible meaning in this space (to be deemed basic, the element must carry a meaning, but as few meanings as possible, otherwise it is considered composed) – in other words, the basic unit is the smallest elementary entirety that we can define in space.
It can be characterized as fluid because it is up to the storyteller to select the basic units according to the rules: how the smallest possible meaning in space is set and defining the number of units. The system, however, is not as open as it sounds, given that we could, with some effort, reach a consensus regarding the representation of the smallest possible meaning in space. However, it is open in such a way that it does not provide a fixed set of characters, only their definition, which allows for some leeway according to the ingeniousness and creativity of the storyteller.
After the selection of the basic elements that could form the representation of an urban scene comes the important moment of eliminating all the elements that will not contribute to the narration, with the aforementioned purpose of avoiding clutter, misunderstandings and distractions – in short: to avoid the noise. The remaining elements serve the purpose of explaining or reinforcing a significant piece of the story and are meaningful (essential) to the storyteller as well as the reader.
Here we finally hit two inevitable questions: what seems to be essential for the story and what is the ultimate urban story? The answer to the first one is inherent in our idea: the decision lies with the storyteller based on his or her expert intuition and the envisioned aim. The answer to the second question is discussed below.
6. The Elusive Ultimate Urban Story
We have already mentioned the general aims the urban story wants to achieve, namely to explain how a specific place works, what it offers, how it is organized, and the cause and effect relationships of urban elements. It also represents the users: who they are and what they do. Finally, the story shows how the place is meant to be used, either presenting the above mentioned elements one by one or all at once. We have shown by which means and instruments – through a visual urban vocabulary – we strive to attain those aims.
The ultimate urban narrations, as well as the ultimate understanding of such a story, of course, do not exist. The vocabulary and diction proposed allow for limitless narrations varying in the presentation of complexities of urban spaces and multiplicity of their phenomena, limited only by the skills of the presenter and comprehension skills of the observer. Instead of further discussing what stories the visual urban vocabulary might tell us, our thoughts can be better followed visually – starting with Figure 11, following the caption instruction, and then proceeding to Figures 12, 13 and 14.

In this ‘textbook’ case scenario, the observer is encouraged to analyse the place in the image: what does it offer, where is it, how does it function, who are the users, is it a representation of a real or a generic place?

The urban story starting at the level of letters and words of the visual urban vocabulary. The sentences, paragraphs and the urban story they amount to, are shown in Figure 13.

The urban story (in grey) as intended with sentences and paragraphs explained. There are several levels the observer can achieve, depending on time, thoroughness, previous experience and expertise one is prepared to invest.

The depicted place is close to the image of its actual twin but is in no way its exact replica or a representation. The plausibility and the similar feel – of an actual and generated place – their stories radiate is what we aim for.
7. Discussion: The Reasons, The Use and Further Research
7.1 Why do we want to tell a visual urban story?
There are multiple impetuses for the story of urban places to be told.
For the purposes of general public education – experiential but, to a large extent, generic presentation of research, studies, statistics, urban entities and phenomena, cause and effect relationships, present states and future trends;
For developing, planning and implementation purposes, especially when non-experts are involved – intuitive representation of different scenarios that can be faster and more easily perceived by non-experts but still possess depth; different comparisons and evaluations of given urban places, before–after effects, etc.;
For commercial and social purposes – use of these principles in different professional tools for creating non-expert friendly presentations; in public relations and public participation processes, where visualizations address target publics and groups with the visual language that is intuitive, likeable and informative but also in which they can recognize themselves planning.
With each impetus emerges also the question of the audience – to whom it will be shown, specifically whether it is going to be presented to an expert or non-expert public. Depending on the audience and based on our research, the stories and presentations can be tuned to best resonate with the viewers.
Research shows that the non-expert can be approached through narration using elements that are not only visually recognizable to the observer, but also indicate benefits or detriments for the user of the place (Verovsek et al., 2013a). This means observing the place through the prism of expectations that need to be satisfied in such places and results in different perceptions and reactions of the user (presence/absence of the sense of orientation, sense of comfort, safety, etc.). The binding layer here are the basic human needs (Bradshaw, 1972; Maslow, 1943) translated into the context of open public spaces, as a leverage to guide one’s perceptions in a particular place and further one’s reactions in terms of impressions, sensations and behaviour. Thus, the user understands a place and its qualities from his or her own experiences: its level of comfort and attractiveness are as perceived by the potential of the space to satisfy and fulfil one’s expectations. The urban qualities, or lack of them, therefore define the quality of living in the broadest sense, as they represent the stage and scenery for human activities and are thus an important factor in the decision-making process, while also being a crucial measure in terms of representation for the general public.
The flexibility of the vocabulary is not limited to the perception of urban spaces and their intra-relationships through anticipated use but can also represent notions of the higher order and abstraction, explaining the design or composition, given that the reader is familiar with the topic or expert in the field. Here again Cullen’s (1971) principles and notions come to mind (i.e. possession, occupied territory, deflection, closed vista, intimacy, etc.).
It is also imaginable that an expert prepares the narration that will explain the space to non-experts through the means of its use and will also use the same presentation for experts to demonstrate not only a narration of its use but also of composition and design.
7.2 What are the further uses of visual urban vocabulary and why has it been developed?
The vocabulary has emerged from the idea of developing a descriptive computational model of urban spaces. This consists of a symbiotic pair of both the descriptive model and the visualization model that would be capable of recreating, generating and representing the output data in a manner understandable to experts and non-experts alike, in an experiential, generic, but also likeable way. Initially the model was set up as an instrument to demonstrate and interpret the distinctive spatial features that contribute significantly to the quality of the built environments. It is meant as a mechanism for interpretation of qualities in urban space (A model for Interpretation of Qualities in Urban Space – aMIQUS) to assist the urban decision-making process. The model is composed of the input elements (attributes) that visually describe the location, and the outputs that indicate its active use (users, their activities, motives, demographic structure, etc.). The binding layer consists of generalized human needs and expectations that are to be satisfied in such places, and these needs and expectations result in different reactions to urban spaces. As previously described, the idea is to create an insight into the space structure from the perspective of users’ daily experiences – this is how a specific place accommodates users’ needs – and furthers the question of how these can be linked back to the theoretical, more abstractly defined metrics and descriptions (used by experts) of this same place (and its attributes). We want to link the tangible urban elements to these notions as well as indicate the active use of the specific location through them (i.e. accessibility by different travel modes including, pedestrians, life between buildings [see Gehl, 1987], demographic structure, social structure, activities, motives, a programme of the buildings). With these means, we accentuate and support the user perspective, which in many ways we deem essential: we literally use the experiential, first-person perspective and refer to it indirectly by reflecting the user’s tacit knowledge and previous experience of similar spaces. Additionally, the emotional effect of the volume of space as a factor in a user perspective is taken into account to recreate the place from the first person point of view (McMurtrie, 2012; Stenglin, 2004).
When modelling urban places (in terms of demonstrating the cause and effect system, not modelling as in shaping or 3D modelling), we distinguish between the initial visual elements designating the given space (geometry, geomorphological absolutes, a functional division of space, infrastructure, furniture, the function and programme of buildings, centrality of the location, geo-position, etc.) and the external factors (geo-spatial, temporal and societal context). The output elements consist of calculated user simulants and spatial activities but also of modifications (addition or reduction) of some initial urban elements. This grouping and division is not universal and could have been different, but it is proposed for feasible communication and computation between the computational, descriptive model and generative visualization model that each has its own needs and constrictions. There are subgroups of elements divided furthermore in greater details (Verovsek et al., 2013b), and this is where visual urban vocabulary emerged.
Based on the research and testing of the model, three digital tools have been proposed that would incorporate both parts of the model based on visual urban vocabulary. They would allow experts, even those with very limited drawing, sketching, graphic design, photographing skills, or none at all, to create visual urban stories for multiple purposes.
The most promising tool (Vili) revolves around mimicking photographs of urban scenes through substitution of the observed phenomena with spatial elements (silhouettes) from the digital library. An algorithm-based engine is used to place the visual forms appropriately and in scale. Each element placed on the photographic underlay comes with the specific attributes, affects the output values and generates visual representation in terms of the qualitative and quantitative trait. Later on, the modifications of the input attributes made by the story creator are enabled and accompanied concurrently by emerging changes in the experiential view. In this way, any adjustment made in reference to the attributes of the input variables is expressed in the final presentation at the experiential level, which ultimately demonstrates the cause–effect narrative and therefore provides an interpretation of the relations within a certain place.
The tools could be used on their own or as suggested by Jutraz et al. (2011) in a comprehensive system of tools and procedures for the participation and inclusion of non-experts into the formal and informal urban design and planning processes.
Indicating where the visual urban vocabulary stems from and how it emerged does not, however, limit its usefulness to the urban design field alone. Given that we, as professionals, engage daily in discussions, debates and the co-shaping of our urban environments, the vocabulary scope goes beyond our work and research, and can be used or at least thought of whenever the need for communication among different actors in space takes place.
7.3 Whereto from here?
By expanding and further developing the initial idea of visual urban vocabulary in our future work, several promising research paths emerge. Most logically followed is the testing of the vocabulary in actual use by means of empirical inquiries. This can be accomplished independently of the particular urban places and their issues. The inquiry could also be to a greater or lesser extent embedded in the actual and ongoing processes of education, municipality urban decision-making, civil initiatives work, advertising, etc. Bringing inquiry closer to the definite, mundane circumstances or merely dragging it slightly from the hypothetical research fields could bring deeper and more realistic insight into people’s responses to the vocabulary. However, greater methodological creativity is needed to establish the approaches of testing to avoid noise effect factors, such as declarative answers, strong identification of the respondents with the given location and its (perennial) issues, existing deeply rooted opinions or attitudes, impact of preliminary knowledge and skills of respondents, etc.
Another point of interest is the inclusion of the spatial identity into generically created representations. Although the generic contextual identity of the places has been included in our vocabulary and images, the exact, particular, exclusive and place-given elements have clashed with the generic vocabulary idea. Nevertheless, we are aware that in any serious debate about places and their qualities, evaluation and development must include their particularities that have to be either hand-placed in images or recreated in some other way.
Pointing out the advantages of the vocabulary earlier, we are aware of some limitations to the static visual formulation of urban stories. The two most perennial are the limited density of elements, due to the format and legibility constrictions and the occlusion of objects in the background by those placed in front of them. Both can be solved by sequencing the story through several images or storyboards, but additional value can be achieved by the dynamic and interactive nature of such stories providing additional and in-depth information, which at this point has not yet been researched.
8. Bringing Our Urban Story to a Conclusion
We have begun by building our argument on the photograph that could be generic in nature and have shown why its generic nature of shown places, elements and users would be beneficial for urban storytelling and explanations of complex urban environments. The place in Figure 11 is based on a photograph of the actual place (Wolfova Street in Ljubljana). Everything else in the image, beyond the geometry of the street and the buildings: urban furniture, trees, shops, bins, lamps, people, is generic and placed into the picture with the sole purpose of telling the story. The depicted place is similar to the image of its actual twin but is in no way its exact replica or its accurate representation. The image is computer generated and is thus, in part, generic – created and visualized in a generic way with the use of the proposed visual urban vocabulary. We aim for plausibility and for a similar impression from the generated narrative, giving it the credibility and homeliness the reader can relate to. On the other hand, the clear advantage of such a vocabulary and its generic nature is that they provide limitless possibilities for telling infinite stories either based on hypothetical or actual scenarios regardless of the actual urban reality. The story can thus be simulated based on the real inputs, tweaked, focused or exaggerated, depending on the aims and purposes of the story-teller and the intended audience. The additional benefit of a visual urban vocabulary’s generic nature is its ability to represent different fictive or actual places with the same language, equalizing the ‘wow’ effect of some compared to others, reducing visual clutter, the effect of amateur compared to professional photography, recognition and identity ballast, etc. and thus making it possible to evaluate them and study the relationships in them on more equal (visual) grounds.
A brilliant story can be told in any number of languages and there are many who have mastered those languages. Yet only some of them are also award-winning writers, proving that the highest level of storytelling is a combination of a careful selection of words, an art of composition and good command of the alphabet, vocabulary and syntax, with a fair amount of creativity, imagination and talent. However, often not just brilliant but a decent, clear, intuitive and intelligible visual urban story would suffice to engage the actors in a discussion about space, motivate their participation in design and decision making, explain cause and effect relationships in urban environments or demonstrate to pupils the complexities of urbanity. This is where the visual urban vocabulary and the proposed tools find their practicability.
Footnotes
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors and there is no conflict of interest.
Biographical Notes
MATEVZ JUVANCIC primarily focuses on architectural education of general public and public participation, specifically dealing with new means of transmitting and communicating space- related issues. His research time is distributed between visualizations and presentations of urban spaces as well as educational architectural tools for general public, their design, use and evaluation. Using these means, he sees the opportunity to communicate both natural as well as cultural aspects of sustainable spatial development.
Address: University of Ljubljana, Zoisova c. 12, Ljubljana, 1000, Slovenia. [email:
SPELA VEROVSEK is currently employed as a researcher in the Faculty of Architecture, University of Ljubljana, where her work is primarily focused on the research and development of visualization techniques in urban design intended for the non-expert public. Her research aims at developing novel approaches to understanding the complex information and logics of urban spaces, also to interpreting different spatial features and qualities through user perception.
Address: as Matevz Juvancic. [email:
