Abstract
In contrast to the widespread approach of the notion of “virtuosity” highlighting mechanical dexterity, I advance the view that a crucial feature of both technical virtuosity and musical expressivity is a specific ability to readily and quickly adapt the attention to various constantly changing aspects of the musical process while performing. In this sense, virtuosity can be likened to general mental capabilities such as intelligence. Based on concepts taken from performance analysis, pedagogical practice and sports science (especially from recent research of attentional control in sports), I attempt to define and discuss key features of performance virtuosity as “mental dexterity”. This includes the ability to quickly position oneself into different temporal perspectives in real time during performance; the ability to quickly shift the attentional focus, as well as to quickly modulate the depth of attention; furthermore, to promptly position oneself into different empathic perspectives, similarly to projecting oneself into another person’s position.
Keywords
Popular tutorials such as School of the Virtuoso (Czerny), Technische Virtuosenstudien (Köhler), The Virtuoso Pianist (Hanon), The School of Virtuosity (Thomán), Technique moderne du pianiste virtuose (Bosquet), The Violin Virtuoso (Bachmann), alongside countless other titles from the 19th and early 20th centuries, either explicitly or implicitly expressed the view that the development of virtuoso skills is founded on technical (e.g., finger) dexterity. Even today, common understanding of the notion of virtuosity among musicians tends to centre on the physical and mechanical, not unusually conjuring up the picture of “an elegant and rather diaphanous man with agile fingers and an empty head” (Berio, 1985, p. 90, quoted by Ginsborg, 2018, this issue). It emphasizes instrumental technique, rapidity and agility related to motor skills as it is often evinced in scholarly reflections on virtuosity. On the other hand, many insightful analyses (such as that of the anthropologist Royce, 2004, see esp. pp. 18–19) and a significant number of musicians’ own definitions, as revealed by Jane Ginsborg’s research in the present issue, highlight that mere physical showing-off does not give rise to “true” virtuosity: according to these, virtuosity closely linked with, or indeed defined by, mechanical dexterity is considered as a means to an exceptionally skilful and convincing expression of musical content. 1 In my paper I argue that the nature of virtuosity, defined as a means to both this end and technical flawlessness, deserves more attention from the point of view of cognitive psychology than has hitherto been the case.
The fascination with “technique”, as opposed to “expression”, has a long history in music, especially in pedagogy. Most notably, for the major part of the 19th century, music pedagogy was predominantly concerned with technical training on instruments, leaving aside the teaching of musical “poesy” (the nurturing of musicality and inspiration), commonly deemed to be impossible to teach (Stachó, 2018b). As the institutionalization and social dissemination of music pedagogy spread, extensive instrumental technical training became the norm: note the burgeoning of etudes and finger exercises that formed a major part of teaching materials from the mid-19th century onwards. Later, towards the end of the 19th century, music education turned its focus to the technique of expression as well, and particularly on the rules of expressive performance determined through descriptive analyses of expressive performances (cf. the many treatises on musical expression written in the era, cited below). These approaches to pedagogy typically prioritized instrumental technique over musical content, or focused on the musical output (i.e., how the end result of the learning process should sound) rather than on the mental processes involved in producing the resulting sound. It seems that the technical approach dominated both instrumental pedagogy and solfège, right up to the highest levels of musical expertise, and that the domain of the alternative, “cognitive” approach was as a rule left to the discretion of the performer and to the act of performance itself.
In a strikingly similar way, empirical research on expressivity in music performance tended for a very long time to focus on the analysis of expressive end-products, producing structural-level descriptions (for an overview of empirical research into musical expressivity see Gabrielsson & Lindström, 2010), leaving aside the phenomenological aspect (including generative imagery, as well as issues in cognitive and attentional control) in the production of performance expressivity. Thus despite the relative longevity of this research period, which basically started in the 1920s (for an overview of early research see especially Gabrielsson, 1999), most studies in this area so far have focused on a structural-level analysis of performance expression, as both theoretical knowledge and empirical data on musicians’ cognitive and attentional processes and strategies in the act of performance have been lacking.
Performance expressivity: A cognitive view
It is a widely agreed fact in music analysis and psychology, as well as in performance practice, that expressivity in music is shaped by two fundamental feature groups: (a) compositional features (i.e., tonal, metrical, and grouping structures [or, depending on the analyst’s point of view, processes]), and (b) performance cues which build up the performer’s interpretation and realization of a musical composition, such as overall tempo variability, 2 patterns of microtiming (small-scale temporal variability), volume, articulation, tone attacks, intonation, vibrato, and timbre (see, e.g., Gabrielsson & Lindström, 2010; Jackendoff & Lerdahl, 2006; Juslin & Lindström, 2016). 3 We must note, however, that although the two fundamental feature groups can be separated on a theoretical basis, listeners almost never actually experience “composed” structure independently from performers’ expressive choices (cf. Cook, 2003).
Performance expression is usually described through generative rules of timing and dynamics that produce the timing and dynamic patterns of an expressive performance, resulting in a structural-level description. One of the most straightforward examples of this kind of description, based on both expert intuition and informed (expert) deduction of performance principles from recorded performances, is the model of performance rules involved in phrasing, microtiming, intonation, the execution of metrical patterns and grooves (including metrical styles in Baroque or folk dances, among others), the expression of tonal tension, and ensemble timing, formulated by a Stockholm-based research group (Friberg, Bresin, & Sundberg, 2006). This line of inquiry (characterized by Langner, Kopiez, & Feiten, 1998 as the “synthetic” approach) has a long tradition, rooted in pedagogy: the reader should recall that whereas instrumental education was predominantly concerned with technical training on instruments during the greater part of the Romantic era in music history, towards the end of the 19th century it shifted its focus to the technique of expression as well, and particularly to the rules of expressive performance determined through descriptive analysis of expressive renditions. Intriguing systematic descriptions of performance expressivity, then thought to be exhaustive or nearly comprehensive, were produced by several German, French, British and American authors most of whom faded into obscurity by the second half of the 20th century (e.g., Christiani, 1885; Fuchs, 1885; Goodrich, 1899; Klauwell, 1890; Lussy, 1884; Riemann, 1903). On the other hand, generative rules of expressive performance can be extracted from large corpora of real, audio-recorded performances automatically, using computer-assisted analyses and including meta-learning algorithms, as in Gerhard Widmer’s research (Widmer, 2002, 2003; for an overview see Widmer & Goebl, 2004). 4
At the cognitive level of the description, expressivity is shaped by the actual feelings and thoughts of the performer, together with the performer’s direction of attention, which are at work behind the generative rules of timing patterns and other performance cues. Influential theories in music aesthetics and psychology acknowledge the fact that cognitive and affective components are highly intertwined during processing of musical materials as musical structures are “impregnated” with emotions (for three converging approaches from different domains see: Meyer, 1956 and Huron, 2006 [aesthetics and psychology of music]; Dahlhaus & Eggebrecht, 1985 [history of music aesthetics]; and Dobszay, 2012 [music analysis]). In effect, every piece of structural information in music is subjectively linked to feelings: on one hand, musicians express and communicate structures by feeling the elements of structural processes (e.g., performers are able to predict and feel the length of a musical unit to be performed, or subjectively link feelings to components of the tonal structure of a composition such as chords or chord progressions); on the other hand, emotional expression is highly structured through musical composition. Thus feelings and emotions are not merely “performed” or simulated but felt during performance, enabling musicians to achieve appropriate performance cue variability to successfully create expression.
In sum, in order to produce expressivity musicians during performance typically do not focus on performance cues as such – which serve ultimately as means to produce a desired expressive effect – but on the generative sources thereof, that is, on emotions and feelings associated with the music by virtue of “understanding” it (meaning the performer’s conscious or unconscious cognitive representations and structures constructed in relation to the music in question; see Jackendoff & Lerdahl, 2006, which draws on Raffman, 1993 and Davies, 1994).
To illustrate the performance phenomena I intend to analyze in the reminder of the paper, I invite the reader to watch Video example 1 in the Supplemental Material section. Here pianist Maria João Pires makes use of the metaphor of space to characterize intensive mental imagery related to the feeling of the musical process. In the following sections, I provide a preliminary formulation of a new theory of music performers’ cognitive–attentional strategies that is hypothesized to produce the mental imagery enabling performers to feel the musical process.
Performance expression and mental imagery
To date, performers’ feelings and thoughts during musically expressive performances have been investigated in depth only in very few studies, despite the fact that the expressive quality of a musical execution heavily relies on the performer’s cognitive-level activity, including the actual feelings and thoughts of the musician and their direction of attention. Some of the most interesting empirical findings on performers’ feelings and thoughts during playing have been obtained by Van Zijl and Sloboda (2011), among others (such as Clark, Lisboa, & Williamon, 2014), through musicians’ self-reports and interviews made with them. However, a unified theory and detailed empirical investigations on musicians’ mental processing and strategies in the act of performance have so far been lacking (cf. also Persson, 2001). Especially noticeable by its absence is theoretical and experimental research on attentional processing, which could stimulate further research and pedagogical applications.
On the other hand, research on musical imagery in relation to performance has usually focused on auditory imagery (e.g., Hubbard, 2013; Repp, 2001), the role of executive functions such as working memory (for the role of working memory in anticipatory timing see Colley, Keller, & Harpen, 2017), the neuropsychology of the imagery process (e.g., Meister et al., 2004; for a related overview see Zatorre & Salimpoor, 2013), and memorization (Bernardi et al., 2012; Holmes, 2005); but to date, regarding mental imagery there has definitely been more empirical research in the domain of sports (cf. Clark, Williamon, & Aksentijevic, 2011) than in music. Another, though still sparse, pedagogically motivated line of inquiry related imagery during performance to expressivity, enjoyment, anxiety, and success (Clark, Lisboa, & Williamon, 2014). Applied sport psychology has inspired studies and psychological methodologies to enhance performance expressivity and success through general psychological variables such as self-efficacy (McPherson & McCormick, 2006; cf. also Green & Gallwey, 1986), growth mindset (O’Neill, 2011), self-talk (Clark Lisboa, & Williamon, 2014; Weiss, 2008), goal setting, and non-judgmental awareness (Green & Gallwey, 1986). However, except for a very few notable but in the context of music performance, rarely cited theories in the phenomenology of time (Husserl, 1991; see also Clarke, 2011) and music analysis (Dobszay, 2012), no inquiry to date has tackled the issue of mental strategies and imagery, specifically related to the temporally unfolding musical process, that lies at the bottom of performance expressivity.
Performance expressivity and the real-time representation of musical meaning
Due to the dearth of empirical research into the phenomenology of the performer related to performance expressivity, we have to resort to research into linguistics, the psychology of acting, as well as everyday – anecdotal – evidence in music performance and pedagogy to advance the theory that in the act of performance the musician, in order to successfully produce expressivity, actively represents the subjective meaning of the music (s)he performs (Stachó, 2016; Stachó & Holics, 2011). In fact, established theory in acting (such as Stanislavski, 1937/1980), both everyday and scientific evidence in music performance (cf. one of the very first works in music psychology to tackle the topic: Seashore, 1938; for a short contemporary overview see Keller, 2012; for a more recent, empirically founded approach cf. Globerson & Nelken, 2013) and pedagogy (one of the best-known textbooks embracing this stance: Green & Gallwey, 1986) suggest that successfully executed performance expression relies on vivid imagery related to the meaning of the text/music. Also, there is empirical evidence that performance expressivity is negatively correlated with thoughts and feelings unrelated to musical meaning (e.g., Clark Lisboa, & Williamon, 2014) and that professional musicians employ imagery to limit distractions in the act of performance (Gregg, Clark, & Hall, 2008). However, thorough research on the nature of this imagery is still lacking (see Clark Williamon, & Aksentijevic, 2011, for a review that remains one of the most up-to-date on the topic).
Although the notion of meaning in music is far from having an established understanding, in both everyday and scientific discourse it is a commonplace to regard music as a language or a language-like system (e.g., Kraut, 1992; Lerdahl & Jackendoff, 1983; Raffman, 1993; for a noteworthy early review see Feld & Fox, 1994). However, whereas music cannot express or communicate propositional content, it is commonly considered to be an ultimate means to convey affect, comprising emotions and feelings. Also, while we may argue that affect is the “meaning” of music, it is so in a different manner from how meaning is conceived in language, since a piece of music does not denote a certain emotion in the same way a verbal utterance denotes its meaning. In music, the process of understanding works rather differently to that in everyday language; indeed, it is more comparable to poetic communication where (a) the process of understanding is “framed” (Jackendoff & Lerdahl, 2006; Juslin, 2013b), opening the way to aesthetic experience, (b) and on the other hand, formal properties such as structure (e.g., metrical structure both in language and music), and fractions thereof, become ultimate sources of meaning (regarding literary texts see Pilkington, 2000).
Previously, I proposed to define music as signalization, in a broad sense of the term referring to the ability to formulate subjective associations in a more or less systematic way to musical stimuli (see its most recent formulation in Stachó, 2018a, in press). This concept of the musical sign has a broader scope than the classical Peirceian definition of sign: in order to capture the essence of how people understand music, instead of suggesting that a sign is something which stands to “somebody for something in some respect or capacity” (Peirce 1931–1935, Vol. 2, §228), I proposed to define a musical sign as a sonorous object which for somebody is associated with something in some respect or capacity in order to achieve an understanding of it, that is, constructing cognitive representations and structures in relation to that object (for a related definition of “understanding” see Jackendoff & Lerdahl, 2006) and relating, or integrating, them (in)to prior knowledge. Note that this definition emphasizes the subjective nature of musical meaning, proposing that it is created in the listener’s/performer’s mind rather than placed in fact or merely based on structure. Relying on aesthetic theory, empirical research into music perception, and pedagogical practice, the following sources of meaning (or “content”) in music may be identified, starting from basic ones, available from early infancy without necessitating a vast array of musical experience, and proceeding towards more complex sources requiring ample exposure to musical stimuli.
Movement patterns (“musical gestures”)
A basic source of musical meaning is the physical dynamism of music. Following the theory of vitality affects outlined by Daniel Stern in the 1980s, the physical dynamism of music yields vitality affects and physiological reactions. Available from early infancy and originating from the dynamic cross-modal attunement in mother–infant interaction, vitality affects can be typically described by dynamic, kinaesthetic terms, such as “surging”, “fading away”, “fleeting”, “bursting”, “drawn out”, and so on (Stern, 1985). Postulated as precursors to the later appearing emotions, these qualities of experience are thought to be most certainly sensible to infants. Further to physical dynamism, several additional features of music, such as pitch contour, are hypothesized to be connected with physical patterns of posture and gesture, conveying affect (Jackendoff & Lerdahl, 2006). In pedagogical practice, these motion patterns, relying on the physical dynamism of the musical flow, are often referred to as “musical gestures” (see, e.g., Gritten & King, 2006, and more recently, through the introduction of the notion of “shape”, Leech-Wilkinson & Prior, 2018). This layer of meaning relates to performance features shaped by the so-called motion principles in Juslin’s (2003) influential theory of performance expressivity, the “GERMS” model.
Immediately expressed emotions (“musical characters”)
Alongside gestures experienced during listening to music or performing, the direct expression of more static affective states (Juslin, 2000, 2013a; Juslin & Timmers, 2010) constitutes a further layer of musical understanding. In pedagogical practice, immediately expressed emotions are usually referred to as “characters” – such as nobility, gloom, fear, pain, and countless other emotions and emotional states. These may be expressed by means of musical gestures; however, they are often not based on them but occur relatively independently from gestures. In performance, this layer of musical meaning is related to the expression of emotions independent from the expression of structure (Juslin, 2003).
Tonal and temporal structure
Emotions, hence meanings, related to the perception of musical structure typically emerge from both active and passive musical experience: they may result from the fulfilment or unfulfilment of momentary expectations about the continuation of music. These experiences are guided by learned rules about musical styles. In his path-breaking book on emotion and meaning in music, Meyer (1956) provided one of the first comprehensive and widely known theoretical accounts of emotions resulting from expectations during music listening (for a more recent integrative account cf. Huron, 2006). First, expectations are linked to tonal structure, the hierarchical framework of the pitch and harmonic content of music which unfolds in real time. Second, expectations follow metrical structure, the hierarchical temporal framework of beats that organizes the musical flow into regularly recurring bars of stressed and unstressed units of pulse (i.e., beats) which in turn are hierarchically organized into larger units. A crucial source of musical meaning for both listeners and performers is thus the metrical process which unfolds in real time, eliciting not only expectations related to length correspondences of larger units, but also well-definable feelings of various temporal lengths. Finally, grouping structure fills out the metrical structure with thematic material, yielding the segmentation of the musical flow into motives, phrases, and larger sections; as grouping structure unfolds over time, it evokes thematic, rhythmic and length correspondences and expectations. In performance, tonal, metrical and grouping structures shape Juslin’s (2003) structural expressivity.
The narrative–dramatic structure
The perception of narrative–dramatic structure – or, as it is often called by narrative theorists, the “affective curve” (e.g., Pasler, 2008) – of a musical process relies on the empathic projection of feelings anthropomorphically onto dynamic processes like music (Walton, 1990) and their ordering according to a narrative–dramatic plan (Levinson, 2004). In performance, both emotional expression and motion principles contribute to the expression and communication of the narrative–dramatic structure (the concatenation of gestures and characters according to a narrative plan).
The psychodynamic layer of musical meaning
A further, psychodynamic layer of musical meaning, accounting for probably the predominant part of our musical experiences, is shaped by individual associations and recollections of thoughts and feelings, yielding emotionally coloured meanings. As these private musical meanings are developed through a person’s own occasional experiences related to actual circumstances of music listening, they fall beyond the general frame of aesthetic experience (Jackendoff & Lerdahl, 2006; cf. also Juslin, 2013b) and are typically beyond the performer’s control. 5
Imagery and consciousness: Dynamic processes
Mental imagery supporting the perceived expressivity, intelligibility and individuality of a performance relies on the performer’s real-time mental representation of musical meaning during playing, that is, their own understanding of gestures, direct emotional expression, narrative and drama, and the tonal and temporal structural processes. In particular, cognitively representing the tonal and the temporal structure in the act of performance facilitates the appropriate use of performance features, tailored to the interpretive process of a presumed listener. 6 Based on these considerations, I embrace the view that the performer’s mental representation of the musical process which unfolds in time, together with their attentional process and strategy related to the expression of musical meaning, define performance features – especially related (but not limited) to timing – and determine the expressive quality of a performance.
Relying on concepts taken from pedagogical practice and theory (Stachó & Holics, 2011), sports science (including recent research into attentional control in sports, e.g. Jackson & Mogan, 2007; Savelsbergh et al., 2002; Williams et al., 2011), as well as from analyses of video-recorded performances, I recently proposed the outline of a novel model of musicians’ mental processes and attentional strategies during performance (Stachó & Holics, 2011; Stachó, 2016). My approach accommodates the hypothesis that these strategies and processes underlying performance expressivity involve expressing and empathizing feelings in real time, and this activity is connected to an intense mental imagery process. Typically, this imagery builds on moments of deep immersion and involves a specific kind of attentional processing. Although my model of mental/attentional processing and strategies underlying performance expressivity was developed independently from related theoretical models in general psychology (such as Zimbardo and Boyd’s time perspective theory: Zimbardo & Boyd, 1999), philosophy (such as Husserl’s theory “on the phenomenology of consciousness of internal time”: Husserl, 1991; see also Clarke, 2011) and music analysis (Dobszay, 2012), it strongly resonates with these approaches.
The theory of moments of focused immersion (MFIs)
It is a well-known phenomenon that in order to achieve expressivity musicians induce altered states of consciousness whose characteristics closely follow established definitions, acknowledging that they demonstrate “a sufficient deviation in subjective experience or psychological functioning from certain general norms for that individual during alert, waking consciousness” and are represented by “a greater preoccupation than usual with internal sensations or mental processes, changes in formal characteristics of thought” (Ludwig, 1969, pp. 9, 10; in relation to music performance this definition is also embraced by Persson, 1993, 2001). So far the nature of the consciousness associated with music-making and attending to music has been open to investigation but we can easily agree with theorists such as Zbikowski (2011, p. 190) in claiming that it mostly reflects memory systems which are “for the most part much more focused on the salient features of dynamic processes than on lexical knowledge or relationships between objects and events”. In line with this claim, I advance the hypothesis that in the act of performance, moments of deep attentional immersion, which are connected to the imagery process and in which the musician generates quick and transient, very brief shifts of consciousness, embody the three time perspectives – the past, the present, and the future – and have multiple functions related to psychology and music theory. Henceforth, I shall refer to these moments of intensive mental imagery which are correlated with an in-depth attentional immersion as MFIs (“moments of focused immersion”).
Present-focus: Enjoyment
A crucial function of MFIs is to allow the musician – and, possibly through direct empathy, the musician’s audience – to achieve a highly focused, mindful perception of the present sounding moment, without breaking the performance process. In their absence both a performance and a listening experience tend to be perceived by listeners as superficial and feebly expressive, and the performer may not be able to capture the listener’s attention. This kind of momentary immersion usually appears to last for less than a second, and it is likely to have specific functions related to music theory such as marking tonally important moments. This function is seminal to achieve meaningful quickness – as in a virtuoso performance which is felt by listeners as “true” (or “meaningful”, not merely technically focused). Although present-focused MFIs are usually associated with tonal moments – enjoying, and consequently allowing listeners to enjoy, either the tonally stable or the unstable points with the implicit aim of making sense of the tonal process –, present-focus extended in time allows for not merely cognitively representing but also mentally immersing into gesture and character.
Past-focus: Recollection
A second function of MFIs during performance is recollection (or retrospection): at certain points of the musical process, such as at structural boundaries, the performer, in order to form a mental representation of the previous musical unit, is reflecting back to the segment that has just ended. Typically, past-focused MFIs involve tonal and temporal retrospection on the previous musical unit: at the end of a structural unit the performer recalls in their imagination the feeling of the length and tonality of that unit. The musical unit to be recalled can be of any length, including a pair of notes (or even one single note), which is in fact the shortest grouping unit. Further to this, prompt attentional immersion with a focus on the past has a marked role in cognitive processing as it contributes to clearing the performer’s working memory. 7
Future-focus: Anticipation
The third vital function of MFIs is to position into the future, by anticipating the duration of upcoming – usually hierarchically embedded – structural units through imagining and feeling their length together with their affective colour (including the immediately expressed emotions and the gestures). Usually, anticipation occurs when the performer cognitively measures subsequent units to the previous ones in order to keep the length ratios of structural units and cognitively position the subsequent unit within the full temporal process (for a measurement of differences between expert and amateur performers related to the shaping of length ratios see Langner et al., 2000; for an empirical investigation of differences between performers’ timing strategies perceived by listeners as “exceptionally expressive” and “average” see Stachó, 2015). The intelligibility and the expressive quality of a performance is hypothesized to be correlated with the pre-imagining and pre-feeling of the length of ensuing structural units (notes, motifs, phrases or larger sections) in the moment before attacking them.
Anticipation is instinctively used by musicians to various degrees through mood induction. Mood induction, producing transient shifts of consciousness, has been found to be actively applied by performers to evoke particular emotions during performance that form the basis of their understanding of the music. An enlightening, though typical, example is given by Persson (1993, p. 197, emphases added in the original) as part of his wide-ranging qualitative research into performers’ phenomenology: I find often that to get me into the mood of a piece, say the Pathétique Sonata [by Beethoven] and the opening of that, you think something sad. You think sadly, not necessarily something that has happened to you, but you think of the experience of sadness before you play that chord. I suppose one could also think of something specific. But that “feeling of sadness” – and that’s what I’m trying to say – is a subconscious thing and you are just trying to bring it out.
Mental navigation and virtuosity
Performance excellence is closely linked to the ability to cognitively control the process of recollection and anticipation at any time scale, regardless of the actual awareness of the act of control. The above-described “navigating” mental imagery, which includes directing the attention forward (anticipating), backward (retrospecting), and to the present moment at well-definable points of the musical process, not only significantly contributes to the perceived expressivity, intelligibility and individuality of a performance, but also helps the musician to feel security and ease during performance. Furthermore, performers’ navigating imagery significantly contributes to technical security through an enhanced cognitive control of fine motor movements.
Attributes such as quickness, mastery, ease and sophistication, together with the quality of “being transported to another world” (cf. shifts of consciousness) are typically related to virtuosity, as has been revealed by Jane Ginsborg’s recent questionnaire study (Ginsborg, 2018, this issue). In fact, it is especially the mastery, sophistication and ease of mental navigation with its very frequent and quick shifts of attention and consciousness that is hypothesized to define the quality of a performance. And similarly to virtuosity as understood by Ginsborg’s respondents, performance-related navigating mental imagery appears to be a skill that can be mastered through practice (note that respondents tended to adopt the view that virtuosity in music performance results from “hard work” rather than a natural gift).
Thus common understanding among contemporary musicians, as well as conceptualizations by outstanding musicians of the past, and historians and theorists from disciplines ranging from music and aesthetics to anthropology (reviewed in Ginsborg, 2018, this issue) profess that virtuosity is not only about being quick and technically polished (i.e., making sound as many notes as possible) in a time frame but is very much about how to fill the time. One can be fast without being expressive (that is, without really reliving the subjectively conceived musical content and conveying it to a potential listener), and performers’ expressivity depends not on bodily quickness but rather on their ability to direct attention and their mastery of imagination. To illustrate this point, compare the performance of a virtuoso piece par excellence (Liszt’s La campanella) by two players, a presumably 12-year-old child (Rachel Su) 8 and a more experienced, 24-year-old pianist 9 whose Liszt performances have been portrayed by music critics as “highly virtuosic” (Miller, 2018), Haochen Zhang (see Video examples 2 & 3 in the Supplemental Material section).
The length of these performances is absolutely identical: both last 4 minutes 53 seconds. However, the latter, compared to the performance of the less experienced child, feels not only less monotonous but more virtuoso as well. This casual comparison can eloquently illustrate the claim that virtuosity is not merely about playing quickly but rather how to fill a given time frame, for what best contrasts the two performers is definitely not that Zhang has better or faster fingers. The difference lies in the mental realm rather than in the mechanics: in contrast to the child, the mature pianist operates a special skill to “interestingly” fill the time. A thorough examination of the contrast between the two performances at the cognitive level, including a tentative in-depth analysis of how differently the two pianists direct their attention, might reveal that the expressive features used by Zhang result from the specific mental navigation defined in the previous section.
It is worth noting that a similar navigating mental imagery, including directing the attention forward (that is, anticipating) and backward (i.e., retrospecting), is a core ability leading to excellence in sports as well. As a remarkable example, recent research into attentional control in sports revealed that an outstanding football or tennis player, compared to a less experienced player, develops faster eye movements in order to be able to anticipate where the ball is going to move rather than looking only at the ball (e.g., in soccer: Savelsbergh et al., 2002, Vestberg et al., 2012, cf. also Wimshurst, 2012; in tennis: Singer et al., 1996, Rowe & McKenna, 2001, Jackson & Mogan, 2007, Williams et al., 2011; for a review on 40 years of research on anticipation in tennis see Crognier & Féry, 2007). At the same time, while dribbling towards a defender, the player automatically watches for movement clues (see an illustrated report of Zoe Wimshurst’s eye tracking research on Cristiano Ronaldo’s attentional strategies during soccer playing: McDowall, 2011).
Quick attentional shifts: The expression and communication of narrative–dramatic structure
Strikingly similarly to top sportsmen’s skills underlying temporal and spatial awareness, the mental imagery skill related to expressive music performance centres on MFIs yielding intensive and very quick attentional focusing and re-focusing. A revealing illustration of the role of quick attentional shifts in the expression and communication of the musical meaning is provided by a pair of video-recorded excerpts from a masterclass with Maxim Vengerov teaching a Beethoven sonata movement (Op. 23, see Video examples 4 & 5 in the Supplemental Material section). 10
While working on the consecutive two-bar musical motifs in the opening movement of the Beethoven sonata, bearers of different (indeed, conflicting) gestures and characters that build up the narrative–dramatic structure of the movement, Vengerov shows and tries to explain to the student the importance of the attentional shifts required here, which can be likened to a very quick, in fact virtuosic, positioning into the perspective of another character (note that the metaphor Vengerov spontaneously used refers to the perspective of another “person”, which brilliantly portrays character). The second excerpt is a testimony to a spectacular boost (presumably a result of Vengerov’s explanations) of the violin student’s attentional processes: she has learned to shift her attention more quickly and to control this process in a more solid and adjusted way which is essential in expressing an important layer of musical meaning, the narrative. In fact, in this sense, virtuosity can be likened to general mental capabilities such as intelligence, encompassing planning and quick cognitive adaptation to the environment (Gottfredson, 1997).
Anticipating and recollecting: The expression of temporal structure
Quick future- and past-directed MFIs (that is, quick attentional positionings into the future and past of the musical process) are typically required to express and communicate changes in both metrical and grouping structure. Evidence from cognitive music theory (see especially Dobszay, 2012; but partly also Lerdahl & Jackendoff, 1983), empirical investigations (Hannon et al., 2004) and pedagogical practice point to the fact that in order to make it salient (viz., accented), the beginning of a metrical or melodic/thematic group is marked for consciousness, usually resulting in a temporal delay. Relying on observation of outstandingly expressive performers, I concur with the stance (elaborated by Dobszay, 2012) that cognitive processing at work behind the generation of temporal accents expressing/communicating the starting point of a metrical or thematic group involves active anticipation. In a performance perceived by listeners as highly expressive, active anticipation is generally achieved through MFIs in which the musician “positions” herself into a subsequent metrical or grouping unit, typically by anticipating its duration by projecting the feeling of its length. Usually, this occurs by cognitively measuring the subsequent units to the previous ones to successfully concatenate them (also consistently with Levinson’s [1997] understanding of this latter concept).
Active anticipation is often linked to visuo-spatial imagery and gestural metaphors (Stachó & Holics, 2011; Stachó, 2016). This can be ideally illustrated by the following excerpt from a masterclass with Maxim Vengerov where the student is invited to use the metaphor of throwing a ball to a prefigured distance in order to appropriately feel and express the difference between the pair of metrically stressed and unstressed beats that start the solo melody of the first movement of Mozart’s G major violin concerto, KV 216 (see Video example 6 in the Supplemental Material section). 11 The gestural metaphor employed by Vengerov not only guided the student’s attention but also provided a cognitive framework for the act of anticipation. For the student, the act of envisioning with full attentional concentration the image of throwing the basketball into the basket resulted in a clearly perceptible MFI, a subtle but well-detectable momentary change of consciousness.
Similarly to anticipation, the act of active recollection is associated with structural boundaries; however, it relates to the closing moment of a structural unit, irrespectively of its length (Dobszay, 2012; Stachó & Holics, 2011). Usually, active recollection is achieved through an MFI involving an instantaneous cognitive reflection on both the tonal trajectory and the feeling of the length of the unit. Note, however, that the speed of the attentional shifts, hence the length of the MFIs, related to the cognitive recollection usually depends on the overall tempo of the excerpt. For example, in a masterclass led by Steven Isserlis on Rachmaninov’s Sonata for cello and piano in G minor (Op. 19), the slowly fading section endings require longer moments of recollection in order to bring out full expressivity by transcending, as Isserlis claims, “commonplace” renditions of the concluding phrases of this Lento introductory section of the opening movement (see Video example 7 in the Supplemental Material section). 12
In the following excerpt (see Video example 8 in the Supplemental Material section), Isserlis explicitly points out the act of “looking backwards” and, as he puts it, “pushing forward”, relating to both cognitive recollection and anticipation at a significant structural boundary, producing very clearly perceptible MFIs. 13 Finally, a further excerpt from a masterclass with Swedish cellist Frans Helmerson (see Video example 9 in the Supplemental Material section) provides a perceptive demonstration of how structural concatenation effectively works through cognitive recollection followed by active anticipation. In the first movement of Dvořák’s Cello concerto (Op. 104), at an important structural boundary separating two sections bearing contrasting characters, the recorded excerpt illustrates the difference between a poorly and a capably executed retrospection and anticipation during performance. At the moment of the long closing note (D#) of the cello melody (which can be felt at the same time as an upbeat to the following section), the teacher invites the student to reflect back on the previous musical unit, taking advantage of a gestural metaphor to induce cognitive recollection, followed by an anticipation of the subsequent musical material by actively envisioning both its character and starting metrical position, giving rise to an enhanced sense of coherence for both performer and listener. 14
Prompt positioning into the present moment: The expression of tonality and character
Present-focused MFIs produce momentary immersion into the present musical instant, involving well-discernible, subtle shifts of consciousness. This type of MFIs allows the musician – and, arguably through empathy or emotional contagion, the musician’s audience – to fully enjoy the present sounding moment, thus fulfilling one of the pivotal functions of music-making. Present-moment focused MFIs on a salient note (or chord) of a grouping phrase usually last for a fraction of a second depending on the actual length of the note. These MFIs’ cognitive purpose can be approached from music theory: the marking of both tonally stable and unstable moments (see e.g., Bigand & Poulin-Charronnat, 2016) makes the tonal process of a composition intelligible and meaningful – in fact, felt – at various hierarchical levels of the tonal structure/process. In addition, very often these intensive present-focused MFIs allow for a short-lived affective immersion into the character of a musical section (and partly the gesture inherent in it); this is hypothesized to help the musician to express and communicate the character (and the gesture) to the listener. Finally, present-moment focus is frequently followed by re-focusing involving cognitive retrospection and anticipation, thus helping the performer efficiently control the process of expressing and communicating the metrical and grouping structures.
The following two demonstrations taken from music performance masterclasses can tellingly illustrate present-moment focused MFIs. In the first video, the masterclass leader (Steven Isserlis) immerses into the highest and longest note, which is both melodically the most accented and tonally the most stable note in the grouping phrase in question from Schumann’s first Fantasiestück from Op. 87. Here a present-moment directed MFI brings expressivity and individuality to the melodic phrase and allows the performer to transcend the “commonplace” (as Isserlis often puts it). Note that it is not the mere prolongation of the D that matters; rather, the quality of the performer’s relation to it is likely to produce specific (psycho)acoustic patterns resulting in a refinement of the note offset (see Video example 10 in the Supplemental Material section). 15
In an earlier excerpt already seen, pianist Maria João Pires explicitly makes use of the metaphor of space to characterize the feeling of immersion into the present moment (see Video example 1 in the Supplemental Material section). Remarkably similarly to the previous excerpt, at the end of the Beethoven variation (bar 6 of variation No. 31 from WoO 80) here it is not the mere lengthening of the Ab alongside its subdominant chord that makes Pires’s rendition so expressive, creating a feeling bordering on “endlessness”, but rather the depth of attentional immersion into that moment, allowing one to relive the character of the variation and to fully enjoy the tonal moment. 16
Further to the MFI that makes the subdominant so memorable, the act of cognitive anticipation and retrospection is well audible and observable on this video-recorded excerpt in many instances. Note that the process of how Pires directs her attention throughout the entire eight-bar period of the variation is strikingly similar to the mental processing observed in outstanding soccer players who are nearly always automatically (and instinctively) able to anticipate where the ball is going to move rather than only looking at it. At the same time, they direct their attention to the surroundings, as well as where the ball has started from in order to mentally plan its trajectory. In a conspicuously similar way, in music performance the performer sets the musical goals for herself and feels them: she feels in advance where the motif, the phrase or the larger section will end before she attacks it. Note that this is more than just knowing where to aim at in the musical process: here there is an intelligence-like, non-conscious, procedural knowledge at work – a kind of “mental dexterity” (in contrast to, e.g., “finger dexterity” so often associated with pointless virtuosity).
Rapid modulation of the depth of attention
A further phenomenon related to the virtuoso control of attention is the rapid modulation of its depth. The depth of attentional focus in MFIs is hypothesized to be particularly associated with highly expressive performances, and the ability to virtuosically manipulate it is a vital element of “mental dexterity”. In one of the most captivating music videos on piano playing from the middle of the 20th century, Alfred Cortot characterizes a little Schumann piano piece, The Poet Speaks (Der Dichter spricht, Op. 15 No. 13) as “a kind of intimate reverie”. When the camera shows him playing the piece, it is possible to observe how Cortot, while being in constant attentional immersion, deepens his attentional focus at certain moments – typically, in order to trigger anticipation at the beginning and cognitive retrospection at the end of the musical phrases. At such moments, the performer’s gaze is able to vividly reveal MFIs related to the cognitive processes of anticipation and retrospection (see Video example 11 in the Supplemental Material section). 17
Conclusion: The concept of “mental virtuosity”
This is the first theoretical study on music performance to argue for a concept of mental acuteness, indeed “mental virtuosity”, in music performance claiming that rapidity, vividness, passion and intensity associated with a virtuoso performance is hypothesized to be based on a well-definable “mental dexterity” at work in real time in the act of performance. The mental–attentional processing specified and illustrated in the present paper involves the ability to quickly position into different temporal and empathic perspectives (e.g., similarly to projecting oneself into another person’s position – cf. the Vengerov masterclass excerpts [Video examples 3 & 4 in the Supplemental Material section]), thus producing MFIs, in order to mentally represent the subjective meaning of music in real time. MFIs inducing subtle and brief shifts of consciousness associated with present-focus typically allow for a momentary but focused enjoyment of both character and tonally salient points, usually associated with either tonal stability or departure from the tonal context. While momentary future-focused MFIs mark the starting points of units of the temporal structure, past-focused attentional absorption (MFI) allows for an active momentary recollection of the length and the tonal trajectory of the previous musical unit – usually within a fraction of a second during performance. Smart and quick attentional shifts producing MFIs mark expressive and convincing execution of changes in gesture and directly expressed emotion.
The theory of performers’ attentional processes and strategies presented here suggests that the key qualities of a virtuoso performance in fact pertain to the cognitive domain – the domain of imagery and attention – rather than the mechanical. They rely on a specific and well-definable mental technique that produces both heightened expressivity and technical brilliance, defining elements of “true” virtuosity. Indeed, to become a “true virtuoso”, a musician needs to virtuosically manipulate their own attention in such a way as to be able to predict, reflect, and enjoy, so that the audience can also enjoy their performance.
Footnotes
Acknowledgements
I would like to thank the two reviewers of the first draft, Daniel Leech-Wilkinson and Kai Köpp, for their valuable comments and suggestions to improve the presentation of the argument.
Funding
The author was a recipient of a “New National Excellence Programme” award (Hungary) during the preparation of the paper.
