Abstract
Studies in musical improvisation show that musicians and even children are able to communicate intended emotions to listeners at will. To understand emotional expressivity in music as an art form, communicative success needs to be related to improvisers’ thought processes and listeners’ aesthetic judgments. In the present study, we used retrospective verbal protocols to address college music students’ strategies in improvisations based on emotion terms. We also subjected their improvisations to expert ratings in terms of heard emotional content and aesthetic value. A qualitative analysis showed that improvisers used both generative strategies (expressible in intramusical terms) and imaginative, extramusical strategies when approaching the improvisation tasks. The clarity of emotional communication was found to be high overall, and linear mixed-effects models showed that it was supported by generative approaches. However, perceived aesthetic value was unrelated to such emotional clarity. Instead, aesthetic value was associated with emotional complexity, here defined as the heard presence of “nonintended” emotions. The results point toward a view according to which the expressive content of improvisation gets specified and personalized during the very act of improvisation itself. Arguably, musical expressivity in improvisation should not be equated with the error-free communication of previously intended emotional categories.
Introduction
Emotions are assumed to be at the heart of what music conveys, and the successful expression of emotions is generally considered an essential part of musicianship. Although musical expressivity is a multidimensional phenomenon (Juslin, 2003, 2019), musicians themselves often understand expressivity as “communicating emotions” (Lindström et al., 2003). Sloboda (2005) even suggests that managing the link between musical structures and human emotion could be the central defining factor for all musical expertise. Not surprisingly, then, there is evidence showing that composers are able to convey intended emotional qualities to listeners (Thompson & Robitaille, 1992). Likewise, research shows that musical performers’ intentions to express given emotional qualities through previously composed melodies may affect the measured parameters of performance and that such intentions may thus be successfully communicated to listeners (e.g., Akkermans et al., 2019; Gabrielsson & Juslin, 1996; Juslin, 2000; for reviews, see Juslin, 2019; Juslin & Timmers, 2010).
Communicating emotions may be even easier in musical improvisation because the performer is able to influence the compositional structure. Accordingly, there are studies showing that musicians (Behrens & Green, 1993), music therapists (Bunt & Pavlicevic, 2001), and children (Volland & Hofmann, 2003) may successfully communicate intended emotions to listeners through freer forms of improvisation even though there may be differences in the communicability of some emotions (Gilboa et al., 2006). In fact, music therapists often assume that even musically noneducated patients can be helped to “articulate their emotions” (Bunt & Pavlicevic, 2001, p. 182) through musical improvisation (see Pavlicevic, 2000). Likewise, music educators frequently take students’ capability for such communication for granted. Consider, for instance, improvisation games for classical musicians in which participants are expected to depict distinct emotions on their instruments (e.g., Agrell, 2008, 2016). In such pedagogical contexts, we can say that emotion terms are used as points of departure for improvisation.
Improvisational tasks can be defined by various kinds of generative ideas, structures, or materials that researchers have variously called “models” (Lortat-Jacob, 1987; Nettl, 1974), “referents” (Pressing, 1984), or “points of departure” (Nettl, 1998, 2009). In a pioneering article, the ethnomusicologist Bruno Nettl (1974) defined improvisational models largely in intramusical terms, using concepts from music theory. However, just as composers may derive inspiration from programmatic contents or other conceptual frameworks (Katz & Gardner, 2012), so can improvisers sometimes rely on extramusical points of departure. These might range from the copperplate prints that the Viennese author Franz Grillparzer used to guide his keyboard improvisations (Keller, 1908, p. 42) to the texts used by the composer Karlheinz Stockhausen (1978) to guide his “intuitive music.” Educators have found such “semantic tasks” to be especially useful in guiding children’s improvisations (e.g., Baldi et al., 2002), but extramusical imagery can also be used with adult learners to “invoke a particular sort of mindset” (Ford, 1995, p. 108) or to imbue cross-artistic improvisatory work with direct experiential meaning (Huovinen & Manneberg, 2013). Music therapists, too, have distinguished between “referential” and “nonreferential” improvisation, depending on whether the music is “organized in reference to something other than itself” or “according to strictly musical considerations” (Bruscia, 1989, p. 10). Notice that it is an open question whether the heard improvisation will—or even need—be recognized as referring back to the ideas or schemes that the improviser used; although people can easily attribute extrinsic meanings to heard music (Huovinen & Kaila, 2015), the communication of such meanings from the improviser to the listener is an empirical question.
Points of departure, in the aforementioned sense, are only one aspect of the improvisational process and do not encompass the totality of physical, cognitive, or sociocultural constraints for improvisation (e.g., Ashley, 2016). Between any point of departure and the resulting musical sound, improvisers almost by definition may take considerable liberties to choose their own approaches, often called strategies (e.g., Norgaard, 2011; Williams, 2017). The freedom of choosing one’s own strategies may explain the appeal of musical improvisation games, where game rules serve as points of departure (e.g., Agrell, 2008, 2016; Friedemann, 1969, 1973, 1974, 1983; Schwabe, 1992). Indeed, even free-improvisation pedagogues often define points of departure—ranging from “relatively specific prompts to get the music going” to “whole plans for how an improvisation might unfold” (Hickey, 2015, p. 434). In such contexts, the “freedom” and individuality of improvisation lie at least partially in the specific strategies chosen by the participants. Notice that what appears as an individual strategy is relative to what is publicly negotiated with other musicians. For instance, an improviser’s private plan for structuring music by dynamic contrasts might be embraced by the other musicians and collectively determined as a shared point of departure for the next performance.
Empirical studies of improvisatory strategies have mostly focused on jazz. By subjecting professional jazz musicians to stimulated-recall interviews after their performances, Mendonça and Wallace (2004) and Norgaard (2011) found improvisers to engage in temporal planning or monitoring their performance and in creative thinking that resulted in the generation and development of the musical material. In particular, Norgaard found that musical material was managed through four generative strategies: using previously learned patterns, giving priority either to the harmonic context or to the emerging melodic aspects, or repeating material played earlier in the solo (see also Norgaard, 2017). Moreover, in a rare study of classical musicians’ improvisational thinking, Després et al. (2017) identified five classes of strategies: preplanning strategies, conceptual strategies (organizing pitch, rhythm, or tonality), structural strategies, atmospheric and stylistic strategies, and real-time strategies.
An improvisatory strategy represents the specific attentional focus that an improviser chooses in a musical situation. In Williams’s (2017, p. 54) words, “strategies can involve placing attentional emphasis on structure, individual musical parameters . . ., ensemble dialogue, and metaphoric imagery . . . among other variable elements.” Improvisatory strategies might thus emerge from subconscious processes of audiation and motor generation when musicians pay attention to them. From a pedagogical perspective, students’ domain of improvisational strategies might expand simply as they become aware of their tacit procedures (see Hargreaves, 2012). The broadening of strategic thought could thus be seen as a major advantage in a pedagogical approach such as Cahn’s (2005), which is based on consecutive cycles of collective improvising and intervening discussion. In the present study, we used this approach to explore the communication of emotional expressivity through improvised performance.
By employing emotion terms as points of departure for improvisation, we hoped to turn our participants’ strategic attention to their own, perhaps otherwise unspoken approaches to emotional expressivity in performance. Research in music and emotions shows that musicians generally have a lot of implicit knowledge of how to use musical features to communicate emotional expression (see Gabrielsson & Lindström, 2010). The “lens model” presented by Juslin (1997, 2000, 2019) explains that performers encode emotions using acoustic cues that are partly redundant in the sense that the same emotional character may be conveyed through different sets of cues. Successful emotional communication ensues when similar cues are utilized by the listener to decode the emotional expression. There is some evidence that “decoding accuracy” might be especially high for such basic emotions as anger, sadness, or fear (see Juslin & Laukka, 2003). This model implies that emotional communication “will be more effective if the cues are used consistently” (Juslin, 2019, p. 152), and Juslin suggests that such consistency may be enhanced in professional performers (see Juslin, 2000; Juslin & Laukka, 2000).
Notice, however, that consistent use of expressive cues could easily be exaggerated. When individuals are asked to play short improvisations to convey given emotions (as in Behrens & Green, 1993), the task instructions might also encourage them to secure accurate communication by emphasizing stereotypical cues. For instance, banging a keyboard with the fists might guarantee success in communicating anger. Interestingly, Volland and Hofmann (2003) found that listeners were able to decode intended emotions more effectively in children’s improvisations compared to adults’ improvisations. The authors suggested that adults may less often experience emotions “in pure form,” but another explanation could be that the adult improvisers sought to produce aesthetically pleasing improvisations, which require more complex use of emotional cues. In more realistic, extended music-making, using the “right cues” might not suffice to sustain musical interest for very long. In a study with jazz pianists, McPherson et al. (2014) thus did not find a simple one-to-one correspondence between intended emotions and musical features but instead found that the pianists employed greater structural diversity that also implicated a wider range of emotion.
Aims
In the context of higher music education, the musical communication of emotion categories provides a useful starting point for discussing and developing emotional expressivity in improvisation. In this study, we addressed three aspects of this topic. First, the previous discussion of improvisatory strategies implies that Juslin’s lens model does not yet consider what lies in the musician’s attentional focus. If an expressive intention is understood as an extramusical point of departure, the musician may still freely choose different strategies to approach making music in the situation. In the language of phenomenology of action, one might say that the distal intention of conveying an emotion still leaves unspecified the proximal intentions (Pacherie, 2008) that are situationally connected to producing the musical sounds (see also Fidlon, 2011). Hence, our first research question was as follows: In college students’ improvisations on emotion terms, what sorts of strategies do the students use to approach the improvisation task based on emotional points of departure, and do their strategies differ between basic emotion terms?
Second, we also addressed emotional communication. While assuming that the communication of intended emotions to the listeners might be affected by the performers’ level of musicianship, our focus in this article was on the role of improvisatory strategies in emotional communication: Do the improvisers’ chosen strategies have an impact on the successful communication of emotions? If so, how?
Third, it seems that the “functional achievement” (Juslin, 1997) in which the listener grasps the performer’s expressive intention still does not guarantee that the music is deemed successful as music. It is an open question whether listeners’ aesthetic judgments of the music might reveal any simple relationship with their judgments of emotional expression or, more importantly, with their decoding accuracy of the intended emotions. The complex nature of listeners’ aesthetic judgments has been well acknowledged by Juslin (2019) himself, but we would like to urge that such judgments should more frequently be considered in studies of expressive communication in music. Hence, our third research question was: In college students’ improvisations on emotion terms, is perceived aesthetic value positively correlated with the successful communication of specific emotions?
Method
Participants
The participants were 16 music majors (10 females, 6 males) from two Finnish universities, with a mean age of 24.2 years (SD = 3.7) and 14.7 years of experience playing their primary instrument, on average (SD = 4.0). Ten of the participants listed piano as their primary instrument, four participants listed piano as a secondary or tertiary instrument (with an average of 9.3 years playing experience), and the remaining two participants had some experience playing the piano, being concurrently enrolled in a pop/jazz piano course. All participants except one reported having previous improvisation experience (pop, jazz, free improvisation, or classical), resulting in an average of 5.4 years of improvisation experience among these participants (SD = 5.0). Thirteen out of 16 participants had actively studied improvisation, with an average of 4.3 years among these participants (SD = 3.4). The participants received course credit for participation in this study.
Emotion Words
As points of departure for the improvisation sessions, we chose two sets of emotion words. The first set consisted of the basic emotion terms joy, sadness, love, fear, and anger (see Shaver et al., 1987), which have been used in numerous studies on music and emotion (sometimes, joy is replaced with happiness and love with tenderness; see Juslin & Laukka, 2003). The second set consisted of 10 nuanced emotion terms selected from the hierarchical emotion scheme in Shaver et al. (1987). Here, each primary emotion (e.g., joy) was represented by one milder (e.g., contentment) and one stronger emotion (e.g., triumph; see supplemental file, Table S1, included with the online version of this article). The words were translated to their Finnish equivalents.
Setting
The participants were assigned to groups of four based on their availability. For each group, three 90-minute sessions were scheduled on consecutive weeks. One group consisted of four females, one group of four males, and two groups with three females and one male each. The sessions occurred in a recording studio featuring a Yamaha C7 acoustic grand piano. The sessions were facilitated by the first author, and the second author was present as an observer and technical assistant. In all of the sessions, the participants were seated in a semicircle, facing the facilitator, but with their backs to the piano to eliminate the influence of visual cues on the listeners. The sessions were video recorded, and for the third session, an additional audio recording of the improvisations was made by using two microphones placed near the soundboard of the piano. The participants also filled out a questionnaire on their musical experience and background. The study conformed to the ethical guidelines of the Finnish university at which the work was carried out. The participants filled out informed consent forms in which their full anonymity was guaranteed.
Practice Sessions
Both during two practice sessions and an ensuing performance session, each participant performed five improvisations on the piano. The facilitator abstained from giving any guidelines about emotional expression in music and from showing any judgmental attitude toward the improvisations. The goal of the first two sessions was to familiarize students with the setting, to allow them to experiment with using the emotion words as a basis for improvisation, and to encourage an open, positive group dynamic that would minimize performance anxiety and promote discussion.
In the first session, the participants were shown a list of the basic emotion terms and told that they would be asked to take turns and play short solo improvisations depicting these emotions, taking any approaches to improvisation that they wanted. They were instructed to choose a word from the list and improvise a piece on the piano based on the emotion. While listening to each improvisation, the other participants attempted to identify the emotion among the given alternatives, writing down brief explanations on handouts. This was always followed by a brief discussion on the explanations, led by the facilitator. Finally, the performer was asked to name the intended emotion and to explain how he or she had approached the task. Five rounds of this guessing game were completed with each group.
The second practice session was identical to the first except that participants picked their points of departure from among the 10 nuanced emotion terms (see Table S1 in the supplemental file included with the online version of this article). Our pedagogical purpose was to familiarize participants with the possibly unclear boundaries between emotional categories and give them room to explore creative, less stereotypical ways to approach the emotions. Each participant performed five times (i.e., on half of the words available), and each performance was followed by a discussion on the listeners’ emotional conjectures and the performer’s improvisatory strategies.
Performance Session
The purpose of the third session was to collect data for the study (we call it performance session here, but the participants were not informed of any difference in importance between the sessions). Once again, participants were asked to perform based on the five basic emotions from the first session, but now one emotion term at a time was openly announced by the facilitator, and each of the performers took turns improvising based on that emotion. To minimize order effects, the order of emotion words was randomized for each group, and the performance order of the participants was rotated between successive emotion terms. After each round of performances on a given emotion, the facilitator conducted a brief group interview with stimulated recall. Here, sections of the recorded improvisations were listened back in turn, and the respective performer was asked to describe his or her approach to improvising on the emotion term in question. The other participants were encouraged to comment on the others’ improvisations, identify emotional nuances, and provide the performer with feedback.
Expert Ratings
The improvisations recorded in the performance sessions varied in length from 21 seconds to 2 minutes and 46 seconds (M = 1 min 11 s). These 80 improvisations—16 on each of the five basic emotions—were subjected to ratings by a panel of experts who were oblivious to the task used to generate the music. The four expert reviewers were lecturers at a university music department, two of them with a PhD in musicology and two in music education. Two of them were ethnomusicologists with extensive experience from improvising in various musical styles, whereas the third was an experienced music analyst and the fourth a music psychologist. Two separate review sessions (in a classroom setting) were organized for two experts at a time, with the first two experts hearing the improvisations in a random order and the second two hearing them in reverse order. Each expert listened to the 80 recorded improvisations, filling out responses to four questions for each improvisation. The first question, “How well do the following emotion words, in your opinion, describe this piece of music?,” required separate ratings for each of the five basic emotion terms, each of them on a scale from 0 to 4. In addition, three aesthetic ratings (on 0–4 scales) addressed the personal quality of the performance, the musical originality of the improvisation, and how well the improvisation worked “as a compositional whole” (compositional wholeness).
Data Analysis
Retrospective verbal protocols from the performance sessions were transcribed and subjected to qualitative coding by the second author, following the pattern coding method of Miles and Huberman (1994) with a single level of abstraction. The coding was concentrated on statements that clarified the participants’ approaches to the points of departure and/or to their choices of what and how to play based on the task. Statements addressing other topics, such as processes of monitoring and evaluation (Norgaard, 2011), were ignored. Based on the resulting succinct criteria for emerging categories, the first author then carried out an independent coding of the data, after which some disagreements were resolved together by refining the criteria for category membership.
Quantitative analyses, including the construction of linear models for the expert ratings, were carried out in the R statistical environment (R Core Team, 2019) using the packages emmeans (Lenth, 2019) and MuMIn (Bartón, 2019) for calculation of predicted values and the coefficient of determination, respectively. In all cases, qq-plots were used to confirm that the residuals of the models followed a normal distribution. Additionally, we used the psych package (Revelle, 2019) in R for principal component analysis of expert ratings, synchrony (Gouhier, 2019) for Kendall’s (1975) coefficient of concordance, and rmcorr (Bakdash & Marusich, 2021) for repeated measures correlation. Repeated measures correlation is used to assess the common within-individual association when pairs of observations are done repeatedly for the same individuals; it avoids violating independence assumptions and avoids first averaging the data (Bakdash & Marusich, 2017).
The experts’ interrater reliabilities were moderate both for the emotion and the aesthetic ratings (mean Kendall’s Ws = 0.55 and 0.41, respectively; corrected for ties). The mean ratings used in our statistical modeling thus represent expert listeners on a group level and may hide some differences of judgment that seem unavoidable in an aesthetic setting.
Results
Improvisers’ Strategies
The qualitative coding of the verbal protocols yielded five categories. As shown in Table 1, these were later grouped in two higher-order categories—generative strategies and imaginative strategies—depending on whether the spoken accounts centered on music-theoretical or extramusical concepts. Generative strategies thus concerned “ways of creating improvised material—making choices about which notes to play” (Norgaard, 2011, p. 118), whereas imaginative strategies concentrated on further developing the emotional point of departure. Regarding generative strategies, the performers often approached the improvisation task with certain music-structural features or certain technical means of performance in mind. This category of music theory included references to scales, keys, melodic motives, rhythmic patterns, and other similar aspects. At times, the participants also constrained their performances in more specific terms, using individual pieces of music as reference points—a strategy that is here called musical reference. These two sorts of strategies might occur together, too, as when one student reported approaching anger by thinking about Stravinsky’s Rite of Spring and by using a chord from the work, “E major, with Eb7 on top” (male, 23 years).
Generative and Imaginative Strategies.
The three imaginative strategies emerging from the discussions can likewise be arranged in terms of their specificity or, more appropriately, evocativeness—the extent to which there seems to be a contact with lived individual experience (see Petitmengin, 2006). In the approach called conceptual nuance, the participants simply specified the given emotion term by one or more adjectival qualifiers—such as in the Rite of Spring improvisation in which the participant noted that this approach, for him, “represents pure rage, a kind of instinctual rage.” As such, mere conceptual specification may be a relatively objectifying stance, not necessarily indicating closeness to a lived-through experience. At times, participants also went beyond mere linguistic specification of the emotion category by referring to general imagery, which concerned some visualized or situational content. Concentrating on an imaginative illustration of the emotion term might help channel the expressive intent to the improvisation: [I approach improvisation] pretty intuitively. I usually get a strong idea revolving in my mind—a kind of visualized thought of what is happening. It always becomes like a story in my mind. Then, the feelings come through that. (Female, 21 years)
Finally, the most experientially specific of the imaginative approaches were those that evoked personal experience—not relying on generic images but rather on individually recollected situations or personal encounters as sources of evocative imagery. Consider the following discussion in which a 21-year-old female student, in accounting for her anger improvisation, not only referred to a real-life situation but also switched perspective between her own experienced feelings, recollected movement qualities, a certain linguistic register, qualities of vocal expression associated with the situation, and even her attitudes toward herself:
Did you have a specific thought in mind?
Yes. I broke my computer’s hard drive yesterday. I shouted a lot, so I somehow tried to play that. I got so angry that I just shouted and cursed, and I was really clumsy.
Is it easier [to express anger] when you envision such a concrete situation?
Yes, when you just try to think about it. You try to describe how you were in the situation—like trying to shout. Like that bass line was as clumsy as [I was] when I dropped it [i.e., the computer]: so that’s how it was. I somehow expressed how I moved and how I felt.
Such highly specific, personalized imagery might conceivably restrict the musical results as well. Thus, a 20-year-old male participant suggested that in using more “general” imagery instead, “you are freer to do things.” By contrast, “If you think up a certain situation [from your life], or if it strongly comes to your mind, you easily stick to it and don’t go anywhere else.” Despite some such views, it also seemed that relating the music to specific personal experiences often helped the students gain in emotional depth and variety—especially when the experiences involved other significant persons of their lives. In the following 24-year-old male student’s account of his sadness improvisation, the memory of a dear family member does not seem to restrict the expressive possibilities in any negative sense. Rather, it evokes a meaningful personal perspective that allows empathetically reenacting and projecting experiential contents. The improviser assumes the “mental landscape” of his grandmother and his father—apparently not so much as a set of precise mental contents but more as a type of familiar perspective that one can take on:
The feeling state [of sadness] was rather topical for me. My grandfather died last week’s Saturday.
We express our condolences. Was he close to you?
Yes, rather close.
Does it affect [your playing] somehow? Did you have this on your mind right away?
Yes, clearly.
[After listening back to the recorded improvisation:]
Did you think about your grandfather’s personality when you played this?
No. Actually, I thought more about my grandmother, and in a way lived through her—lived her mental landscape. And my father’s as well.
Based on coding the retrospective verbal protocols, generative strategies were present in 32.8% of the improvisations, whereas imaginative approaches were seen in 62.5% of them. Most of the participants used both generative (13 of 16) and imaginative (15 of 16) accounts, but the mean number of improvisations with generative strategies per participant (1.9) was lower than the corresponding number for imaginative approaches (3.1). The distribution of strategies as responses to the various emotion terms is shown in Table S2 in the supplemental file included with the online version of this article. It should be borne in mind that within both main categories, the codes might sometimes reflect strategic choices incompletely. For instance, using a specific musical reference might have prompted the conscious use of music theory to carry out the task even if this was not mentioned by the improviser, and a mention of general imagery might itself be just an incomplete description of personal experience. It may thus be safest to observe counts of when any of the strategies in a given class were present. According to a χ2 test, there was no significant difference in the use of generative strategies between emotional categories, χ2(4) = 8.11, p = .088, although results show that anger mostly did not evoke generative strategies (see Table S2 in the supplemental file included with the online version of this article). By contrast, the presence of imaginative strategies did vary significantly between the emotion words, χ2(4) = 16.00, p = .003. Anger and fear, in particular, found the participants emphasizing extramusical imagery over and above musical materials or reference points (see Table S2 in the supplemental file included with the online version of this article).
Emotional Communication
In the first practice session, 81.8% of the guesses matched the intended emotion. In the second session, decoding accuracy was lower for the longer list of nuanced emotions (56.5%), and it remained at 64.4% even when the guesses were categorized according to the parent basic emotions. Given this categorization, a Wilcoxon signed-rank test on the mean percentages of right guesses for each participant’s improvisations showed a significant difference between the two sessions (W = 135, p < .001). This supports our informal observation concerning a change of approach between the two sessions. In the first session, discussions often revealed a match between the performer’s simple cue-based approach (e.g., using major seventh chords for “love”) and the listeners’ reported reasons for their guesses. In the second session, the participants appeared to be less content with signal-like communication of the emotion term as their primary focus, seeking more elaborated and “musical” approaches. That said, we turn back to the main results.
The expert listeners’ ratings suggested that emotions were also successfully communicated in the performance session, considering that in 49 out of 80 cases, the intended emotion received the highest mean rating among the five emotions, (one-sample proportions test with null probability 0.2: χ2[1] = 85.08, p < .001). Often, however, other emotions in addition to the intended ones received notable ratings. Let us define emotional complexity as the mean rating of nonintended emotions. For instance, if the expert gives a sadness improvisation a rating of 4 for sadness but ratings 3, 3, 1, and 0 for the four other emotions, the improvisation has an emotional complexity rating of 1.75. In this case, the intended emotion still rates higher than the other rated emotions by a score of 2.25 (4 − 1.75 = 2.25). We consequently define this difference between the rating of the intended emotion and the emotional complexity rating as emotional clarity. (Here, we basically follow Behrens and Green’s [1993] measures but use terminology with less emphasis on the idea of “correct” guesses.) Notice, therefore, that even if the ratings show some amount of heard emotional complexity, emotional clarity indicates how clearly the intended emotion is perceived compared to the other emotions. In our data, the mean clarity of 1.28 (SD = 0.86) shows that the rating of the intended emotion typically exceeded the other ratings by a rather large difference compared to the mean complexity of 0.61 (SD = 0.26). According to repeated measures correlation, these two measures were slightly negatively correlated (rrm = −0.24, 95% confidence interval [–0.46, 0.01]), but the small effect size supports retaining two separate measures.
To address our second research question regarding the effects of improvisational strategy on emotional communication, we constructed statistical models for emotional complexity and clarity. Initially, we attempted to build linear models also involving the musical background variables (reported in the “Participants” section) as explanatory variables, but because none of them could be used to improve the models, we changed the approach (the same is true for the models concerning aesthetic value reported in the following). We chose to use linear mixed-effects models, taking the participant as a random effect and considering the intended emotion term and the presence of various improvisatory strategies as explanatory variables. This included binary variables corresponding to the presence of the five strategies identified in Table 1 and two similar variables indicating the presence of any generative strategy and any imaginative strategy. The strategy variables were first screened by correlation analyses with emotional complexity and clarity; variables showing significant or near significant correlations with one of them were considered in likelihood-ratio tests to be included as fixed effects for the result variable in question. Two-way interactions between the predictors were also considered.
Model summaries are shown in Table S3 in the supplemental file included with the online version of this article. In the case of emotional complexity, we successively added to the null model (with only the random effect) fixed effects of emotion term, likelihood-ratio test: χ2(4) = 13.83, p = .008, and the personal experience strategy, χ2(1) = 4.91, p = .027. No other fixed effects or interactions could be added to improve the model. Regarding the effect of emotion, estimated marginal means suggested lower emotional complexity for joy improvisations (0.51) than for the other emotions (sadness: 0.61, love: 0.65, fear: 0.74, anger: 0.76). In addition, the model predicted higher emotional complexity for the use of the personal experience strategy (0.74) than when this strategy is not used (0.58). According to the conditional R2, the model accounted for 32% of the variance, suggesting a “moderate” fit (Ferguson, 2009; marginal R2 = .19).
The chosen model for emotional clarity included only the random effect of participants and a fixed effect of generative strategy, χ2(1) = 5.59, p = .018. No other fixed effects (or interactions) could be added to improve the model. Predicted values suggest that the use of generative strategies yields higher emotional clarity (1.56) than would be the case without such strategies (1.10). This is in line with the idea of cue-based emotional communication described by Juslin’s lens model. However, model fit was quite poor: According to conditional R2, only 12% of the variance was accounted for by the entire model (marginal R2 = .07).
Based on the results in this section, it appears that although using musical cues may to some extent support conveying the intended emotion, there is also another aspect of emotional communication—emotional complexity—that should be taken into account. In particular, the emotional complexity of improvisations may significantly differ between various emotion terms, and such complexity may also increase due to the improvisers’ imaginative reference to their own personal experiences.
Perceived Aesthetic Value
We now turn to the experts’ aesthetic ratings: personal quality, musical originality, and compositional wholeness. To assess the effect of emotional communication on perceived aesthetic quality, we first compared the group of 49 improvisations where the highest emotion rating corresponded to the intended emotion with the 31 improvisations where this was not the case. According to preliminary one-way analyses of variance, for none of the three direct aesthetic judgments was there a significant (p < .05) difference between the two groups of improvisations. This suggests that successful emotional communication did not significantly predict perceived aesthetic value. This result was also supported by repeated measures correlation analyses, relating the previously defined variables of emotional complexity and clarity to experts’ mean aesthetic ratings and emotion ratings (see Table 2). Although the aesthetic ratings showed high mutual correlations, they were not similarly correlated with emotional clarity. In other words, the transparency of expressive intention apparently did not lead to higher aesthetic judgments. Instead, the perceived personal quality of the improvisations turned out to be correlated with emotional complexity.
Repeated Measures Correlations Between Aesthetic Ratings, Emotional Ratings, and Variables Regarding Emotional Communication (n = 80).
Note. Significance levels are adjusted for multiple comparisons: *p < .05/45, **p < .01/45, ***p < .001/45 (df = 63).
Table 2 also highlights two other aspects regarding listeners’ reactions. First, in the ratings, emotions were grouped into pairs anger/fear and joy/love. Second, sadness ratings also showed significant positive correlations with aesthetic ratings. Both of these aspects received support from principal component analyses using the original expert ratings (n = 320). A first principal component analysis (with varimax rotation) on the emotion ratings revealed three distinct dimensions, loading on joy/love (proportion of variance explained 30.2%), anger/fear (28.6%) and sadness (21.4%; cumulative variance explained 80.2%). In another principal component analysis, we also included the aesthetic ratings (see Table S4 in the supplemental file included with the online version of this article). Now the aesthetic ratings were grouped together in one principal component, together with sadness ratings (proportion of variance explained 29.4%), followed by two other components comprising joy/love (18.8%) and anger/fear (18.7%; cumulative variance explained 67.0%).
Based on the latter analysis, it appeared that the three aesthetic ratings converged on a more or less unified value dimension. We thus defined a synthetic variable of aesthetic value, corresponding to the component scores of the first principal component (Table S4 in the supplemental file included with the online version of this article), normalized between 0 and 1. We concluded our analysis by constructing a linear mixed-effects model for this variable, aesthetic value. Here, we used the same approach and considered the same explanatory variables as in the previous section. Corresponding to our third research question, however, we also considered emotional clarity and complexity as potential predictors. A summary of the chosen model is shown in Table S5 in the supplemental file included with the online version of this article. In addition to the random effect of participants, it included fixed effects of emotion term, χ2(4) = 41.23, p < .001, and emotional complexity, χ2(1) = 22.58, p < .001. No other fixed effects or interactions could be used to improve the model. In particular, emotional clarity could not be shown to affect aesthetic value. After removing one outlier (based on a qq-plot), the coefficient of determination suggested “strong” model fit (Ferguson, 2009) for the full model (conditional R2 = .73; marginal R2 = .41). Predicted values from the model were in line with our main observations from the preliminary correlation analysis previously described. First, our model predicted higher aesthetic value for the sadness improvisations than for improvisations based on the other emotion terms and lower aesthetic value for the joy improvisations than for others (see Figure 1a). Second, the main effect of emotional complexity suggested that aesthetic value increases with emotional complexity of the improvisations (see Figure 1b).

Aesthetic value in the improvisations. (a) Predicted values for the five emotion terms, with significance levels in Tukey post hoc tests. (b) Predicted values for different levels of emotional complexity, plotted against the original data (“x” = outlier). The whiskers in Figure 1a and the dotted lines in Figure 1b indicate standard errors. *p < .05. **p < .01. ***p < .001.
Discussion
If managing the expressive link between musical structures and human emotion is central to musical expertise (Sloboda, 2005), it should also be a key concern for the teaching and learning of musical improvisation. Previous research has demonstrated that when asked to do so, musical improvisers are able to reliably convey intended emotional categories to listeners (Behrens & Green, 1993; Bunt & Pavlicevic, 2001; Gilboa et al., 2006; Volland & Hofmann, 2003). Arguably, such emotional communication might also take place rather simplistically by using stereotypical cues. In our present study, we thus also addressed what strategies improvisers use when intending to convey emotional characters through music and what determines listeners’ aesthetic judgments in such a context. In the study, higher education music students produced free-form piano improvisations based on basic emotion terms, explaining their creative strategies after each performance. The improvisations were rated by an expert panel for five basic emotion categories and for three evaluative aspects: personal quality, musical originality, and compositional wholeness.
Qualitative analyses of the improvisers’ strategies showed that, quite like jazz improvisers (Norgaard, 2011), our participants used some generative strategies, focusing on musical materials to convey the intended emotional qualities. Even more frequently, however, they adopted imaginative strategies. In these cases, the emotional points of departure were further specified in extramusical terms—by evoking more specific emotions, general imagery, or personal experiences. Interestingly, imaginative strategies were more frequently utilized in improvisations conveying “negative” emotions (anger, fear, sadness) than “positive” emotions (joy, love). Our generative and imaginative strategies correspond to the two main approaches to developing musical expressivity distinguished by Woody (2000). In his study, college music students’ views differed on whether “instruction should address concrete musical concepts and physical sound properties, or instead focus on performer[s’] felt emotions or moods, trusting that these naturally translate into expressive performance devices” (p. 21). Although we could not explicitly address the extent to which the participants actually felt the emotions they specified through imagery, our coding suggests that imaginative recourse to prior subjective experiences was one of the main imaginative strategies. This would be compatible with Sloboda’s (2005, p. 288) suggestion that expressive musical performance centrally takes place through the use of templates learned from extramusical experience.
In the practice sessions of our study, the participants were highly successful in communicating intended emotions to each other. Likewise, the improvisations from the performance session were successfully “decoded” in this sense by the expert listeners. For more detailed analysis, we defined measures of emotional complexity (mean ratings for nonintended emotions) and emotional clarity (the mean winning marginal of the intended emotion in the ratings). As in Gilboa and colleagues’ (2006) study of music therapists’ improvisations, we were unable to explain effectiveness of emotional communication with variables regarding the participants’ musicianship. Instead, we found that focusing on generative strategies tended to slightly enhance the communicative clarity of emotions. More notably, perhaps, the complexity of emotional communication was supported by improvisers’ use of a personal experience strategy. Gilboa and colleagues (2006), focusing on emotional communication, interpreted a similar finding by suggesting that revivification of emotional memories hindered emotional communicability. A more positive interpretation might be to say that reference to personal experiences apparently relaxed the improvisers’ attentional focus from the plain communicative task, allowing room for a broader spectrum of expressive qualities to emerge.
Our study also provided an opportunity to explore the aesthetic consequences of perceived emotional clarity and emotional complexity. It is well known that music may be perceived as being expressive of “mixed emotions” due to conflicting cues in various musical parameters (Juslin & Timmers, 2010). For instance, fast music in a minor mode or slow music in a major mode may elicit mixed perceptions of happiness and sadness in listeners (Hunter et al., 2010). An account of musical value based on the communication of emotion might hold that such conflicting cues are detrimental to musical value. However, in our study, expert listeners’ aesthetic ratings revealed no significant relationship with emotional clarity. Instead, emotional complexity—the heard presence of nonintended emotional expressions—emerged as a significant predictor of aesthetic value. In particular, emotional complexity appeared to be connected with hearing the music as having personal quality. The pedagogical implication is that higher education music students may often be well beyond the point where the improvement of communicative unambiguity should be seen as a central component in developing expressive musicianship (cf. Juslin et al., 2006). In our practice sessions, the opportunity to engage in matching emotion labels with musical cues quickly led the students to push beyond the simple goal of communicating emotion labels—toward the goal of creating expressive, interesting, good music. Although our results should obviously be replicated with a larger sample, they do suggest caution in equating aesthetically compelling improvisatory expressivity with error-free, maximally obvious communication of specific emotional intentions.
We were also unable to find differences in the emotional clarity of improvisations based on the five basic emotion terms. In other words, there was no difference in the decoding accuracy between the emotions, as found in some other studies (see Juslin & Laukka, 2003). Yet the choice of emotion term did affect the emotional complexity of the improvisations as well as their aesthetic value. Most notably, sadness was the emotion term that best supported aesthetic value when taken as a point of departure. In designing pedagogies of musical expressivity, it may be good to have such a possibility in mind. Another finding with possible pedagogical consequences is the fact that joy improvisations showed both the lowest emotional complexity and the lowest aesthetic value. Interestingly, joy was also the emotion that most often led the improvisers to use generative approaches and least often inspired them to use imaginative approaches. It appears, then, that the participants approached joy in a more cue-based manner than the other emotions and that this line of action led to expressively narrow performances (i.e., with low emotional complexity), which in turn may have led to rather modest aesthetic judgments.
Although it was not our main concern, the experts’ ratings indicated that aesthetic value was also related to perceived emotion. Logically speaking, this is a separate matter from the relationship between intended emotion and aesthetic value. For instance, even though joy appeared detrimental to aesthetic value as a point of departure, correlations between experts’ emotion and aesthetic ratings did not indicate that joy would have been particularly disfavored as a perceived emotion. Indeed, the only perceived emotion that we found significantly connected to the aesthetic judgments was sadness: Ratings of personal quality and compositional wholeness were positively correlated with sadness ratings quite apart from their correctness with respect to performers’ intentions. This is in line with previous research showing high ratings on preference and beauty in music heard as sad (Eerola & Vuoskoski, 2011; Vuoskoski et al., 2012) and the relative pleasantness of sadness-inducing music (Vuoskoski & Eerola, 2012). Given the high emotional clarity achieved, our results concerning intended and perceived sadness are, of course, interconnected. Using our present methods, we cannot securely distinguish experts’ possible preference of sad music from the question of whether improvisers in some more objective sense might have played “better” music in response to sadness.
As in any research relying on retrospective verbal protocols, one may ask whether the improvisers’ verbal accounts might just represent later rationalizations of prior action (see Ericsson & Simon, 1993). That is, instead of accounting for the strategies they had utilized to convey a given emotion, our participants may simply have described what they after the event realized to have accomplished. Rather than outlining plans existing prior to the musical event, some of the verbal accounts indeed mentioned what “started happening” during the improvisation (see Table 1). One might compare this with how Fidlon (2011), interrupting jazz improvisers and asking them what they were about to play, found that more experienced improvisers reported less musically detailed proximal intentions than novices. Such findings highlight how improvisers’ situational decisions may depend on the moment-to-moment evaluation of their own prior actions (see Pressing, 1988). To avoid any necessary implication of rational preplanning, one might want to replace the term strategies by speaking more broadly of the approaches taken.
In fact, however, the aforementioned criticism may point toward an important interpretation of the expressive situation. According to the music-therapeutic conception of Hegi (2010), improvisation can be seen as a subjective experiment that helps the improviser to approach and elucidate an unclear emotion. The philosopher Andrew Bowie (2019) suggests something similar: “The expressive possibilities in music . . . can enable us to experience emotions that did not exist before the music that discloses them” (p. 755). Retrospective reports of imaginative strategies might thus be understood as more precise accounts of an expressive achievement given when the participants already had clarified for themselves in the improvisation what there was to express. Far from being just inaccurate later rationalizations of improvisatory strategies, imaginative accounts thus suggest that the task of conveying an emotional character had turned into real subjective expression. In this view, musical expressivity is not just about conveying to the listeners musicians’ distal intentions that were pronounced prior to the situation but also about realizing proximal intentions as they arise in the activity itself (see Pacherie, 2008). In the course of such expressive action, the expressive message is personalized. Instead of just successfully communicating previously chosen emotional categories, improvisation may come closer to “a validation of the emotional and spiritual sides of ourselves,” as one free improviser put it (Nunn, 1998, p. 76). For music education, the challenge will be to develop approaches to improvisation pedagogy in which specific expressive intentions—announced as fixed points of departure—are gradually complemented by more complex and dynamic, life-like notions of human expressivity.
Supplemental Material
sj-pdf-1-jrm-10.1177_00224294211044676 – Supplemental material for Improvising on Emotion Terms: Students’ Strategies, Emotional Communication, and Aesthetic Value
Supplemental material, sj-pdf-1-jrm-10.1177_00224294211044676 for Improvising on Emotion Terms: Students’ Strategies, Emotional Communication, and Aesthetic Value by Erkki Huovinen and Aaro Keipi in Journal of Research in Music Education
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Author Biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
