Abstract
Numerous studies have investigated the effects of pitch structures on perceived emotion in music, but the emotional effects of rhythm and meter have received far less attention. In the experiment reported here, we manipulated the metrical framework of music by Robert Schumann and asked participants to judge perceived emotion in the resulting excerpts. The distinction between metrical dissonance and consonance offered by Harald Krebs (1999) was the theoretical basis of this study. Stimuli were 10 metrically dissonant excerpts from Schumann’s Carnaval and 10 metrically consonant recompositions of these excerpts. Recompositions maintained original tempi, global meters, and harmonic frameworks. Participants—graduate-level pianists and non-musicians—heard all 20 excerpts in randomized order. On each trial, they chose a cluster of emotion words, based on Schubert (2003), that best described the excerpt, a cluster that worst described the excerpt, and rated their level of interest. Results indicate significant effects of both metrical character and expertise on perceived emotion. There were also significant differences in the likelihood of participants changing their responses between metrically dissonant and consonant versions as a function of musical excerpt. This last finding leads to suggestions for future investigations on degrees of metrical dissonance.
The current exploratory study investigates the effects of metric and rhythmic change on the perception of emotions. In other words, we address the emotional connotations of the music and how changes in musical structure affect such connotations, rather than the felt or induced emotions that participants might experience through music listening (Gabrielsson, 2002; Gabrielsson & Lindström, 2010). While the effect of pitch structures on perceived emotion has been reliably demonstrated (Crowder, 1984; Peretz, Gagnon, & Bouchard, 1998; Temperley & Tan, 2013), fewer studies have addressed the emotional effects of rhythm and meter. Rhythm is acknowledged as an important factor in emphasizing certain notes with affective implications (Lindström, 2006, p. 109) and a contributing factor in the composition of music with distinct emotions (Thompson & Robitaille, 1992). There is evidence that rapid rhythms are generally rated happier than slower rhythms, and happiness ratings strongly correlate with ratings of subjective tempo (Motte-Haber, 1968; Thompson & Robitaille, 1992; Schellenberg, Krysciak, & Campbell, 2000). Studies have also shown that exaggerated durational contrasts (specifically long-short patterns) contribute to differences in perceived emotion (Lindström, 2006) and can be used by a performer to modulate emotional expression (Gabrielsson & Juslin, 1996). Finally, “complex” rhythms have been associated with anger or instability (Lindström, 2006; Thompson & Robitaille, 1992). In all of the aforementioned studies, however, rhythm and meter are only two among several factors under consideration.
The study reported here focuses on the emotional connotations of music containing two or more layers of rhythm, each with separate accentual and grouping patterns (Roeder, 2001). The theoretical basis for the study is Harald Krebs’s (1999, 2003) conception of “metrical dissonance.” Building on the work of Yeston (1976) and Lerdahl and Jackendoff (1983), Krebs (1999) describes meter as a union of multiple layers of motion, each comprising regularly recurring pulses. The pulse layer is “the most quickly moving pervasive series of pulses, generally arising from a more or less constant series of attacks on the musical surface” (Krebs, 1999, p. 23). Interpretive layers are slower moving “perceptible phenomena” that organize the pulse layer into larger units. When two or more interpretive layers align with each other and with the pulse layer, the metrical state is “consonant”; a misalignment of layers, in turn, results in metrical “dissonance” (Krebs, 1999, p. 29). Figure 1 is an abstract representation of two specific kinds of metrical dissonance: a. displacement and b. grouping. In a displacement dissonance, two interpretive layers contain pulse units of the same size, but the layers are offset from one another; this is commonly known as “syncopation.” In Figure 1a, interpretive layers of 3-pulse units (ii. and iii.) are offset from each other by one pulse (“D3+1” in Krebs’s nomenclature). In a grouping dissonance, at least two interpretive layers contain differently sized pulse units; the most typical of these, a group of two against a group of three, is commonly known as a “hemiola.” In Figure 1b, an interpretive layer of 3-pulse units (ii.) occurs against an interpretive layer of 2-pulse units (iii., “G3/2”).

a. Displacement dissonance (D3+1): i. is the pulse layer, ii. and iii. are interpretive layers, each comprising 3 pulse units; iii. is displaced from ii. by one pulse. b. Grouping dissonance (G3/2): i. is the pulse layer, ii. and iii. are interpretive layers; ii. contains groups of 3 pulses, and iii. contains groups of 2 pulses.
Metrical dissonance can be understood as effecting a disruption to an established or implied hierarchical metrical pattern. Such disruptions are perceptible to most people (Ladinig, Honing, Háden, & Winkler, 2009; Honing, 2012; Biamonte, 2014; Parncutt, 1994; Hannon & Trehub, 2005), and rhythmic patterns that disrupt an established metrical pattern are more difficult to replicate through bodily movement than those that fit the established pattern (Povel & Essens, 1985; Patel, Iversen, Chen, & Repp, 2005; Ellis & Jones, 2009; Jones, 2016). The implications of these disruptions on perceived emotion have yet to be explored empirically, though theoretical writings support the idea that metrical dissonance affects perceived emotion. Rosen (1995) notes that metrical dissonance contributes to a sense of “shock value,” or “anxiety,” that may otherwise be absent in metrically consonant music (cited in Krebs, 1999, p. 20). His characterization may come from the perspective of a listener or a performer; as a pianist, Charles Rosen had first-hand knowledge of the mind- and hand-bending required to perform the metrical dissonances that so pervade Schumann’s work. Additionally, London (2004) associates metrical dissonance with conflicting cues (p. 100). Conflict may have an emotional association, depending on the perceiver. Such descriptions of conflict, anxiety, and shock value lead us to believe that the perceptions of emotions will change if metrical dissonance is eliminated from the music.
Metrical dissonance can be found in the music of many composers, including John Dowland, Arcangelo Corelli, Ludwig van Beethoven, Johannes Brahms, and many others in the 19th century, and as Biamonte (2014) has demonstrated, metrical dissonance can also be found in rock music. Within art music, however, Robert Schumann’s compositions are a particularly rich source of metrical conflict. Moreover, according to Krebs (1999), metrical dissonance serves a particular function in Schumann’s music: to reflect emotional or psychological states (p. 161). While Krebs suggests that metrical dissonance can convey “violence” (p. 161) or “excitement” (pp. 102, 180), exactly which emotions are conveyed is an open question and one that we examine here.
The present study investigates the emotional connotations of metrically dissonant excerpts from Schumann’s Carnaval and metrically consonant recompositions, as perceived by two groups of participants: trained pianists (musicians) and listeners without any musical training (non-musicians). All excerpts were performed by a concert pianist, and care was taken to make sure that paired original and recomposed excerpts were performed at similar tempi and with similar levels of rubato. Participants were asked to select emotion words that best conveyed the emotions they perceived in the music and words that were the worst in conveying the emotions they perceived; they were also asked to rate their level of interest in each excerpt. We predicted that the original compositions would yield different emotional content from their corresponding metrically consonant recompositions (more specific predictions are discussed further below).
It should be noted that our study is somewhat exploratory, employing ecologically valid musical excerpts and a wide variety of emotion terms. Additionally, we use a broad, binary distinction between dissonance and consonance. It is possible that either displacement or grouping dissonance can occur together or alone; indeed, our stimuli, compiled in the appendix (available online in both score and audio formats), contain examples of many kinds of dissonance layering. It may thus be possible to treat metrical dissonance in terms of varying intensity or degrees, but we did not do so here; we will revisit this issue in the Discussion.
Reporting emotions
The current study builds on existing studies on music and perceived emotion in which participants are asked to choose an emotion word that best corresponds to each musical stimulus. The majority of previous studies have employed basic emotions, such as happiness, sadness, fear, and anger, or opposing emotions, such as happy versus sad (see Gabrielsson & Lindström, 2010, for an overview). As theoretical writings suggest, however, the musical emotions conveyed by metrical dissonance may be complex, perhaps incorporating conflicting cues (Krebs, 1999; Rosen, 1995; London, 2004; Hunter, Schellenberg, & Schimmack, 2008). Thus, we decided that limiting emotion words to five or six more general terms was too reductive. The clusters of emotion words first presented by Hevner (1936) and later updated by Schubert (2003) suited our purposes best. To make the Schubert (2003) clusters uniform in size, we eliminated two that contained fewer than four words, and for all remaining clusters, we picked the four words with the highest means. The final clusters used in the present study appear in Figure 2. Future studies based on the same premise may do well to use the simpler lists of emotions, as suggested by Wedin (1972) and Rigg (1937), and used by Lindström (2006), but in this exploratory study we wanted more extensive options for our participants. (In an earlier pilot study, we asked participants to choose individual adjectives from a modified Russell circumplex model (1980; also in Posner, Russell, & Peterson, 2005; and Yik, Russell, & Steiger 2011). Participants reported difficulty navigating the model, as there were too many ungrouped descriptors.)

Final clusters, based on Schubert 2003, formatted as used by participants.
We predicted that participants would select different emotion clusters for metrically consonant and dissonant excerpts. Given the aforementioned associations between complex rhythms and anger (Lindström, 2006), conflict (London, 2004), and anxiety (Rosen, 1995), we predicted further that for metrically dissonant excerpts, participants would choose high-arousal clusters such as F (dramatic, passionate, exciting, triumphant) and G (agitated, angry, restless, tense) most frequently as the best descriptors of emotions perceived. We also asked the participants to select which cluster “worst describes” the excerpts. The main motivation for this relatively idiosyncratic question was to determine whether there are emotional opposites among the clusters, such that they could be arranged around Russell’s (1980) circumplex model—as suggested by Schubert (2003) and reflected in the cluster arrangement of Hevner (1936). This question, then, relates to studies that deal with proposed opposites of individual emotions, such as happiness versus sadness (as in Lindström 2006).
Musical training and emotion perception
Participants in our study were graduate level pianists at the Indiana University Jacobs School of Music (“musicians”), and members of the Indiana University community at large who did not play a musical instrument and had no formal musical training (“non-musicians”). Formal training can provide musicians with a distinct skill set for both creating and perceiving expression or emotion in music (Rowher, 2001; Taruffi, Allen, Downing, & Heaton, 2017). The pianists in our study took private lessons with distinguished pedagogues, who themselves likely used a variety of descriptive language when communicating with students about music. Such language has instructional or analytical implications, and it can also be used for expressive purposes (Karlsson & Juslin, 2008). Formal training may thus provide musicians with a more varied vocabulary with which to describe music they are performing or hearing. Spitzer and Coutinho (2014) provide further evidence that one’s level of expertise in music analysis, music theory, performance, music history, or composition affects one’s perception of emotion in music. The researchers asked participants with low and high levels of expertise to choose an emotion for each movement of Bach’s Violin Sonata, BWV1001. Results showed that high-level experts, who were actively engaged in classical music and were familiar with the style of Bach, based their choices on form, a structural feature. In contrast, low-level experts, undergraduate students who listened primarily to popular music, based their choices on acoustic features, such as pitch, duration, speed, and loudness. While our study did not manipulate formal properties, we predicted that the high-level expert musicians would have different responses than the non-musicians because of their exposure to expressive language in their musical education and because of their relative familiarity with Schumann’s music.
Method
Participants
Two populations from Indiana University participated in this study: musicians and non-musicians. The study received ethical approval from the Indiana University Institutional Review Board; all participants gave their informed consent. “Musician” participants were graduate-level piano performers who had an average of 17.95 years of training on piano (n = 22; average age = 24.5, range = 21–31). “Non-musician” participants were members of the wider Indiana University community who had never received formal music lessons (n = 27; average age = 24.85, range = 19–41). Musicians were recruited through various announcements within the Indiana University Jacobs School of Music. Non-musicians were recruited through an online ad posted on “IU Classifieds.” All participants were paid $10 for completing the study. A post-test question asked participants to name the genres of music they preferred (as many as they wanted). Musicians’ top preferences were classical (n = 21) and pop (n = 5), with one occurrence each of rock, blues, and jazz. Non-musicians’ top preferences were pop (n = 10), rock (n = 9), and rap and R&B (n = 7). Other genres reported once only included jazz, classical, “all,” punk, indie, country, gospel, Egyptian, and Motown. Another post-test question asked participants to name the composer of the excerpts they heard. Among musicians, 19 correctly named Schumann while three named Frédéric Chopin (1810–1849), Schumann’s contemporary. Among non-musicians, composer responses were more varied: Bach (n = 4), Mozart (n = 4), Beethoven (n = 3), Chopin (n = 2), and one response each of Debussy, Tchaikovsky, and Rachmaninov; the remainder of responses were “I don’t know” (n = 6) or uninterpretable. None of the non-musicians correctly named Schumann.
Materials
Stimuli were recorded performances of metrically dissonant excerpts from Robert Schumann’s Carnaval and corresponding consonant recompositions. Ten excerpts were chosen from Carnaval that contained some form of metrical dissonance; these were drawn from several movements: “Préambule” (excerpt Nos. 1–3), “Valse Noble” (No. 4), “Coquette” (No. 5), “ASCH SCHA (Lettres dansantes)” (No. 6), “Chiarina” (No. 7), “Estrella” (No. 8), “Paganini” (No. 9), and “Promenade” (No. 10). Scores for these original excerpts can be found in the online appendix with specific measures and type of dissonance (displacement, grouping) indicated. The 10 excerpts were recomposed using the music notation software Sibelius; these recompositions are also in the appendix, alongside the originals. The only intended manipulation in the recomposed excerpts was changing the metric framework from a dissonant to consonant state. This was carried out by the first author, who reworked harmony, melody, and articulation—including pedaling—so that accents landed on the downbeats of the measures. Displacement dissonances were eliminated by aligning interpretive layers. Grouping dissonances were eliminated by adding or subtracting notes or harmonies so that interpretive layers contained similarly sized pulse units (recall Figure 1). There was one practice trial before the experiment proper. For this, we used an excerpt from Franz Schubert’s Trauerwalzer, D.365. This excerpt did not include metrical dissonance.
A brief discussion of the first original excerpt and its recomposition will further illustrate the two types of metrical dissonance and the approach to eliminating the dissonance (see online appendix). In the original excerpt, a D3+1 displacement dissonance begins in m. 3; the same melody is heard simultaneously in the left and right hands—but offset by a beat. A G3/2 grouping dissonance begins in m. 4 as a result of accents in the left hand: low bass notes recur every two beats rather than every three, creating a hemiola within the underlying 3/4 meter. Several contextual cues contribute to the sense of multiple interpretive layers in this excerpt, namely the placement of slurs, sforzando accents, and registral emphasis in the bass. Other cues are possible, too, including harmonic rhythm, dynamic accents, agogic accents, and textural density. The first author’s consonant recomposition of the excerpt mitigates the metrical dissonance; the melodic motives, the sforzando accents, and the low bass notes now occur on the downbeats of measures. This metrically “consonant” version retains the melody, harmony, and dynamics of the original while altering the rhythm and meter.
All 20 excerpts were performed by a doctoral-level pianist at the Indiana University Jacobs School of Music who had recently performed Carnaval in concert. For each pair of original and recomposed versions, the pianist was instructed to keep tempi and rubato as similar as possible. We also encouraged the pianist to closely follow the accents notated in the scores, that is, the off-beat accents in the original excerpts and the strong-beat accents in the recomposed excerpts. The pianist recorded several takes of each excerpt, and the first and third authors chose the closest matching pairs. The excerpts were performed on a Steinway model A piano within a recording studio on campus. A pair of DPA 4099 condenser microphones were placed inside the piano, and a Neumann TLM 107 microphone was placed a few feet away from the crook of the piano. The audio from the close and mid-distance microphones was mixed with a Yamaha O1V96i console and recorded using ProTools Digital Audio Workstation. The files were trimmed, and level was adjusted for consistency via Wavelab audio editing software. All excerpts were 20–30 seconds long (audio online).
Procedure
Participants were seated in front of Macintosh computers in a quiet room and listened to the stimuli at a comfortable sound level with over-ear headphones. Instructions, questions, and stimuli were displayed using the survey tool Qualtrics, and participants recorded their responses using this tool. In an overview of the study procedure, participants were told that they would do the following tasks in order: a) study several “clusters” of emotions provided on the piece of paper; b) complete a practice trial in which they would listen three times to a musical excerpt (Schubert’s Trauerwalzer) and answer questions about it using the emotion clusters; and finally c) complete 20 test trials in which they would listen three times each to different musical excerpts, answering the same questions as in the practice trial.
Each participant completed all 20 test trials in a unique, randomized order. In each trial, participants listened to the excerpt three times in a row before answering the following questions: a) “Which cluster best captures the emotions in this example?”, b) “Which cluster is the worst at capturing the emotions in this example?”, and c) “How interesting is this example?” For the first two questions, participants responded by selecting a radio button (A through F). For the third question, participants used a sliding scale from 0 (extremely boring) to 100 (extremely interesting). Participants progressed through the experiment at their own pace, using the Qualtrics interface. They were prevented from skipping any questions.
Results
In analyzing our data, we were interested in two broad issues: a) whether metrical dissonance influenced participants’ choice of the emotion clusters that best and worst described the music, as well as their interest in the music; and b) whether musical training played a role in these responses. We conducted three sets of tests to address these issues. First, we focused on ratings of interest. Second, we examined the particular emotion clusters chosen in relation to metrical dissonance and expertise. Finally, we used a generalized estimating equation to test whether specific excerpts or expertise influenced changes in emotion between metrically dissonant and consonant pairs of excerpts.
Ratings of interest
To examine the effect of metrical dissonance and expertise on participants’ ratings of interest, we performed a mixed ANOVA (with the Bonferroni correction) with metrical dissonance as a within-subject factor and expertise as a between-subject factor, ω2 = .44. There was neither a main effect of metrical dissonance, F(1, 929) = 3.25, p = .072, nor a main effect of expertise, F(1, 47) = .604, p = .441. In addition, there was no significant interaction between the two independent variables, F(1, 929) = 1.91, p = .168. This suggests that the two groups of participants did not differ in their average ratings of interest for metrically dissonant (original) and consonant (recomposed) excerpts. In all cases but excerpt 4, musicians rated metrically dissonant versions slightly more interesting than metrically consonant ones. In contrast, non-musicians rated the metrically consonant versions higher than the dissonant versions in half of the excerpts.
Emotions that best and worst described the excerpts
To examine the effects of metrical dissonance and expertise on participants’ choice of emotion clusters that best described and worst described the excerpts, we performed a series of Pearson’s chi-squared tests. We found a significant effect of metrical quality (dissonant or consonant) on the specific emotions used to best describe each excerpt, X2(6) = 14.92, p = .021, φ = .12. There was no significant effect of metrical quality on emotions chosen as worst in describing each excerpt, X2(6) = 4.23, p = .645, φ = .07. There was a significant effect of expertise (musician or non-musician) on the emotional clusters for both best describes, X2(6) = 29.84, p < .001, φ = .20, and worst describes, X2(6) = 38.61, p < .001, φ = .17.
For each of the two groups of participants, we examined the specific emotions chosen to best describe and worst describe metrically dissonant and consonant excerpts over all trials. Figures 3a and b show the proportion of total trials on which each emotion cluster was chosen. As Figure 3a illustrates, musicians chose cluster B (humorous, light, lyrical, playful) most frequently to best describe both metrically dissonant and consonant excerpts. After this, musicians chose cluster F (dramatic, passionate, exciting, triumphant) second to best describe metrically dissonant excerpts; clusters A (bright, happy, cheerful, joyous) and D (dark, melancholy, sad, solemn) were tied for second best for metrically consonant excerpts, though these proportions were close to chance levels (14%). Like musicians, non-musicians chose cluster B most frequently to best describe both metrically dissonant and consonant excerpts and cluster F as the second choice for dissonant excerpts. Unlike musicians, non-musicians also chose cluster F as the second choice for consonant excerpts.

a. Clusters chosen to “best describe” excerpts. b. Clusters chosen to “worst describe” excerpts.
Figure 3b shows the clusters participants chose to worst describe metrically dissonant and consonant excerpts. For both musicians and non-musicians, cluster D (dark, melancholy, sad, solemn) was chosen the greatest proportion of time; the proportion increased when participants heard the consonant recompositions. For both groups of participants, Cluster C (calm, graceful, quiet, soothing) was the second worst choice; the proportion of times this cluster was chosen decreased in the consonant recompositions.
Changes in perceived emotions
Our final set of tests examined whether specific excerpts (1 through 10) and expertise influenced changes in participants’ responses to metrically dissonant and consonant versions of the same excerpt. Recall that participants were presented with stimuli in random order, not by pairs of metrically dissonant and consonant excerpts. They sometimes heard the metrically dissonant version of an excerpt (the original) first and other times the metrically consonant version first. To account for participants’ individual preferences and proclivities that may have influenced them to respond a certain way across multiple musical excerpts, a generalized estimating equation (GEE) was used. The response was a binary indicator of whether each participant changed their response on which emotions best and worst described the excerpt after hearing the second version of the excerpt. Since the response was binary (change/no-change), the GEE implemented a binary logistic regression, using a logit link function and an exchangeable correlation structure to answer the research questions.
We start with the question of whether specific excerpts or expertise can influence changes in the emotion cluster chosen to best describe the music; the effect size of the GEE model was R2 = .15. There was a significant effect of excerpt on participants’ likelihood of changing their opinion, X2(9) = 47.31, p < .001 (see Figure 4a). Many comparisons can be made, but we found the excerpt with the most frequent change to be excerpt 8 and the least frequent change to be excerpt 4. The likelihood of this change may vary depending on whether the participant is a musician or not, as there was a significant interaction between expertise and each excerpt, X2(9) = 17.42, p = .043. Figure 4b compares musicians’ and non-musicians’ proportion of change for each excerpt. Here, we see that excerpt 8 garnered the most amount of change from musicians and the second-most amount of change from non-musicians; excerpt 4 had the least amount of change for musicians, while it was tied (with excerpts 2 and 5) for least amount of change for non-musicians. We found no significant effect of expertise, X2(1) = 1.62, p = .203, but non-musicians were more likely to change their cluster choice when presented with metrically dissonant and consonant versions of the same excerpt (62% of the time) compared to musicians (54% of the time).

a. For the question “best describes,” the proportion of change between metrically dissonant and consonant excerpts for all participants.
For the question of worst describes (R2 = .05), there was no significant effect of excerpt, X2(9) = 14.61, p = .102, meaning that the 10 excerpts elicited similar amounts of change in emotion clusters between metrically dissonant and consonant versions. There was a significant effect of expertise, X2(1) = 4.12, p = .042: musicians changed their choice of emotion cluster between metrically dissonant and consonant excerpts (and vice versa) 43% of the time while non-musicians changed 33% of the time. There was no significant interaction between expertise and excerpt, X2(9) = 3.79, p = .925, indicating that for individual excerpts, both groups of participants changed their responses approximately the same proportion of time.
Discussion
We predicted that metrically dissonant excerpts and their consonant recompositions would be perceived as conveying different emotional qualities; specifically, we predicted that dissonant excerpts would more likely be best associated with high-arousal clusters F (dramatic, passionate, exciting, triumphant) and G (agitated, angry, restless, tense). We also predicted that musical training would influence perceived emotion. Before addressing these predictions, we discuss the benefits and potential pitfalls of using recorded human performances as the experimental stimuli.
We chose to use a recorded human performance in this experiment for several reasons. While we acknowledge that computational models of expressive performance are beginning to allow for precisely manipulated independent variables (Widmer & Goebl, 2004), we believe that the use of a pianist—especially for an exploratory study that emphasizes ecological validity—is better for understanding the nuances of metrical dissonance through expressive timing. In a pilot version of the current study, we used MIDI recordings of all excerpts. During the post-test debriefing, participants, especially musicians, reported that they had difficulty making judgments about emotion because of the quality of sound. Since we wanted to ensure that participants in the main study would be readily able to make decisions about the emotional content of the excerpts, our stimuli were performed by a human pianist, thereby allowing the participants to access expression that only a performer can convey (Juslin & Timmers, 2010). Human performers are intentional in communicating expression through performance (Juslin, 2000; Evans & Schubert, 2008; suggested in London, 2004, pp. 79–80), and this aspect of the musical experience is lost in computer-generated MIDI recordings. Moreover, the concert pianist in this study was intimately familiar with Carnaval, having just performed it in concert recently, and she was able to perform the excerpts stylistically, adding appropriate dynamic contrast, articulation, and notated tempo changes much more fluidly and naturally than any electronic rendition could.
We acknowledge that there are potential confounds in using a human performer to express the experimental metrical manipulations. Much as two recordings of the same composition by the same performer will never be the same, there may have been slight differences between the original and recomposed versions of the same excerpt, besides the intended manipulation. For instance, a compositional shift of notated accents could have resulted in performative changes of dynamics, articulation, microtiming, rubato, and pedaling (Dodson, 2002; Yorgason, 2009). Research has shown that performers emphasize grouping and points of stability through slowing of tempo in performance (Todd, 1985), and expressive timing may be constrained by the meter and the physical movement involved in performance (Schafer, 1984). In the current study, neither the notated meter nor the notated tempo changed between consonant and dissonant excerpts; however, the pianist’s physical movements were necessarily different between these two kinds of excerpt. In recording all excerpts, we worked with the pianist to control for meter, movement, and timing to the extent possible. Indeed, the distinction between metrical consonance and dissonance may have been further exaggerated through human performance, since we encouraged the pianist to emphasize metrical features and differences while mitigating other differences between the consonant and dissonant excerpts. We believe any differences, besides the variable of metrical dissonance, are small enough to be unnoticeable to all but the most discerning of listeners. Music information retrieval techniques may offer other means for validating the dissonance manipulation in the future (see Lartillot et al., 2013).
The concert pianist (initials “CL”) provided us with an additional perspective on the recomposed excerpts in a post-recording questionnaire. She reported that the hardest aspect of learning the recompositions was “unlearning” the “most challenging excerpts from Schumann’s Carnaval.” She stated: “after having spent so many hours making my brain and body adapt to some of the backwards metrical patterns in the original—the various accentuations in the left hand that differ from the right-hand phrase patterns, the pedalings that went against the time signature—… I had to UNLEARN all of that and play the piece ‘normally’!” However, she did note that she also informally “recomposed” passages of Carnaval when learning the piece in the first place—a practice technique that incorporates displacements, specifically. When asked which recomposed excerpts she experienced (performatively) as most similar to the original, CL cited excerpts 2 and 8. The excerpts she reported as most different were excerpts 3 and 6. All four of these excerpts contain a combination of dissonances.
We also asked CL if learning the recompositions changed her performance of the original composition. She said that she had previously manipulated tempo through a liberal use of rubato, which may have hindered the metrical ambiguity in the score itself. CL reported that through the act of recording metrically dissonant excerpts and consonant recompositions, she realized that the original excerpts had a certain “feel” and “emotion” to them that lay in the metrical dissonance; additional performer-introduced tempo changes might in fact “[take] away from the musical message that Schumann worked so hard, and successfully, to express!”
Finally, CL also reflected on the role of emotion in performing the piece overall. It is important to note here that Carnaval is a character piece, which is defined by programmatic (plot-based) ideas—sometimes with emotional connotations—usually expressed in the titles: both of the piece overall, and of the individual movements (Brown, 2001). She noted the varying and playful character of Carnaval; the many characters that play a role within the piece may have an impact on overall emotion. She emphasized the change of emotion that she perceived in the two versions of the movements, “Chiarina” (excerpt 7) and “Estrella” (excerpt 8): Once you hear the difference the small changes make [between the originals and the consonant recompositions], it becomes clear where the magic of Schumann’s melodies lies: In Chiarina, putting the resolution of the melodic anticipation on the second beat took away some of the urgency of the original (which, is to me, much of the emotional part of the character—urgency, intense longing). The same goes for Estrella: the simplified left hand in the middle section makes that melody feel sing-songy and frivolous, rather than the reaching quality that the original has thanks to the echoes between the hands.
In what follows, CL’s responses will be considered alongside a discussion of the experimental results.
Musical training
There was a significant effect of expertise on chosen emotions for all excerpts. This suggests that, in line with our prediction, musicians and non-musicians chose different vocabulary overall to describe the music. We note that this finding could be due to familiarity with the style of music (Spitzer & Coutinho, 2014) or Schumann’s Carnaval in particular, rather than only formal musical training (Karlsson & Juslin, 2008). Most of the musicians correctly identified the composer, and those who did not, identified his contemporary (Chopin). In contrast, none of the non-musicians correctly identified the composer, though a few named contemporaries (Chopin, Beethoven, Tchaikovsky). Further research would be needed to tease apart a possible distinction between training and familiarity.
Non-musicians were more likely to change their cluster choice overall between metrically dissonant and consonant excerpt pairs, as can be seen in Figure 4b. Familiarity may have played a role: the non-musicians may have assumed all 20 excerpts were different compositions, prompting greater differences in response across the stimulus set, while the musicians may have recognized similarities between the pairs (even though all participants heard the excerpts in randomized order). Musicians and non-musicians also differed in the excerpt pairs that led to changes in emotional clusters. Musicians had a higher proportion of cluster change when they heard metrically dissonant and consonant versions of excerpts 1 and 2. Non-musicians had a higher proportion of cluster change for excerpts 3, 4, 5, 6, 9, and 10. There was a high proportion of cluster change by both groups for excerpts 7 and 8; recall that CL also singled out these specific excerpts for the emotional differences between each of their two versions. The largest differences between musicians and non-musicians were found in excerpts 2 (musicians changed more often) and 9 (non-musicians changed more often). Original excerpt 2 has both displacement and grouping dissonances that mask the underlying meter entirely. Musicians, due to their deeper understanding of the musical structure, may have been more attentive to these relatively complex dissonances (Spitzer & Coutinho, 2014). Original excerpt 9 only has a displacement dissonance, a relatively straightforward and fast-moving D2+1. While this may have been a salient feature for non-musicians, it may have been less notable to musicians.
Notably, the pairs of excerpts that CL performatively experienced as most similar and most different (not relating to emotion) diverged from those that had the highest and lowest proportion of change in musician-selected emotions: CL reported that the original and recomposed versions of excerpts 3 and 6 were most different and those of excerpt 2 were most similar. CL also experienced the two versions of excerpt 8 as most similar in performance, in contrast to the high proportion of emotion cluster change for both groups of participants, and in contrast to her own emotional assessment. These responses suggest that future studies with trained musicians may do well to attempt separating the perception of performance effort from the perception of perceived emotion.
Perception of emotion
We predicted that participants would perceive different emotions in metrically dissonant and consonant excerpts. We further predicted that metrically dissonant excerpts would be best described by high-arousal clusters F (dramatic, passionate, exciting, triumphant) and G (agitated, angry, restless, tense). Our findings offer partial support for these hypotheses. For several excerpts, consonant recomposition caused notable changes in cluster choice (see Figures 4a and 4b); we examine one such excerpt further below. And for musicians, the second-place cluster for metrically dissonant excerpts overall was F, while the second-place cluster for metrically consonant excerpts overall was A (bright, cheerful, happy, joyous). Yet across all participants, cluster B (humorous, light, lyrical, playful) was the most frequent choice for both dissonant and consonant excerpts. The results suggest that the effect of metrical dissonance was not strong enough to significantly influence emotional expression, and other factors, such as mode and tempo, may have had a stronger influence on the perception of emotion. Additionally, the strength of cluster B (humorous, light, lyrical, playful) could be related to the character of the piece itself, since, as mentioned above, Carnaval denotes many characters and emotions in its varied movements, and this relation may have taken away from any other emotional connotation overall.
We note that, based on the design of the experiment, we cannot determine whether recompositions reliably led to specific changes in emotional clusters. We cannot say, for instance, that recomposition significantly shifted participants’ choice of cluster A to cluster B. Nevertheless, we can compare Figures 3a and 3b and point to general trends. Clusters B (humorous, light, lyrical, playful) and F (dramatic, passionate, exciting, triumphant), which have high valence and moderate to high arousal, were chosen on the greatest proportion of trials as best describing metrically dissonant excerpts (Figure 3a), yet they were chosen relatively infrequently, particularly by musicians, for “worst describes” (Figure 3b). Conversely, clusters C (calm, graceful, quiet, soothing) and D (dark, melancholy, sad, solemn), which have low–moderate valence and low arousal, were chosen a high proportion of times for “worst describes” (Figure 3b) yet chosen relatively infrequently for “best describes” (Figure 3a). This suggests that in general, participants consistently treated some clusters as opposite, after Russell (1980), even though the layout of clusters that participants viewed (see Figure 2) did not deliberately encourage such thinking. Vuoskoski and Eerola (2011) also found that different emotion measures generally can be accounted for in the valence and arousal model of Russell (1980).
Case studies: Excerpts 4 and 8
Due to the significant effect of excerpt on cluster change, we further examined the response regarding the metrically dissonant and consonant versions of excerpts with the least difference in cluster choice (excerpt 4) and the most difference in cluster choice (excerpt 8). The graph in Figure 5a shows the clusters chosen across all participants to best describe excerpt 4. Cluster D (dark, melancholy, sad, solemn) and cluster C (calm, graceful, quiet, soothing) were the most popular choices for both dissonant and consonant versions. Examining the music (found in the online appendix), we note that the only metrical dissonance in the original is a displacement dissonance [D3+1] in the left hand: the low bass note is displaced from the downbeat to beat 2. The dissonance is subtle; indeed, the continuous eighth-note pattern and the low register might make it hard to perceive.

a. Emotion clusters chosen for excerpt 4.
In contrast, excerpt 8 had the greatest proportion of cluster change between metrically dissonant and consonant versions. The graph in Figure 5b shows the clusters chosen across all participants to best describe excerpt 8. In general, participants perceived the dissonant version as more negatively valenced than the consonant version. Cluster G (agitated, angry, restless, tense) was chosen most frequently for the dissonant version while cluster B (humorous, light, lyrical, playful) was chosen most frequently for the consonant version. And whereas no participants chose cluster A for the dissonant version (bright, cheerful, happy, joyous), several participants did so for the consonant recomposition. Examining the music, the dissonant version has two levels of displacement (D3+1 and D3+2, written in the example as D3+1+1). Additionally, these dissonances occur prominently in not only the bass voice but also in the upper lines (soprano and tenor ranges). In the consonant recomposition, the displaced melodies are placed on the downbeat, while the upbeat harmony is retained to have some connection to the original. Metrical manipulation had a greater effect on the emotional connotations of this excerpt than in excerpt 4, perhaps due to the multiple layers of dissonance across multiple parts and their more noticeable mitigation in the consonant version.
A comparison of these two excerpts suggests that there are degrees of metrical dissonance in Carnaval. While both excerpts contain displacement dissonance, factors such as register, texture, and number of interpretive layers make the dissonance more or less immediately perceptible and thus the recomposition more or less perceptibly different from the original. As noted earlier, we did not rank our excerpts in terms of degrees, or intensities, of dissonance; we treated them strictly in a binary fashion, dissonant or consonant. Krebs does allude to intensities and levels of dissonance; however, he does not give a systematic formulation of how to measure metrical dissonance in terms of these intensities (Krebs, 1999, pp. 56–57). Additionally, we did not consider other categories of metrical dissonance proposed by Krebs (1999). Different categories and intensities could possibly contribute to the creation of a continuum of metrical dissonance to consonance. Existing compositions—or alternately, newly composed pieces—could be situated along this continuum. This would require systematic consideration of musical parameters that work to amplify or mitigate metrical dissonance, such as accent, range, texture, tempo, and register.
Finally, the presence of multiple interpretive layers in Schumann’s music, each contributing to and complicating the music as a whole, suggests that metrical dissonance might productively be considered in relation to “complexity,” which Lindström (2006) defines as “music with many complicated changes [that] contains much information,” in comparison to repetition and redundancy (p. 90). In Lindström (2006), participants were asked to rate simple melodies (variations of “Frère Jacques”) on two types of bipolar scales: those relating to perceptions of musical structure (stability/instability, simplicity/complexity, relaxation/tension) and those pertaining to perceived emotion (happy/sad, angry/tender, expressive/expressionless). Such an approach should be seriously considered for future research on the perception of metrical dissonance. That is, it might be useful to have a response measure that captures how the perceived complexity created by metrical dissonance (along with other factors such as harmony, rhythm, mode, and tempo) relates to perceived emotion.
Summary
The current study demonstrates, for the first time, how metrical character can influence perceived emotions. Though this study concentrates on a single musical work, Robert Schumann’s Carnaval, it has implications for other music with metrical dissonance, a compositional feature found in many different musical eras and genres. Our results support the conclusion that emotional connotations vary depending on whether listeners respond to metrically dissonant excerpts from the original piece or metrically consonant recompositions of those excerpts, and depending on whether listeners are skilled pianists or have no formal musical training. Notably, differences in perceived emotion between metrically dissonant and consonant versions of the same excerpt may vary according to excerpt. This suggests the presence of degrees of metrical dissonance. A systematic examination of the specific factors that contribute to these degrees—and their emotional implications—would be a fruitful area of future research.
Supplemental Material
MSX836708_supp_mat_1 – Supplemental material for Effects of metrical dissonance and expertise on perceived emotion in Schumann’s Carnaval
Supplemental material, MSX836708_supp_mat_1 for Effects of metrical dissonance and expertise on perceived emotion in Schumann’s Carnaval by Jessica Sommer, Kimberly Simmons and Daphne Tan in Musicae Scientiae
Supplemental Material
MSX836708_supp_mat_2 – Supplemental material for Effects of metrical dissonance and expertise on perceived emotion in Schumann’s Carnaval
Supplemental material, MSX836708_supp_mat_2 for Effects of metrical dissonance and expertise on perceived emotion in Schumann’s Carnaval by Jessica Sommer, Kimberly Simmons and Daphne Tan in Musicae Scientiae
Footnotes
Acknowledgements
The authors would like to thank Clare Longendyke, Anthony Tadey, and Sebastiano Bisciglia for their assistance in conducting the study and Michael Frisby and Peter Miksza for their assistance in data analysis.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
