Abstract
Philosophers have long wrestled with the apparent paradox of the enjoyment of negative emotional portrayals in the arts. An example of this apparent paradox is the enjoyment among some listeners of nominally sad music. An experiment is reported in which 39 participants listened to sad and happy music while serum prolactin (PRL) concentrations were measured. The purpose of the experiment was to test an a priori theory, proposed by Huron, that liking sad music is mediated by elevated PRL levels. Contrary to the theory, sad music did not result in a significant increase in PRL; nor was the pleasure of listening to sad music associated with increased PRL. Nominally happy music did result in a decrease of PRL, especially for those participants who most prefer happy music over sad music. The effect was greatest for those who score high on a measure of loneliness. Consistent with other studies, the degree of liking sad music over happy music was found to correlate with trait openness to experience, although this effect was not echoed in PRL levels. Post-hoc analyses indicate that PRL decreases were most marked for male listeners and those who score high on a loneliness measure. In general, the results are not consistent with the theory proposed by Huron.
Why do people enjoy listening to sad music? Sadness is generally regarded as a negatively valenced emotion, so people normally endeavor to avoid feeling sad. As philosopher Stephen Davies has pointedly asked, if sad music makes people feel sad, why would they bother to listen? (Davies, 1997, p. 242). This question might be viewed as a specific instance of a general question posed more than two millennia ago by Aristotle, who pondered how people could enjoy watching Greek tragedy with its portrayals of horror and terror (Aristotle, 1922). Philosopher Carole Talon-Hugon (2014) has summarized the paradox via the following self-contradictory syllogism, expressed as three statements:
Negative emotions are psychologically unpleasant.
Some art causes us to experience negative emotions.
We find pleasure in this experience of negative emotions.
At least one of these three must be wrong. For example, some emotions we regard as “negative” affects are not actually unpleasant (challenging the first statement). Or it could be that the emotions evoked by art are not the negative ones we presume (challenging the second). Another possibility is that the pleasure we experience comes from some other source and is not related to the emotions represented or evoked by the art (challenging the third). In the case of sad music, a prominent idea has been to claim that nominally sad music does not actually evoke feelings of sadness and that listeners are feeling something else instead (e.g., Davies, 1997; Kivy, 1980; Levinson, 1997).
Aesthetic philosophers have offered a variety of theories in an effort to explain this apparently paradoxical situation (e.g., compilation by Levinson, 2014). However, few of the extant theories have been tied to proposed biopsychological mechanisms. In this study we report a test of one such theory—the “prolactin theory of sad music” proposed by Huron (2011). Since publication, Huron’s paper has been widely cited, yet no empirical studies have explicitly tested the proposed theory.
Empirical studies of the enjoyment of sad music point to high individual variation. Some people enjoy listening to nominally sad music whereas others exhibit a notable dislike of sad music. Taruffi and Koelsch (2014), for example, conducted a large survey which included the question: “How often do you actively select sad music to listen to?” They found that roughly 20% responded “never” or near never (i.e., responses 1 or 2 on a 7-point scale from “never” to “always”), whereas roughly 10% responded “always” or near always (i.e., responses 6 or 7 on the same scale).
One possible factor influencing enjoyment of nominally sad music is personality: in particular, openness to experience (Cotter, Silvia & Fayn, 2018; Colver & El-Alayli, 2016; Nusbaum & Silvia, 2011; Vuoskoski, Thompson, McIllwain, & Eerola, 2012). Vuoskoski et al. (2012) found that empathy and openness to experience are significantly associated with liking for sad music and with the (self-reported) intensity of response when listening to sad music. Garrido and Schubert (2011) found that absorption and music empathizing are associated with the enjoyment of sad music, proposing that this association arises because sad music creates a dissociation from pain or displeasure.
Despite the wide variation in liking for nominally sad music, sad-music likers/dislikers do not appear to constitute distinct classes of listeners. In the Taruffi and Koelsch (2014) study, for example, self-reported frequency of listening to sad music exhibits a unimodal rather than a bimodal distribution. Most of the philosophical theories proposing to explain the paradox of sad music listening do not posit a range of listening behaviors. That is, most of the extant theories are formulated as an explanation for all or the majority of listeners. However, the empirical research suggests that any theory purporting to explain the enjoyment of nominally sad music must account for such individual differences.
A number of studies have examined the effect of music on various hormones. Pertinent reviews are found in the work of Gangrade (2012), Chanda and Levitin (2013), and Fancourt, Ockelford, and Belai (2014). Endocrine music studies have typically focused on cortisol (CS) (e.g., VanderArk & Ely, 1992), testosterone (e.g., Fukui, 2001), human growth hormone (HGH), oxytocin (OT), and adrenocorticotropic hormone (ACTH). The few extant studies examining the effect of music on prolactin (PRL) have produced complex results. Möckel et al. (1994) observed a decrease of PRL after exposure to relatively stressful modern music with irregular rhythms. Comparing classical music with more energetic techno-music, Gerra et al. (1998) found no significant changes in PRL but observed an increase in cortisol levels in the techno-music condition. Evers and Suhr (2000) studied the influence of pleasant and unpleasant music on ACTH, PRL, and serotonin. They found no significant changes in either PRL or ACTH.
Like CS, PRL is implicated as a stress hormone. Increases in serum PRL have been observed in response to stressful circumstances (Armario, Marti, Molina, de Pablo, & Valdes, 1996; Corenblum & Taylor, 1981; Meyerhoff, Oleshansky, & Mougey, 1988). Frey (1985) reported high levels of PRL in psychic tears (caused by emotional states) but not in irritant tears (caused by physiologically painful states). Studies such as those by Dinan, Scott, Thakore, Naesdal, and Keeling (2001) have found a close relationship between CS and PRL.
By contrast, other research has found a PRL decrease after a stressful task (Gerra et al., 2001; Mutti et al., 1989). Moreover, several studies have observed no changes in PRL concentration in response to stress (Gerra et al., 1996; Semple, Gray, Borland, Espie, & Beastall, 1988). Theorell (1992) found that serum PRL reacts in different ways connected with the stress coping mechanism: in subjects who experience passive helplessness, PRL concentration increases, whereas in subjects with increased anxiety and active coping, PRL concentration decreases. Sobrinnho et al. (2003) suggest that there is no general stress response, and that CS and PRL may be mutually exclusive responses to different emotions.
Other research suggests that CS and PRL are not exclusively associated with negatively valenced experiences, but with positively valenced high-arousal experiences as well. Codispoti et al. (2003) systematically examined if there are differences in the neuroendocrine responses for pleasant and unpleasant stimuli when both have a very high arousal level. For unpleasant stimuli, CS concentrations increase whereas PRL concentrations decrease. For pleasant stimuli, in turn, PRL concentrations increase. The reasons for the decrease in PRL levels after unpleasant and stressful emotions are not entirely clear. The authors speculate about the relationship between stress and neurochemical activity of dopaminergic neurons and dopamine release in humans (Finlay & Zigmond, 1997; Pani, Porcella, & Gessa, 2000). Noradrenaline may increase dopamine, which could be responsible for prolactin suppression (Arbogast & Voogt, 1996; Franklin et al.,1999; Kojima, Arita, Kuwana, & Kimura, 1995).
Changes in PRL concentrations are also associated with feelings of consolation or comfort. For example, PRL has been observed to be released following sex, and the amount of PRL released is correlated with judgments of sexual satisfaction and relaxation (Brody & Krüger, 2006). Anecdotally, many people commonly report that crying makes them feel better (Bylsma, Vingerhoets, & Rottenberg, 2008; Vingerhoets & Bylsma, 2016). Taruffi and Koelsch (2014) note that many listeners report using sad music to facilitate “venting” of negative emotion or mood. Since crying is linked with the release of PRL (Frey, 1985), it is possible that PRL may have a causal role in the reported positive feelings.
In light of these and other observations, Huron (2011) conjectured that the comforting feelings associated with PRL might have a homeostatic function. That is, when a sad or grief state occurs, PRL may act to inhibit or limit the physiological and psychological effects. In cases of music or dramatic portrayals, for example, a sad-like or grief-like state may be artificially induced even though there is no objective reason to experience psychic pain. Accordingly, Huron (2011) proposed that the psychological effect of PRL might explain why some listeners experience sad music as enjoyable. Such a state could be experienced as phenomenologically pleasant in an analogous way to the administration of exogenous opiates when no physical pain is experienced. If correct, the theory predicts that, for individuals who report sad feelings due to music, those who enjoy the experience will tend to exhibit higher levels of PRL than those who do not enjoy the experience. In this study, we test this theory. In brief, the experiment involves measuring serum prolactin levels in response to happy and sad music for participants who tend towards the two poles along the continuum from sad-music lovers to sad-music haters.
Hypotheses
We state our hypotheses as follows:
Methods
In brief, the experiment involved exposing listeners to sad and happy music conditions while making baseline and treatment blood draws in order to trace potential changes in serum PRL. Two controls are appropriate in this study. The first contrasts musical stimuli with silent periods. This allows us to compare PRL levels arising from musical exposure to baseline PRL levels in the absence of music. A second control contrasts sad music with non-sad music. That is, we need to ensure that any observed change in PRL concentration is specifically related to sad music, and not merely a response to music in general.
Participants
In order to increase the possible effect size, we recruited participants who reported they either really enjoyed or really disliked listening to sad music. All participants were recruited through a three-stage process and provided informed, written consent before participating. First, potential participants were solicited through ResearchMatch, a national electronic web-based recruitment system maintained by Vanderbilt University. Volunteers register with ResearchMatch, providing basic demographic and other information; e-mail solicitations are then broadcast to those registered participants whose demographic information fits the eligibility criteria of the researcher. To be eligible for the current experiment, participants had to (a) be 18 years of age or older, (b) reside within a reasonable distance of the Ohio State University Medical Center, (c) be in good health and not pregnant or breast feeding, and (d) be someone who feels truly sad when listening to sad music and who either really enjoys or really dislikes sad music. Potential participants received an initial solicitation email, and ResearchMatch Volunteers who responded to this first email received a more detailed recruitment message (see Appendix A in Supplemental Material).
For those prospective participants who responded positively to the second detailed recruitment email, we arranged a telephone interview. The goal of the interview was to ensure that they complied with the selection criteria. We employed a semi-structured interview format. The interview involved a set sequence of questions (Appendix A in Supplemental Material).
Using this recruitment method, 88 prospective participants responded in an email exchange, 78 provided telephone contact information, and 50 participated in a telephone interview. Of the interviewees, we selected 40 participants, emphasizing the selection of those who most identified themselves as either liking or loving sad music or as disliking or hating sad music. One prospective participant dropped out so, in the end, 39 volunteers (25 females) took part in the study. The research protocol was reviewed by the Biomedical Sciences Institutional Review Board of the Ohio State University and received approval number 2010H0335.
Musical stimuli
Two types of musical stimuli were employed—nominally “sad” and “non-sad” music. In the case of non-sad music, one might simply select music representing or exhibiting a comparatively neutral or flat affect (e.g., “elevator music”). However, if “unemotional” neutral music of this kind were to be used as a control, any observed PRL increase in the case of the sad music would also be consistent with the hypothesis that PRL is released by listening to “emotional” music, whether sad or otherwise. Accordingly, we elected to use nominally “happy” music as the musical control.
The selection of stimulus materials aimed to maximize the sad and happy feelings experienced by each participant. Research in music and emotion has established that self-selected music is typically more likely to evoke a target affect than music selected by experimenters or others (e.g., Blood & Zatorre, 2001). A disadvantage of using participant-selected music is that it can introduce unknown factors, such as familiarity effects or personal associations. Despite these disadvantages, in order to maximize a possible effect size, we nevertheless chose a method in which the stimuli differed from participant to participant.
Participants could choose any music at all for their mood-induction program. This included recordings from their own collections or playlists of works deemed “happy” or “sad” by the participant. In order to assist participants in finding music that is especially effective for them, the experimenters also provided participants with two CDs containing a variety of works deemed by others (the experimenters and research colleagues) to be especially sad or happy. Having selected their sad and happy works, participants then arranged the works into 20-minute sad and happy programs. In creating their programs, participants could elect either to arrange a particular sequence of works or to repeat a single piece multiple times.
Procedure
Participants took part in two sessions, involving exposure to either their happy or sad music programs. The order of the sessions was randomized for each participant. Table 1 summarizes the time-line for a single session. For those sessions involving sad music, participants were offered the opportunity to undergo the happy music mood induction program before leaving, in order to induce a more positive mood.
Procedure timeline.
Note. ULS-8 = short version of the UCLA Loneliness Scale; REMAP = Resilience Resources Measure for Prediction and Management of Somatic Symptoms.
Participants were instructed not to smoke or consume caffeine on the morning prior to the experiment. The two experimental sessions were arranged on different days. Due to diurnal changes of PRL levels, the two experimental sessions for each subject were scheduled within a clock-hour of each other (Djursing, Hagen, Møller, & Christiansen, 1981; Sassin, Frantz, Weitzman, & Kapen, 1972). Thus, a participant might experience session #1 on Monday at 10 a.m. and session #2 on Thursday at 9 a.m. The sessions were scheduled between 1 and 7 days apart.
Each session was roughly 1 hour in duration and was conducted in a quiet, private hospital room. After a consent briefing, participants sat in an upholstered phlebotomy chair tilted to a position deemed comfortable by the participant. Venipuncture and catheterization were performed by a registered nurse. Subjects then completed a 44-item version of the Big Five personality inventory (BFI) (John & Srivastava, 1999), followed by a short version of the UCLA Loneliness Scale (ULS-8) (Hays & DiMatteo, 1987), and the Resilience Resources Measure for Prediction and Management of Somatic Symptoms (REMAP) (Malarkey et al., 2016). The UCLA Loneliness Scale is an 8-item questionnaire on loneliness. The REMAP questionnaire is a 25-item resilience instrument which provides six individual scores representing an overall score, and scores for five individual factors (relational engagement, emotional sensibility, meaning engagement, awareness of self and others, and physical health behaviors). Finally, we asked participants to indicate their general preferences regarding nominally happy and sad music. These self-report trait assessments were in response to two questions: “How much do you like sad music?” and “How much do you like happy music?” Responses were given on a 7-point scale with three anchor points labeled “hate it” (−3), “neutral” (0), and “love it” (+3).
The experimenter engaged the participant in “small talk” for roughly 20 min, to encourage the participant to relax and feel comfortable with the catheter in place. In order to reduce the possibility that the conversation topic might unduly influence the participant’s mood, the participant was left alone for five minutes before the first mood and pleasantness assessments were done. Participants received the following instructions: I’m going to leave you alone now for about 5 minutes, then I’ll come back, ask you a couple of questions, and then we’ll start the music. For the next 5 minutes, I’ll leave you alone with your thoughts. Try to avoid thinking about anything emotional—either especially happy or sad thoughts. Do you have any questions?
The mood assessment consisted of a single question with the sad/happy versions asked at appropriate sessions: On a scale of 1 to 100 where 1 is not at all [sad/happy] and 100 is the [saddest/happiest] you’ve ever felt, how [sad/happy] would you say you feel right now at this moment?
Similarly, the pleasure assessment consisted of a single question: On a scale of 1 to 100 where 1 is no pleasure at all and 100 is the most pleasure you’ve ever felt, how much pleasure would you say you feel right now at this moment?
Immediately following the mood and pleasure assessment, the first blood draw occurred. Notice that this occurred approximately 30 min after catheterization, so the stress of catheterization would have been minimal at this point.
Sad and happy moods were induced through a multi-pronged procedure involving a combination of guided imagery vignettes, and music presented via headphones. The mood induction procedure (both happy and sad) followed a method used by Mayer, Allen, and Beauregard (1995). In both conditions, the first four minutes consisted of music program alone. Then the volume of the music was lowered to a background level and the participant read a series of statements (guided imagery vignettes) written on small cards, modified to use first person rather than second person (see Appendix B in Supplemental Material). The participant was previously instructed to imagine as vividly as possible each scene or scenario described in the vignette. The eight statements were read slowly over a period of about 4 minutes. After this, the music was returned to full volume. The music then continued for a further 12 min or so, for a total mood induction period lasting 20 min. Mayer et al. (1995) have shown that guided imagery vignettes and music are additive in their mood-induction effects; the combination is better able to induce a given mood than either the vignettes or the music alone.
After the personality inventory and mood and pleasure assessments, but prior to the mood-induction music program, participants were instructed as follows: Although the situation here is rather artificial, please do your best to get into a [sad/happy] mood. We’re going to give you as much privacy as possible. The nurse will not be in the room during the music program. I will be behind this curtain, in case you need something or have any questions. The nurse will draw a blood sample right now, and then we’ll start the music. First you will hear 4 minutes of music to help you get into the appropriate mood. Then we will fade out the music for a few seconds and fade it back in, to give you a signal that you should start reading this series of statements. They are numbered, please read them in consecutive order. Read one statement at a time and try to relate the statements as much as you can to your own life. Take your time. After each statement, pause for about 30 seconds, and imagine as vividly as possible the scene or scenario described. Then move on to the next statement. After the final statement, please answer the two questions about your mood and your pleasure again that you filled out just before. After you are done with that, the music will continue for another 10 to12 minutes. Try as best as you can to stay in the mood that the statements and the music induced. You might find it useful to close your eyes, but that’s not necessary. Do whatever you think will help you get into the right mood and stay in that mood throughout the procedure. [For the sad music procedure: “Ideally, it would be great if you are able to cry.”] After the music stops the nurse will enter again and take a final blood sample and remove the catheter from your forearm. Do you have any questions?
The music program continued after the second mood and pleasure assessments, for a total exposure of 20 min. This was then followed by the second blood draw. Finally, a third mood and pleasure assessment was made.
Assays
For each participant three mL blood samples were collected into red-top tubes (no anticoagulant). Samples were allowed to sit at room temperature until clotting was complete. Serum was harvested from each sample following centrifugation and placed in two labeled screw-top plastic vials, the first receiving 0.5 mL serum and the second receiving the remaining serum.
Diagnostic procedures were performed by certified clinical laboratory personnel. Serum concentrations of PRL were measured using a DPC Immulite 1000 immunoassay instrument (Diagnostic Products Division of Siemens Healthcare Diagnostics, Deerfield, IL). This instrument uses a photon-counting immunoassay methodology to measure the concentrations of various serum constituents. The 0.5 mL samples were thawed at room temperature and mixed to insure homogeneity. Hormone concentrations of the samples were measured twice. If results varied by more than 10%, then they were reevaluated.
Results
In analyzing our results, it is useful to note that we gathered three mood/pleasure assessments, and two PRL measures. Mood/pleasure assessments were carried out prior to the music program, during the music program, and immediately after the music program. In the end, we ignored data for the second mood/pleasure assessment and focused exclusively on the assessments prior to and immediately following the music program. For convenience, we will refer to the measure prior to and the measure immediately following the music program as T1 and T2, respectively.
In the first instance, we conduct a manipulation check and ask whether the music programs induced the intended change of mood (see Figure 1(a)). Since Shapiro-Wilk tests found that the difference scores between mood ratings provided at T1 and T2 were normally distributed only for sad music, W(39) = 0.970, p = .369, but not for happy music, W(39) = 0.928, p = .016, non-parametric Wilcoxon signed ranks tests were used in both cases for consistency and to ease comparison across related measures. For the sad music condition, sad mood increased significantly from T1 (Median = 10.0, interquartile range (IQR) = 19.0) to T2 (Median =55.0, IQR = 36.0), Z = −5.2, p < .001, r = −.590.Similarly, happy mood ratings increased significantly in the happy music condition from T1 (Median = 70.0, IQR = 15.0) to T2 (Median = 80.0, IQR = 25.0), Z = −4.4, p <. 001, r = −.504. Moreover, in post-experiment debriefings, 10 participants reported actual crying or tearing episodes during the sad music condition. During the happy music condition, many participants reported tapping their feet or humming along to the music.

Plots of means as well as individual data points for (a) mood, (b) serum prolactin (PRL) concentrations, and (c) self-reported pleasure ratings for the sad and happy music conditions. T1 identifies baseline measures prior to exposure to the sad or happy music programs; T2 identifies measures following 20 min of music exposure. The sad music program resulted in marked increased reported sadness, and a marked decrease in pleasure. The happy music program resulted in significantly happier mood, significantly increased reported pleasure, and significantly reduced PRL compared with an elevated baseline PRL. In this figure, results have been pooled for both happy-music likers and sad-music likers. Error bars represent confidence intervals computed using Cousineau’s (2005) method for within-participant designs. Statistical significance is inferred from non-parametric Wilcoxon signed-rank tests with *** designating p < .001. In Panel (b), serum PRL was log-transformed merely for visualization purposes.
Having established that listener moods changed in response to the sad and happy music programs, we then examined whether the PRL levels also changed as a result of the manipulation (Figure 1(b)). Once again, nonparametric tests were appropriate since the data were not normally distributed for difference scores between PRL measurements provided at T1 and T2 for the happy, W(39) = 0.872, p < .001, and sad, W(39) = 0.869, p < .001, music conditions. Contrary to the motivating hypothesis, the sad music program did not result in a significant increase in serum PRL level from T1 (Median = 8.4, IQR = 5.7) to T2 (Median = 8.5, IQR = 5.5), Z = −0.483, p = .636, r = −.055. However, the happy music program did result in a significant decrease of serum PRL concentration from T1 (Median = 10.3, IQR = 7.0) to T2 (Median = 8.4, IQR = 5.6), Z = −3.79, p < .001, r = -.429. Curiously, baseline PRL levels were significantly higher prior to the happy music condition than they were prior to the sad music condition, Z = -2.927, p = .003, r = -.331.
Finally, we test directly the principal motivating hypotheses related to the prolactin theory (i.e., H1 and H2). Hypothesis 1 predicted that for those listeners who report increased pleasure when listening to sad music, sad music would cause an increase in serum PRL concentrations; conversely, those listeners who report no increased pleasure from sad music would show little or no increase in PRL when listening to sad music. As reported above, changes in PRL in response to happy and sad music were both non-normally distributed. Changes in pleasure ratings were also non-normally distributed for happy music, W(39) = 0.936, p = .027, whereas changes in pleasure ratings for sad music did not significantly violate the normality assumption, W(39) = 0.958, p = .152. Accordingly, non-parametric test procedures were employed. Spearman’s rank-order correlations between change in PRL and change in pleasure ratings were non-significant in both the happy (rs = .094, p = .570) and sad music conditions (rs = .182, p = .268). In short, the results are not consistent with the main motivating hypothesis.
A second operationalization of sad/happy preferences and PRL
A second approach to testing the “prolactin theory of sad music” considers whether those participants who generally prefer mostly sad music are more likely to exhibit serum changes of PRL (i.e., H2). Recall that participants provided self-report (trait) values regarding their preferences for happy and sad music. Separate 7-point ratings were given for “how much do you like sad music” and “how much do you like happy music” (both ranging from −3 to +3). From these data, it is possible to characterize individual listeners as either predominantly either a happy-music-liker or a sad-music-liker by calculating the difference between happy- and sad-liking ratings. Extreme positive and negative values are indicative of a listener who prefers one kind of music much more than the other. Table 2 shows the number of participants who exhibited various difference scores, segregated by sex. Plus (+) signs indicate difference scores favoring happy music over sad music, whereas minus (−) signs indicate difference scores favoring sad music over happy music.
Number of participants and degree of preferring happy(+) or sad(−) music.
Since the difference scores between PRL measurements provided at T1 and T2 were not normally distributed for happy-music-likers listening to sad music, W(39) = 0.778, p < .001, and for sad-music-likers listening to happy music, W(39) = 0.820, p = .047, it was inappropriate to use an ANOVA. Instead, for consistency and to facilitate direct comparisons across measures, we conducted four non-parametric Wilcoxon signed-rank tests with a corrected alpha level of .0125. With regard to those participants deemed “happy-music-likers” (n = 20), for the sad music program, no significant difference was evident between PRL concentrations at T1 (Median = 8.15, IQR = 5.4) and T2 (Median = 8.15, IQR = 5.7), Z = −1.088, p = .289, r = −.172. For the happy music program, we observed a significant decrease in PRL concentrations from T1 (Median = 9.60 IQR = 4.6) to T2 (Median = 7.80, IQR = 5.0), Z = −3.174, p = .001, r = -−.502. With regard to those participants deemed “sad-music-likers” (n = 8), for the sad music program, no significant difference was evident between PRL concentrations at T1 (Median = 9.20, IQR = 10.3) and T2 (Median = 10.25, IQR = 8.3), Z = −1.101, p = .328, r = −.275. For the happy music program, no significant difference was evident between PRL concentrations at T1 (Median = 6.35, IQR = 16.8) to T2 (Median = 6.00, IQR = 18.6), Z = −0.351, p = .758, r = −.088. By way of summary, those participants who most preferred happy music over sad music exhibited a significant reduction in serum PRL when listening to happy music (although see below). Such a decrease was not observed for the sad-music-liker group.
A third operationalization of sad/happy preferences and PRL
Finally, one could argue that characterizing listeners as predominantly sad- or happy-music-likers by subtracting one set of ratings from the other is misguided. One can imagine a nominally sad-music-liker who also enjoys listening to happy music. We thus considered only the sad-music preference assessment asking if those listeners who self-reported a high preference for sad music (whatever their enjoyment of happy music) exhibited elevated levels of PRL in the sad music condition. For this test, we split participants into sad-likers (n = 31) and sad-dislikers (n = 7) according to whether the sad music self-reported preference was positive or negative (one participant was excluded due to scoring zero on preference for sad music). Shapiro-Wilks tests showed that difference scores for PRL measured at T1 and T2 were not normally distributed for sad-likers listening to sad music, W(31) = 0.831, p < .001, or happy music, W(31) = 0.864, p = .001. For sad-dislikers, however, PRL difference scores did not violate the normality assumption in the sad music, W(7) = 0.870, p = .185, and happy music conditions, W(7) = 0.953, p = .760. Thus, non-parametric tests were employed throughout for consistency and to ease comparison across measures. However, to check whether non-significant results could be attributed to low statistical power due to the use of non-parametric test procedures, further parametric tests were conducted for the sad-disliker data. Tests were adjusted for multiple comparisons using a Bonferroni-corrected alpha level of .0125.
For sad-likers, a Wilcoxon signed-rank tests showed that, for the happy music condition, PRL decreased significantly from T1 (Median = 10.8, IQR = 8.7) to T2 (Median = 9.0, IQR = 7.8), Z = −2.97, p = .002, r = −.377. In the sad music condition, no significant change in PRL concentration was found from T1 (Median = 9.6, IQR = 5.9) toT2 (Median = 9.1, IQR = 6.2), Z = −0.32, p = .753, r = −.041. For the sad-dislikers, a Wilcoxon signed-rank test showed that, for the happy music condition, PRL decreased non-significantly from T1 (Median = 9.8, IQR = 6.0) to T2 (Median = 8.0, IQR = 4.8), Z = −2.20, p = .031, r = −.587. Similarly, in the sad music condition, decreases in PRL concentration between T1 (Median = 8.4, IQR = 4.7) and T2 (Median = 7.7, IQR = 2.0) also did not reach the statistical significance threshold, Z = −2.37, p = .016, r = −.634. The parametric test procedures justified by the non-significant Shapiro-Wilk tests for these data yielded results such that sad-dislikers were indeed found to experience significant PRL decreases after listening to happy, t(6) = 4.18, p = .006, r = .863, and sad music, t(6) = 3.61, p = .011, r = .827. Thus, PRL appeared to decrease in sad-likers listening to happy music as well as in sad-dislikers listening to happy or sad music. It is worth keeping in mind that these effects were not hypothesized, and our sample size of sad-dislikers was quite limited.
Openness to experience
Recall that hypotheses H3 and H4 proposed that those participants who showed the highest liking for sad music and those who showed the greatest change in serum PRL in response to sad music would be more likely to score high on openness to experience. Because changes in PRL were not normally distributed (see above) and because liking for sad music over happy music resulted from data provided on ordinal rating scales, the calculation of non-parametric Spearman rank-order correlations was justified.
Consistent with the extant literature we found that those listeners who scored higher on openness exhibited a greater liking for sad music over happy music, rs(37) = −.367, p = .021. However, we found no relationship between openness to experience and changes of serum PRL, rs(37) = .072, p = .663). The non-significant result for the change in serum PRL suggests that PRL is not part of the biological underpinnings of the relationship between openness and a preference for sad music. In short, our results are consistent with H3 but not consistent with H4.
Post-hoc tests
Aware of the problems of multiple tests, we can nevertheless cast a wide net and explore the findings of some post-hoc tests related to possible sex- and trait-related variables.
Sex differences
In general, research regarding strong emotional effects induced by music has identified a variety of differences between males and females. Some studies suggest higher emotional reactivity in females compared with males (Kamenetsky, Hill & Trehub, 1997; McFarland & Kadish, 1991; Webster & Weir, 2005; Wells & Hakanen, 1991). Other studies, however, have shown no effects of gender (Lundqvist, Carlsson, Hilmersson, & Juslin, 2009; Robazza, Macaluso, & D’Urso, 1994).
As the PRL difference scores were not normally distributed for females (n = 25) listening to sad, W(25) = 0.867, p = .004, and happy music, W(25) = 0.849, p = .002, and for males (n = 14) listening to sad music, W(14) = 0.861, p = .031, four nonparametric tests were conducted. With regard to females, for the sad music program, no significant difference between PRL concentrations at T1 (Median = 8.40, IQR = 5.4) and T2 (Median = 9.20, IQR = 5.3), Z = −1.172, p = .249, r = −.166 was evident. For the happy music program (after correcting for multiple tests), no significant difference was evident between PRL concentrations at T1 (Median = 9.80, IQR = 8.4) and T2 (Median = 8.00, IQR = 8.6), Z = −2.113, p = .034, r = −.299. With regard to males, for the sad music program, no significant difference was evident (after correcting for multiple tests) between PRL concentrations at T1 (Median = 8.00, IQR = 5.5) and T2 (Median = 8.20, IQR = 4.6), Z = −2.343, p = .017, r = −.443. For the happy music program, however, a significant difference was evident between PRL concentrations at T1 (Median = 10.60, IQR = 6.7) and T2 (Median = 8.70, IQR = 5.1), Z = −3.235, p < .001, r = −.611. A parametric paired-samples t-test confirmed this significant decrease in PRL from T1 (M = 12.92, SD = 10.51) to T2 (M = 11.25, SD = 10.30) in males listening to happy music, t(13) = 6.693, p < .001, r = .880. Notice that this result suggests a refinement to our earlier summary result. It appears that the reduction in PRL when some people listen to happy music is mostly or exclusively attributable to the responses of the male participants in our experiment.
Personality
With regard to personality, Table 3 reports correlations for each of 11 attributes, including the remaining four factors from the five-factor personality model (neuroticism, extroversion, agreeableness, and conscientiousness), the ULS-8 loneliness measure, and the five factors in the REMAP resilience instrument (relational engagement, emotional sensibility, meaning engagement, awareness of self and others, and physical health behaviors).
Nonparametric correlations between personality traits and measures of self-reported preference for happy over sad music and changes in serum prolactin (PRL) levels after listening to happy and sad music.
Statistically significant at α = .0015 (Bonferroni-corrected, 95 % confidence interval).
Whereas Shapiro-Wilk tests showed no violation of normality for any of the personality variables, all W(39) ⩾ 0.947, all p ⩾ .063, except for the meaningful engagement subscale of the REMAP resilience instrument, W(39) = 0.893, p = .001, nonparametric correlations were justified due to non-normal distributions of PRL changes (as previously reported) as well as of the sad-music-versus-happy-music preference scores, W(39) = 0.915, p = .006. We also cannot assume interval/ratio scale for self-reported personality measures provided on Likert-like scales. Thus, Spearman’s rank-order correlations are reported for each of three variables: the sad-music-versus-happy-music preference score, as well as the serum PRL changes in the happy and sad music conditions. All reported correlations involve 37 degrees of freedom. Corrected for multiple tests (33), the revised alpha level equivalent to a 95% confidence level was .0015.
Only one correlation was found to be statistically significant, namely the relationship between score on the loneliness scale and changes of PRL for the happy music condition. That is, the results are consistent with the post-hoc conjecture that listeners who scored high on loneliness experienced the greatest reduction in serum PRL in the happy music condition.
Pleasure
Further post-hoc tests focused on changes in self-reported pleasure over the course of the happy and sad music treatments. Does exposure to participant-selected happy and sad music make listeners report greater pleasure? In order to address these questions, we compared baseline self-reported pleasure (i.e., prior to any musical exposure) with self-reported pleasure following the musical exposure (see Figure 1(c)). Because data did not follow a normal distribution for pleasure ratings provided at T2 after sad, W(39) = 0.924, p = .012, and happy music, W(39) = 0.923, p = .011, nonparametric Wilcoxon signed-rank tests were employed. As might be expected, in the happy music condition, reported pleasure increased significantly from T1 (Median = 50.0, IQR = 40.0) to T2 (Median = 70.0, IQR = 30.0), Z = −4.63, p < .001, r = −.524. In the sad music condition, reported pleasure decreased significantly from T1 (Median = 50.0, IQR = 25.0) to T2 (Median = 20.0, IQR = 30.0), Z = −4.41, p < .001, r = −.499. That is, compared with baseline measures taken before the advent of the music treatment, listeners reported less pleasure following the sad music condition.
This latter finding raises the question of whether those participants who purport to prefer sad music over happy music exhibit less of a decline in self-reported pleasure compared with listeners who purport to enjoy happy music more than sad music. Once again, using the distinction between happy-music-likers (n = 20) and sad-music-likers (n = 8) derived from the sign of the difference scores between happy- and sad-music liking, we examined a possible association with the amount of change in reported pleasure for the sad music condition. Since Shapiro-Wilks confirmed that no violations of the normality assumption were present for the difference scores for pleasure ratings provided before (T1) and after (T2) listening to sad or happy music by sad- or happy-music-likers, both W(20) ⩾ 0.906 and both W(8) ⩾ 0.919, all p⩾ .053, parametric test procedures were employed. An independent-samples t-test on the difference scores for pleasure ratings provided before and after exposure to sad music showed no interaction between changes in pleasure ratings over the course of music listening and whether people prefer sad or happy music, t(26) = 0.360, p = .722, r = .070 (nor was there any interaction in the happy music condition, t(26) = 0.080, p = .936, r = .016). That is, a sad-music preference was not associated with less decline in pleasure over the course of the sad-music condition.
Discussion
Contrary to the motivating theory, exposure to nominally sad music did not result in a significant increase in serum PRL concentrations for listeners. Instead, a significant decrease in PRL was observed when listening to nominally happy music. The decrease in PRL concentrations for the happy music condition was found to be driven predominantly by the happy-music likers. Although there was no general change in PRL concentrations for the sad music condition, a post-hoc analysis showed a significant decrease for male participants when listening to happy music. Moreover, this decrease was associated with those participants who scored high on a loneliness scale. Taken at face value, when listening to happy music, people who like happy music, male listeners, and people who score high on loneliness appear to exhibit reduced PRL levels. We should note, however, that none of the effects reported in the preceding sentence were hypothesized a priori. Interestingly, Taruffi and Koelsch (2014) found that a considerable number of participants engage with sad music when feeling lonely. Most importantly, the results are not consistent with the main conjectured effect proposed by Huron (2011). That is, there was no significant positive correlation between changes in serum prolactin and changes in pleasure reported when listening to sad music.
With regard to hedonics, our results are inconsistent with a variety of conjectures relating pleasure to sad-music listening. Listening to sad music caused decreases in self-assessed pleasure. Moreover, the amount of decreased pleasure was not attenuated for those listeners who preferred sad music over happy music. The only increase in assessed pleasure occurred when listeners heard happy music. Consistent with earlier research we found a significant correlation between sad music preference (over happy music) and trait openness to experience. However, openness was not predictive of changes in PRL concentrations in response to sad music listening.
By way of summary, the main effect evident in this experiment is that listening to happy music led to a reduction of serum prolactin concentrations in listeners, and that this decrease in PRL was linked to an increase in self-assessed pleasure. In general, the results are inconsistent with the “prolactin theory of sad music” proposed by Huron (2011).
Potential confounds
As in any experiment reporting negative or non-significant results, one can point to several possible factors that may have confounded the observations. Many scholars (e.g., Davies, 1997) have argued that the “sadness” reported by listeners who enjoy listening to sad music cannot be genuine sadness, otherwise the experience would not be enjoyable. Several theories propose that the hedonic pleasure of negative emotional portrayals relies on a cognitive assessment that transforms the feelings to “aesthetic,” “mock,” or “fictional” sadness (Schubert, 2016). As noted earlier, in the sad-music condition, we augmented the sad musical stimuli with guided imagery vignettes (Mayer, Allen, & Beauregard, 1995). The intention was to enhance or ensure feelings of sadness. However, although the vignettes themselves were also “fictional” (not true statements of the listener’s experience), it may be that the overall experience might have been of genuine sadness rather than any purported aesthetic or mock sadness. Similarly, recall that many of our participants chose to listen to participant-selected music. These works could well be associated with particularly tragic events in the participant’s life, and so further increase the likelihood that participants were experiencing genuine sadness rather than some sort of aesthetic, mock, or fictional sadness. Accordingly, it might be suggested that the experiment failed to evoke the kinds of “aesthetic” feeling states that some aesthetic philosophers argue might render the experience of nominally sad music into a pleasurable event.
Another possible confound could arise from the tendency for participants to move in the happy music condition. In general, movement is associated with the release of dopamine. Research has well established that dopamine and PRL are intimately intertwined (Ben-Jonathan & Hnasko, 2001; Fitzgerald & Dinan, 2008; Van Vugt et al., 1979). Both serve as inhibitors of each other. Recall that in post-experiment interviews, some participants reported tapping their feet and humming along when listening to the happy music. Since decreased PRL occurred only in the happy music condition, it is possible that this change arises, in some way, due to music-induced movement. It is possible that the observed decrease in PRL arose in response to an increase in dopamine linked to the music-induced movement behaviors.
This same movement-related conjecture could potentially account for the observed differences between male and female listeners. A considerable research literature has established large effects of sex on general physical activity levels with males moving more than females. Given that this difference is robustly present already in infancy, it is assumed to have some biological basis (Campbell & Eaton, 1999). Gender-biased expectations and socialization processes, which are especially prominent in Western culture, are likely to magnify this effect during childhood and adulthood (Eaton & Enns, 1986). Note, however, that our experiment involved no systematic measurement of movement, so we have no evidence that our male participants could have engaged in greater music-induced movement than our female participants. Moreover, to our knowledge, sex differences in spontaneous movement have not been investigated for music listening.
Recall, furthermore, that the largest changes in PRL concentrations were found for participants who scored high on loneliness. Lonely people are known to engage in less spontaneous movement (Hawkley, Thisted, & Cacioppo, 2009) which argues against the above discussed “movement artifact” conjecture. On the other hand, it is possible that participants who score high on loneliness have generally higher levels of baseline PRL, and so the effect of movement-engendered dopamine would have a greater impact on PRL decline in participants exhibiting relatively high loneliness scores. A study by Quigley, Judd, Gilliland, and Yen (1980) offers support for this interpretation: specifically, Quigley et al. measured changes in serum PRL levels in response to dopamine infusion in healthy women as well as in female prolactinoma patients showing abnormally high levels of PRL. They found a striking .976 correlation between baseline PRL levels and the magnitude of PRL decline during dopamine administration. Accordingly, we carried out a post-hoc test to determine whether baseline PRL levels were higher for those scoring higher on loneliness. Spearman’s rho correlations between loneliness and baseline PRL concentration (i.e., T1) in the sad and happy music conditions were .218 and .252, respectively. Using a single-tailed test, the corresponding p-values were .091 and .061. Although not significant, the skew in the predicted direction and relatively low p-values are suggestive, lending some credence to the possibility that movement-engendered dopamine release may have confounded the experimental results. Apart from the possible relationship to loneliness, baseline PRL was unusually high in the happy music condition (significantly higher than the comparable baseline PRL in the sad music condition). Following the same logic, this suggests that the observed reduction of PRL following exposure to happy music might simply be a homeostatic artifact. By way of summary, it is possible that the observed changes in PRL in the happy music condition may be artifacts of a combination of high baseline PRL coupled with a propensity to move in the happy music condition, leading to increased dopamine which subsequently inhibited PRL levels.
Notice that, if the conjectured movement confound is the case, then it is likely to conceal or mask any potential increases in PRL arising from listening to sad music. That is, the “prolactin theory of sad music” could still be observable if participant movement were controlled. It should be added that this conjectured confound relies on anecdotal reports in post-experiment interviews as well as post-hoc interpretations and post-hoc tests that do not achieve statistical significance.
It should be recognized that our experimental manipulation focused on the musical condition rather than prolactin, per se. That is, our experiment did not manipulate PRL levels directly through either administering PRL or administering a PRL antagonist or inhibitor. Consequently, the conjectured causal relationship between PRL and pleasure was not directly tested. Moreover, it must be recognized that our mood manipulation involved both musical components and narrative vignettes. This means that the PRL changes observed in the happy music condition for some participants may not necessarily be attributable to the music, but could have resulted from the vignettes alone, or from the combination of the vignettes and the music.
Finally, recall that a preference for sad music was not associated with an increase in pleasure over the course of the sad music condition. More specifically, pleasure did not generally increase for the sad music condition for self-declared sad music lovers. This situation might arise for a number of reasons. For example, it may be that laboratory conditions are not conducive to experiencing nominally sad music as pleasurable, or it may be that the self-report questions were insufficiently sensitive in identifying sad music lovers or in identifying changes in experienced pleasure. Whatever the cause of this situation, it raises many questions concerning what was observed in this experiment.
For at least some people, listening to nominally sad music remains a pleasurable experience. This suggests that the “paradox” of negative emotion in the enjoyment of music cannot be a true paradox: there must be some explanation. Indeed, many candidate explanations have been proposed that appeal in varying degrees to empirical evidence (for partial reviews see Garrido & Schubert, 2011, and Garrido, 2017; for a taxonomy of different theoretical approaches, see Levinson, 2006; also, Levinson, 2014). While the enjoyment of sad music is perhaps not a paradox, it remains something of an enigma worthy of future research.
Supplemental Material
AppendixA_EnjoyingSadMusic – Supplemental material for Enjoying Sad Music: A Test of the Prolactin Theory
Supplemental material, AppendixA_EnjoyingSadMusic for Enjoying Sad Music: A Test of the Prolactin Theory by Olivia Ladinig, Charles Brooks, Niels Chr. Hansen, Katelyn Horn and David Huron in Musicae Scientiae
Supplemental Material
AppendixB_EnjoyingSadMusic – Supplemental material for Enjoying Sad Music: A Test of the Prolactin Theory
Supplemental material, AppendixB_EnjoyingSadMusic for Enjoying Sad Music: A Test of the Prolactin Theory by Olivia Ladinig, Charles Brooks, Niels Chr. Hansen, Katelyn Horn and David Huron in Musicae Scientiae
Footnotes
Acknowledgements
Our thanks to Holly Bookless, Claire Carlin, David Phillips, William Malarkey, Betsy Nini, and the nurses of the Center for Clinical and Translational Science at the Ohio State University Medical Center for their assistance in carrying out this study.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
