Abstract
Achieving a blended timbre for particular combinations of instruments, pitches, and articulations is a common aim of orchestration. This involves a set of factors that this study jointly assesses by correlating the perceptual degree of blend with the underlying acoustical characteristics. Perceptual blend ratings from two experiments were considered, with the stimuli consisting of: 1) dyads of wind instruments at unison and minor-third intervals and at two pitch levels, and 2) triads of wind and string instruments, including bowed and plucked string excitation. The correlational analysis relied on partial least-squares regression, as this technique is not restricted by the number and collinearity of regressors. The regressors encompassed acoustical descriptors of timbre (spectral, temporal, and spectrotemporal), as well as acoustical descriptors accounting for pitch and articulation. From regressor loadings in principal-components space, the major regressors leading to substantial and orthogonal contributions were identified. The regression models explained around 90% of the variance in the datasets, which was achievable with less than a third of all regressors considered initially. Blend seemed to be influenced by differences across intervals, pitch, and articulation. Unison intervals yielded more blend than did non-unison intervals, and the presence of plucked strings resulted in clearly lower blend ratings than for sustained instrument combinations. Furthermore, prominent spectral features of instrument combinations influenced perceived blend.
Keywords
In orchestration, composers may consider several factors when they intend to achieve a blended timbre between two or more instruments playing synchronously. There is the choice of suitable instruments that can yield a blended combination, which depends on the acoustical traits of these instruments. The remaining factors involve more musical considerations: whether instruments will be playing in unison or non-unison, which instrument is assigned to the top voice in non-unison passages, in what registral range the instruments will be playing, and what kind of articulation they will employ (e.g., bowed or plucked string). When it comes to establishing general associations between the perception of timbre blend and its underlying acoustical characteristics, the joint assessment of these factors will assist in predicting the perceived degree of blend for combinations of instruments, pitches, and articulations.
Previous research has defined perceived timbre blend as the auditory fusion of concurrent instrumental sounds, where individual sounds become less distinct. The most common method to measure perceived blend employs rating scales (Kendall & Carterette, 1993; Lembke, Levine, & McAdams, 2017; Lembke & McAdams, 2015; Sandell, 1995; Tardieu & McAdams, 2012). All studies found that spectral features influence blend, but employed different approaches to spectral description. One approach used the global descriptor spectral centroid, i.e., the amplitude-weighted frequency average of a spectrum. The composite (or sum) of the individual sounds’ centroids was found to predict blend in unison dyads best (Sandell, 1995; Tardieu & McAdams, 2012), whereas for non-unison dyads, the absolute difference in individual spectral centroids served as the more reliable predictor (Sandell, 1995).
Another approach to spectral description has considered the influence of prominent spectral features, such as maxima or formants. Similar to the relevance of formants in describing the acoustics of the human voice (Fant, 1960), wind instruments in particular exhibit formant structures that remain largely invariant across pitch (Lembke & McAdams, 2015; D. Luce & Clark, 1967; D. A. Luce, 1975; Schumann, 1929). Their identification and description can be achieved through spectral estimations that are aggregated across an instrument’s complete pitch range (Lembke & McAdams, 2015), and therefore can be considered pitch-generalized. Reuter (1996) has argued that similarity between instruments’ formant structures can explain blend. Hardly distinguishable instrument pairings can exhibit very similar formant locations (e.g., horn and bassoon), whereas the strongly pronounced, unique formant structure of the oboe may hinder it from blending with most other instruments.
Frequency relationships between the most prominent main formants appear to influence blend critically (Lembke & McAdams, 2015). In dyads comprising a recorded wind-instrument sound and a synthesized analogue to that instrument, whose main-formant frequency could be shifted relative to that of the recorded sound, blend decreased drastically as the frequency of the synthesized formant exceeded that of the recorded sound. This relative dependency relates to musical performance, where accompanying musicians adjust their main formants to be lower than when playing as the leading instrument (Lembke et al., 2017).
Apart from spectral properties, differences between temporal features, such as note attacks or onsets, have been found to explain blend as secondary factors for unison dyads (Sandell, 1995). However, their influence becomes more dominant as attacks turn impulsive: shorter durations and steeper attack slopes lead to reduced blend (Tardieu & McAdams, 2012).
With respect to those musical factors unrelated to timbre, blend for unison dyads is perceived as stronger than for non-unison combinations (Kendall & Carterette, 1993; Lembke et al., 2017). Furthermore, the assignment of instruments to the upper and lower pitches in non-unison intervals resulted in differences in perceived blend between instrument inversions in one study (Kendall & Carterette, 1993), but lacked a comparable effect in another (Sandell, 1995). All of these studies on blend are limited to dyadic contexts, leaving open how the obtained results and proposed hypotheses would fare in combinations of three or more instruments. Little work has been published on timbre combinations in triadic contexts (Kendall, 2004; Kendall & Vassilakis, 2006, 2010), and none of these articles address issues directly related to blend.
With the aim of predicting perceived blend between arbitrary instrument combinations, linear correlation or regression can be employed to associate blend measures with single acoustical features (Sandell, 1995; Tardieu & McAdams, 2012), without, however, making it possible to assess how several acoustical descriptors could jointly contribute to the explanation of blend measures. This limitation can be overcome by multiple linear regression (MLR). Past attempts have succeeded in explaining up to 63% of the variance in blend ratings for mixed-instrument dyads (Sandell, 1995). Similarly, MLR models also explained up to 87% of the variance in blend ratings across dyads in which the role of local, parametric variations of the main-formant frequency was studied (Lembke & McAdams, 2015).
Yet, the MLR approach also has clear limitations. High collinearity among independent variables (regressors) or a low number of cases compared to the number of regressors may both lead to less reliable and less valid results as well as mathematically ill-defined solutions. This becomes problematic given the aim of the current article, because many spectral descriptors are known to exhibit a high inter-correlation (Peeters, Giordano, Susini, Misdariis, & McAdams, 2011). For conventional MLR, this leaves two options: 1) disregarding the collinearity, at the risk of obtaining less reliable or invalid results or 2) eliminating regressors that are collinear to a reference regressor, that is, one found to predict blend most strongly in simple linear regression. However, the latter approach risks excluding variables that might perform even better than the selected one once they interact with other regressors.
A viable solution to deal with collinearity is to employ a dimension-reduction technique like principal component analysis (PCA) that reduces a high quantity of regressors to a small number of substitute or latent variables, that is, principal components (PCs), which are orthogonal to one another. These PCs can thereafter serve as regressors that represent the common aspects for groups of collinear descriptors (e.g., Giordano, Rocchesso, & McAdams, 2010).
A promising regression method that uses PCA as an integral part is partial least-squares regression (PLSR), which originates from the discipline of chemometrics, but has more recently been applied within the field of auditory perception (Eerola, Lartillot, & Toiviainen, 2009; Kumar, Forster, Bailey, & Griffiths, 2008; Rumsey, Zieliński, Kassier, & Bech, 2005). PLSR allows analysis of complex correlational relationships among perceptual measures and arrays of acoustical or psychoacoustical variables.
The current investigation uses PLSR in an attempt to predict blend ratings from perceptual experiments. The perceptual data are collected on a diverse set of variables that affect timbral blend and orchestration, including different instruments, pitches, and unison and non-unison intervals, as well as dyadic and triadic contexts. The set of potential regressors consists of a wide range of acoustical measures that, through several stages of PLSR models, are continually refined to retain only the most relevant regressors and, importantly, ones that are independent of each other.
Method
Partial least-squares regression (PLSR)
Predicting a single measure of blend through a set of regressors relating to acoustical descriptors can be expressed mathematically by associating the column vector of blend ratings
PLSR decomposes
Performance, predictive power, and reliability
Regression performance evaluates the variation in
Identifying relevant and independent regressors
The current PLSR analysis aims to reduce the number of investigated regressors in
Perceptual data sets
The regression analysis considers two data sets that originate from listening experiments in which participants provided blend ratings for dyads or triads. The two experiments were unrelated with respect to their original motivation and experimental design, yet they employed similar blend ratings, with the medians across participants taken as the dependent variable
The stimuli were presented over a standard two-channel stereophonic loudspeaker setup inside an Industrial Acoustics Company double-walled sound booth. They were drawn from recorded instrument samples from the Vienna Symphonic Library 2 (VSL), supplied as stereo WAV files (44.1 kHz sampling rate, 16-bit amplitude resolution). In separate pilot experiments, all stimuli had been both adjusted for perceptual synchrony between sounds constituting the dyads and triads and equalized for loudness within the dyad and triad sets independently. Adjustments for synchrony were based on consensus by three people for dyads and two for triads. The loudness equalization was conducted subjectively, anchored to a global reference for all dyad or triad conditions. The equalization was conducted by five people for dyads and six for triads. Gain levels were determined that equalized stimulus loudness to the global reference. These gain levels were based on median values across participants; all corresponding interquartile ranges were less than 4 dB.
For the main experiments, participants with varying degrees of musical experience were recruited from the McGill University community. All participants passed a standardized pure-tone audiogram (ISO 389–8, 2004; Martin & Champlin, 2000) ensuring that thresholds at all audiometric frequencies were less than or equal to 20 dB HL. Informed consent was obtained, and both studies were certified for ethical compliance by the McGill University Research Ethics Board II.
Dyads
Participants
A total of 19 people took part in the experiment (12 females and 7 males) with a median age of 21 years (range: 18–46). Among the participants, nine considered themselves amateur musicians, two as professional musicians, and eight as non-musicians. All were compensated financially for their participation in the hour-long experiment.
Stimuli
The stimulus set comprised a total of 180 dyads that resulted from the combination of several factors. Six wind instruments, namely, (French) horn, bassoon, oboe, C trumpet, B♭ clarinet, and flute, formed the 15 possible non-identical-instrument pairs listed in Table 1. These instrument pairs occurred at two pitch levels: C4 (
Fifteen dyads across pairs of the six investigated instruments.
All VSL samples were sustained, non-vibrato recordings, performed at mezzoforte dynamics, and were limited to the signal in the left channel. Both instruments were simulated as being captured by a stereo main microphone at spatially distinct locations inside a mid-sized, moderately reverberant room. Encompassing a volume of 600 m3, the relatively absorbent room yielded a reverberation time
Procedure
Participants heard individual dyads in randomized order and were asked to rate their degree of blend, employing a continuous slider scale with the verbal anchors most blended and least blended visualized on a computer screen. Ahead of the main experiment, participants had been familiarized with the degree of possible variation in blend among all dyads and had completed 15 practice trials on a separate but comparable stimulus set.
Triads
Participants
In total, 20 people (15 females and 5 males) with a median age of 21 years (range: 19–64) completed the experiment. Of the participants, 13 classified themselves as amateur musicians, with the remaining seven being non-musicians. All were remunerated for the hour-long experiment.
Stimuli
The stimuli comprised 20 triads, representing only a selection of the vast multiplicity of possible instrument and pitch combinations. In order to focus on timbral characteristics, all triads formed the same chord with pitches C4, F4, and B♭4, thus controlling for contextual effects with pitch register, chroma, and height. This chord choice of stacked perfect-fourth intervals avoided standard major and minor triads and generated the same consonances (perfect fourths) and a single dissonance (minor seventh). Such quartel chords were neutral enough not to draw attention to any one melodic voice, while allowing “inside,” middle voices to be easily heard.
In terms of instrumentation, the triads were composed of flute, oboe, B♭clarinet, tenor trombone, and cello sounds, corresponding to the instrument families woodwinds, brass, and strings. The instrument selection for triads (see Table 2) comprised mixtures between two or three instrument families. Furthermore, the selection included all woodwind reed types (air jet, single and double reed) and two different excitation types for strings (bowed and plucked excitation; arco and pizzicato, respectively), with each distinction represented by a single instrument (e.g., oboe for double reed, cello for string instrument). Instruments would only take on pitches based on conventional voice assignments given a particular mixture. For instance, the cello only occurred at the two lower pitches, whereas the flute was always highest in pitch relative to other woodwinds. Each instrument appeared in from six to 10 triads (counting different excitation types as separate instances).
Twenty triads and their constituent instruments and assigned pitches.
All samples were taken as stereo files from VSL, with woodwind samples comprising sustained sounds at mezzoforte dynamics and without vibrato. The trombone samples were similar, but at mezzopiano dynamics. The arco cello samples were recorded at mezzoforte dynamics. Unlike the wind instruments, they decayed after just a brief bow stroke, in order to be more similar to the pizzicato versions, which occurred at forte to allow for a longer sound decay. All cello sounds contained vibrato. The total duration for all triads was limited to 850 ms by applying an artificial 100-ms linear amplitude-decay ramp. A set of 10 representative triad stimuli can be found in the Supplemental Material Online.
Procedure
Participants were asked to sort all triads based on their relative degree of blend along a scale continuum with the verbal anchors most blended and least blended. At the beginning, visual icons for all triads were randomly arranged on a computer screen and could be dragged around or clicked on to trigger sound playback. Participants were first asked to identify two triads perceived as exhibiting the highest or lowest blend, to assign them to the extremes of the visualized continuum and then to position all remaining triads along the continuum. The sorting was conducted twice, the first counting as a practice round meant to familiarize participants with both the experimental task and the triads, the second serving as the main experiment.
Acoustical descriptors
For each data set, a collection of acoustical measures constitute the regressors in matrix
Acoustical descriptors investigated for dyads and/or triads (marked by “x” in the rightmost columns), related to the global spectrum (
ERB: equivalent rectangular bandwidth (Moore & Glasberg, 1983).
Descriptor relationships within dyadic and triadic contexts
As dyads and triads consist of several constituent sounds, their individual descriptor values need to be summarized to a single regressor value per stimulus by an association of some kind. For dyads with the constituent sounds
Timbre descriptors
Spectral descriptors
These descriptors assess properties associated with a time-averaged spectral representation. The investigated descriptors are computed on the output of one of two spectral-analysis methods: 1) analyses of the audio signals (∼) for individual instrument samples from VSL (e.g., oboe at G4) by use of the Timbre Toolbox (Peeters et al., 2011) employing harmonic analysis, and 2) pitch-generalized (∘) spectral envelopes (Lembke & McAdams, 2015), which are estimated by fitting a curve to partial tones aggregated across all available pitches from VSL (e.g., oboe from B♭3 to G6) and therefore allow for the characterization of an instrument’s pitch-generalized formant structure. Furthermore, the spectral descriptors can be distinguished as quantifying global and local spectral properties, as listed in Table 3. The global descriptors (
The formant structure derived from the pitch-generalized spectral envelope for a horn is shown in Figure 1. A set of frequencies indicate formant maxima (solid red lines; two formants are identified for this horn) and delineate their extent through lower and upper bounds (dotted lines) at which the magnitude has decreased by 3 dB. Note that in the case of the horn in Figure 1, there is no lower bound for the second formant, because the minimum point in the envelope between the two formants is not at least 3 dB below the maximum of the second formant. In this study, the focus lies on the main formant

Pitch-generalized spectral envelope of a horn with identified frequencies for formant maxima and 3 dB bounds (see Lembke & McAdams, 2015).
Furthermore, the degree to which wind instruments are characterized by formant structure varies, being strongest for oboe but much weaker for clarinet and flute. The measure

Pitch-generalized spectral envelopes of oboe (blue) and flute (red).
Temporal descriptors
Three descriptors characterize the time course of the amplitude envelope with respect to the attack (
Spectrotemporal descriptors
A pair of descriptors account for spectral variation across time, which the (static) spectral descriptors leave unaddressed. Previous research has not reported specific spectrotemporal (
Other descriptors and variables
The experimental designs involved factors that were likely to explain variance in median blend ratings but were not related to or not reliably measured through timbre features. Their relevance as potential regressors is assessed by several categorical variables (
For triads, a strong distinction was expected beforehand for the presence (D1) versus absence (D0) of pizzicato string sounds
Results
As mentioned under Method, PLSR analysis of a particular data set involves three stages, beginning with the original set of regressors
Performance (
Performance (
Dyads
PLSR models predicting median blend ratings for dyads initially involved 46 regressors (

Dyad model fit of
Figure 4 visualizes the loadings

PLSR loadings
Reflecting the main distinctions in median blend ratings, the scores
Figure 5 suggests that the spectral and pitch influence is independent (orthogonal) on the plane spanning PCs 2 and 3. The spectral regressors involve several composite

Dyad
Overall, the dyad data exhibit a complex structure of underlying factors, involving interval type, pitch level, and spectral features. Across all investigated models, their performance
Unison
A three-PC model on
As shown in Figure 6, the

Unison-dyad model fit of
PC 1 explains 22% of the variance and, as shown in Figure 7, appears to be linked to spectral composite

Unison-dyad
Non-unison
Twenty-three regressors in

Non-unison-dyad model fit of
As shown in Figure 9, PC 1 clearly reflects a grouping of dyads based on pitch level (

Non-unison-dyad

Non-unison-dyad
Triads
The PLSR analysis of triads first involved 61 regressors, which reduced to 30 regressors in

Triad model fit of
The main distinction found in Figure 12 along PC 1, which accounts for 85% of the variance, concerns the presence or absence of pizzicato cello sounds (the categorical variable

PLSR loadings
PC 2 explains the remaining 3% of the variance, appearing to relate to the distribution (
Discussion
Previous research has associated blend with acoustical measures describing spectral features, as well as temporal features like the attacks or onsets of sounds under certain circumstances. The current investigation pursued a correlational analysis using PLSR, modeling two perceptual data sets involving dyads and triads. PLSR loadings allowed us to evaluate the extent to which regressors were collinear or independent of each other. This approach helped select the most effective regressors. Applied to the complete data sets for both dyads and triads, the final models based on optimized regressor sets explain around 90% of the variance in median blend ratings. Notably, these levels of explained variance were still achieved after the elimination of non-essential regressors, that is, more than two thirds from the original set. The variation in both data sets is best explained by a dominant factor that is unrelated to spectral features.
For dyads, the distinction between unison and non-unison intervals explains 91% of the variance, with the fundamental-frequency difference ∆
In addition, even the second-most important factor in explaining the variation among dyads,
With regard to triads, the presence of a pizzicato cello evoked a strong decrease in blend ratings, whereas even triads including cello sounds excited by a single, brisk bow stroke led to substantially more blend. Again, this distinction had been anticipated, given that increasingly impulsive sounds have been associated with comparable decreases in blend (Tardieu & McAdams, 2012). Regarding the description of onset articulations, the difference in attack slopes
With both data sets being dominantly influenced by pitch or temporal features (e.g., attack), spectral descriptors only occur as secondary or even tertiary sources of variation in the modeled blend ratings. In perceptual tasks comparable to those employed in these experiments, participants may focus their attention on the dominant distinctions across stimuli at the cost of perceptual resolution for the less pronounced differences.
As the spectral factors likely only affected blend ratings in these regions of reduced perceptual resolution, the possible role of behavioral noise needs to be considered. Indeed, clear discrepancies between model performance
Three spectral descriptors stand out in explaining the PLSR models for both data sets, namely, the centroid of the pitch-generalized spectral envelope
Non-unison dyads yield a more complex relationship and involved the composite for
The results for triads expand previous knowledge beyond dyadic contexts. Even if spectral features only account for 3% of the variance, some new insight is gained from the distribution (
Overall, the global descriptor
When considering the relative locations of instrument combinations along the PCs that correlate with spectral features, a recurring pattern of dyads or triads including oboe (grey), on the one side, opposed to combinations involving horn or trombone (green), on the other, becomes apparent. Dyads or triads containing oboe are often less blended, whereas combinations with horn or trombone (e.g., bassoon and horn, clarinet and horn, trombone and trombone) are among the most blended ones. If we consider the notion of blendability of a particular instrument, the oboe should be considered a poor “blender,” which can be explained spectrally by its prominent and unique formant structure. Similar observations linking oboe to poor blend have been made in previous perceptual investigations (Kendall & Carterette, 1993; Reuter, 1996; Sandell, 1995; Tardieu & McAdams, 2012) as well as “prescriptions” found in orchestration treatises (Koechlin, 1954; Reuter, 2002). On the other hand, the horn is generally considered an easily blendable instrument, again reflected in perceptual results (Reuter, 1996; Sandell, 1995). The relatively “dark” timbre of the horn could support a general hypothesis of lower centroids leading to more blend (Sandell, 1995), at the same time supporting the argument that similar main-formant locations explain the good blend obtained between horn and bassoon (Lembke et al., 2017; Reuter, 1996).
In addition, Figure 12 illustrates that the distribution (
Conclusion
The present investigation shows that the perception of blended timbres in dyadic and triadic contexts correlates with a number of acoustical factors. Analyses using PLSR converged on an apparently reliable selection of independent predictors. The importance of factors such as pitch interval type, pitch, and articulation (e.g., impulsive vs. gradual note attack) became apparent. In addition, a group of spectral descriptors that exhibit the strongest predictive abilities could be identified from a wide range of descriptors, namely, the global spectral centroid and the upper frequency bound of main formants, which may represent the relevant features informing instrumentation choices. This wide variety of predictors suggests that in blend-prediction applications aimed at realistic musical scenarios, all factors should be taken into account. Given an appropriate acoustical characterization of instruments and details of how they are combined and employed musically (e.g., in unison or non-unison, the articulation and dynamic markings), these properties could suffice to predict the associated degree of blend.
One main challenge for future research is determining the effective weighting between these different factors of influence. Whether the clear dominance of interval type or impulsiveness of attacks over spectral features, which became apparent in the current investigation, would extend to more complex musical contexts remains to be explored. It can be assumed that the growing complexity that a listening scenario involving musical contexts would present, given the simultaneous presence of other musical parameters, could significantly alter the relative importance of factors found in listening experiments employing isolated dyadic or triadic stimuli.
For instance, a composer may assign a unison blend between two instruments to a melodic voice while juxtaposing this against a chordal, non-unison accompaniment layer whose instruments are chosen to blend amongst themselves into a homogeneous timbre. On another level, the melody may become more distinct from the accompaniment due to the distinction between unison and non-unison, which may also be desired. This case scenario illustrates that blend-related factors need not stand in competition with each other like they do in the investigated perceptual data, but instead could operate on independent levels, fulfilling separate functions within the musical context.
For the composer, working with blend is not a matter of favoring unison intervals over non-unison intervals, but being able to employ it at individual levels of the musical scene (e.g., melody, accompaniment, or contrasting the two). Within each level, blend is achieved by relying on the same principles, that is, similarity in spectral description as well as articulatory features (e.g., note attacks). This hypothetical scenario encourages future work on blend-prediction models to rely on perceptual data obtained from stimuli involving musical contexts (Kendall & Carterette, 1993; Lembke et al., 2017; Reuter, 1996), as it provides a more realistic setting from which weights between blend-related factors could be estimated. We thus propose the need for a meta-analytical investigation into a diverse range of perceptual blend data, in an attempt to move toward generally applicable blend-prediction techniques.
Supplemental Material
Supplementary_Material – Supplemental material for Acoustical correlates of perceptual blend in timbre dyads and triads
Supplemental material, Supplementary_Material for Acoustical correlates of perceptual blend in timbre dyads and triads by Sven-Amin Lembke, Kyra Parker, Eugene Narmour and Stephen McAdams in Musicae Scientiae
Footnotes
Acknowledgements
The authors would like to thank Bennett K. Smith for programming the control software for both experiments and Emma Kast for the recruitment and running of participants for the experiment involving triads. We also thank two anonymous reviewers and the editor for their valuable feedback on an earlier version of this article. The findings reported in this article were conducted as part of the first author’s doctoral research and are also featured in his thesis (Lembke, 2015).
Funding
This research was partly funded by an ACN CREATE undergraduate research award to Kyra Parker as well as a Canadian National Sciences and Engineering Research Council grant (RGPIN 312774-2010) and a Canada Research Chair to Stephen McAdams.
Supplemental Material
Representative examples for the dyad and triad stimuli are available online as supplemental material, which can be found as part of the online version of this article at http://msx.sagepub.com. Access the sound files through the hyperlink “Supplemental material”. Their file names follow the naming convention found in Table 1 and
.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
