Abstract
Since Lippius and Rameau, chords have roots that are often voiced in the bass, doubled, and used as labels. Psychological experiments and analyses of databases of Western classical music have not produced clear evidence for the psychological reality of chord roots. We analyzed a symbolic database of 100 arrangements of jazz standards (musical instrument digital interface [MIDI] files from midkar.com and thejazzpage.de). Selection criteria were representativeness and quality.The original songs had been composed in the 1930s and 1950s, and each file had a beat track. Files were converted to chord progressions by identifying tone onsets near beat locations (±10% of beat duration). Chords were classified as triads (major, minor, diminished, suspended) or seventh chords (major–minor, minor, major, half-diminished, diminished, and suspended) plus extra tones. Roots that were theoretically less ambiguous were more often in the bass or (to a lesser extent) doubled. The root of the minor triad was ambiguous, as predicted (conventional root or third). Of the sevenths, the major–minor had the clearest root. The diminished triad was often part of a major–minor seventh chord; the half-diminished seventh, of a dominant ninth. Added notes (“tensions”) tended to minimize dissonance (roughness or inharmonicity). In arrangements of songs from the 1950s, diminished triads and sevenths were less common, and suspended triads more common, relative to the 1930s. Results confirm the psychological reality of chord roots and their specific ambiguities. Results are consistent with Terhardt’s virtual pitch theory and the idea that musical chords emerge gradually from cultural and historic processes. The approach can enrich music theory (including pitch-class set analysis) and jazz pedagogy.
Keywords
Harmony is the art of combining simultaneous tones to create chords. The term also refers to the successive chords that make up harmonic progressions in Western music (Schoenberg, 1911/1978). In these artistic processes, some tones and some chords are perceived to be more important than others, or to function as psychological references for the perception of other tones or chords (an aspect of tonality; Dahlhaus, 1967/1990; Krumhansl, 1990). The present article addresses chord roots, understood as psychological reference pitches for simultaneous tones or chords (Parncutt, 1988; Thomson, 1993). We will present new evidence that, although chord roots are historically specific aspects of music, they depend on ahistorical, non-musical perceptual processes, namely, the perception of pitch in harmonic complex tones, such as voiced sounds in speech (Terhardt, 1982; Terhardt et al., 1982).
Jazz harmony is based on classical Western harmony (McGowan, 2010). That is true despite the contrasting histories of classical music and jazz in European and African-American cultural contexts, and the persistently White racial frame of the academic discipline of music theory and the hierarchical concept of major–minor tonality (Ewell, 2021). In this article, we argue that chord roots play a similar role in both jazz and classical traditions, for two reasons. The first reason is historical: the relationship between a chord and its root emerged in European polyphony and was later developed within jazz. The second is psychological: the relationship is based in both cases on a perceptual universal, namely the perception of harmonic complex tones in speech. In particular, we claim that the pitch-class salience profile or (almost equivalently) pitch-class stability profile of a chord (Parncutt, 2011) is essentially the same in both cultural traditions and can be accounted for by the same theory based on the perception of harmonic complex tones in non-musical sounds. We support this hypothesis by measuring pitch-class salience/stability profiles in jazz in a new way, from computer-based analyses of recorded jazz arrangements.
In this introduction, we address the phenomenon of chord roots and their ambiguity in jazz by first revisiting the historical musical contexts in which chord roots emerged. We then survey current attempts to explain chord roots psychologically, comparing the contrasting ways in which music theorists, psychoacousticians and psychologists have approached the topic. A long-standing debate about the root of the minor triad is tentatively resolved. We then revise existing predictions for the root(s) of familiar chords, and consider how they might be tested.
History of chord root theory
Approaches to understanding the perception and cognition of musical structure can be divided into two main categories: historical (in the humanities) and psychological (in the sciences). Both approaches may be necessary for a full understanding, and both approaches may involve the history of music perception. The perception of musical structure, which depends to a large extent on familiarity with music, changed gradually as music changed historically. In that sense, the musical experience of a modern individual depends not only on the music that person has heard in her or his lifetime, but also more indirectly on a long historical process. Modern discourse about musical structure also depends on the history of ideas on this topic.
To modern ears, Renaissance polyphony comprises chord progressions that are dominated by major and minor triads (more often major than minor, Parncutt et al., 2018). But Renaissance composers had no terminology for triads, and the idea of chord roots did not yet exist (Rivera, 1984). Compositional practice and music-theoretic discourse referred instead to interval combinations.
The idea of chord roots emerged gradually in the 17th and 18th centuries. Lippius (1612) proposed that the basis of harmony was the three-note chord or trias harmonica (Rivera, 1978, 1984). Lippius also investigated voice-leading in triadic chord progressions, noting that chords usually have a third and fifth above the bass (i.e., triads were usually in root position) and intervals between successive bass tones are often leaps (fourths and fifths), whereas melodic intervals are more often steps.
Over a century later, Rameau developed the modern concept of the chord root (basse fondamentale). His first major treatise (Rameau, 1722) built upon Zarlino’s (1558) concept of intervals and their consonance or dissonance. Zarlino had observed that intervals can be derived or generated by dividing a vibrating string (the monochord) into one, two, three, four, five, or six equal lengths, producing the unison, octave, octave plus perfect fifth, double octave, double octave plus major third, and double octave plus perfect fifth. On that basis, Rameau explained that the octave, fifth and third intervals were generated from a single fundamental sound.
A few years later, Rameau (1726) realized the crucial relevance of the harmonic series for music theory. In the history of mathematical and music-theoretical ideas, the harmonic series idea had been emerging gradually (Green, 1970). It was nascent in Mersenne (1636), Sauveur (1701), and Rameau (1722). Rameau (1726) realized that Zarlino’s six intervals were physically present in a single harmonic complex tone, which he called the corps sonore. In the spirit of the French enlightenment, he presented his theory as scientific, based on the laws of physics and mathematics (Christensen, 1993).
Chord roots in 20th-century tonal theory
Since Rameau, music theorists have taken for granted that chord roots are usually voiced in the bass, that is, chords appear more often in root position than inversion. Many music theory texts also recommend doubling chord roots more often than other chord tones, that is, sounding them simultaneously in different octave registers. The approach of Lovelock (1974) seems old-fashioned nowadays, but it does at least give students clear guidelines: In arranging a triad for four voices, one note has to appear in two parts, i.e. it is “doubled.” The best note to double is the root, the next best the 5th. In a minor triad the 3rd may be doubled, but in a major triad this is undesirable (for the present), except in certain special circumstances mentioned below. The leading-note may never be doubled. For the moment it should be taken that the root only is to doubled. (p. 3, italics in original) Doubling in first inversions needs consideration. In a major first inversion, double either the 3rd or the 6th above the bass, but avoid doubling the bass itself; in a minor first inversion any one of the three notes may be doubled, but there is no need to double the bass without sufficient reason. (p. 21)
The authors of music theory texts tend to expound such principles on the basis of their experience of music from the canon. Explaining where rules of this kind ultimately come from is not so easy. Presumably, they involve culture-specific responses to consonance and dissonance (Cazden, 1945). Westerners traditionally like dissonances only when they are suitably surrounded by consonances and connected by appropriate voice-leading; the dissonances must resolve correctly.
If we consider sonorities in isolation, major and minor triads appear more often in root position than in inversion (Eberlein, 1994), suggesting that root position chords are perceived to be more consonant. The more general principle may be that chords with clear roots are perceived to be more consonant. The root can be clarified by voicing it in the bass or doubling it. The effect may be an example of harmonicity (similarity between the spectrum of a sound and the harmonic series)—part of a psychological theory of consonance and dissonance (Harrison & Pearce, 2020; Parncutt & Hair, 2011) that was first clearly expressed in Stumpf’s (1911) theory of perceptual fusion. The pitch ambiguity of a sonority may be reduced, and consonance increased, by increasing harmonicity.
Mainstream harmony texts clarify that the leading tone (the seventh tone of a diatonic scale, which is raised relative to the key signature in minor keys) should never be doubled. The rule can be explained in two different ways. First, a doubled leading tone is likely to rise by a semitone to the tonic. If that happens in two parts at the same time, parallel octaves result. Second, the leading tone may destabilize the current tonality if it is emphasized too much. Of the seven diatonic tones in a major or minor key, the leading tone is the least stable (Krumhansl, 1990). Huron (1993) observed more generally that more stable tones within major and minor keys are more often doubled in mainstream tonal music.
Discussions of this kind assume octave equivalence (Révész, 1913). Tones at octave intervals are understood to be harmonically equivalent, even if different chordal inversions are not treated exactly the same as each other (as illustrated by the quotation from Lovelock, above). Thus, the root of a chord exists in different octave registers (in principle, in all 10 audible registers).
Here, we make the simplifying assumption that throughout the history of Western polyphony the 12 degrees of the chromatic scale (tuned approximately) have been equally available for music performance—ignoring enharmonic differences, tendencies to use some tonalities more than others (because some patterns are easier to notate or perform than others), and dependencies of key profiles on absolute pitch (Eitan et al., 2017; Quinn & White, 2017). In a quasi-objective scientific approach to understanding chord roots, the 12 chromatic scale degrees can be regarded as equally important a priori. That makes pitch-class set theory (Forte, 1973) an appropriate foundation for a psychological theory of chord roots. In such a theory, the root is a pitch class; music psychologists, such as Krumhansl (1990), use the equivalent term chroma.
Simple pitch-class set theory can be used to label all possible sonorities within the 12-tone chromatic scale. Labels based on pitch-class sets have the advantage of neutrality with respect to diatonic, enharmonic, or tonal implications. Similarly, in discussions of Renaissance counterpoint, Schubert (1999) used the term “three-pitch-class sonorities” (p. 210), avoiding the tonal connotations of “triad;” one may similarly use the neutral term “trichord.” But, it is sometimes difficult to avoid anachronistic terminology; Schubert made his chapter on four-part writing more readable for modern students by anachronistic use of terms, such as “chord” (“6/3 chord”) and “diminished triad” (p. 234).
The root of a chord is often ambiguous, and the degree of ambiguity depends on the chord. For example, Rameau was unsure whether the added-sixth (sixte ajoutée) chord on the subdominant (e.g., in C-major: FACD) was in root position or first inversion, coming to different conclusions in different texts (Christensen, 1993, p. 118). In fact, the root of this chord is intrinsically ambiguous. In jazz, if it is called F6, its root is by implication F, whereas if it is called Dm7, its root is D. In octave-generalized music theory, these two chords are one and the same, since both comprise the same pitch classes (D, F, A, and C).
The added sixth or minor seventh chord is not special in this regard. The root of any chord can be ambiguous. Take the major triad. Normally, one would think its root is quite unambiguous. But when a simple C chord appears in second inversion (GCE) and resolves to a dominant triad (GCE–GBD), the chord is (or can be) heard as a cadential 6/4 with root G (Beach, 1967). Relative to that reference pitch, the sixth above the bass (E) falls by step to the fifth (D), and the fourth(C) falls to the third (B). In Schenkerian terminology, this is a G chord with suspended sixth and fourth (“double suspension”)—a prolongation of G. If, however, the same chord resolves differently (e.g., GCE–FCF), GCE is (or can be) heard as a C chord in second inversion. Depending on the preceding and following voice-leading—perhaps also the listener’s background—the chord has two possible roots: C or G. In jazz notation, the chord GCE is called C/G, suggesting that its root is C. However, that kind of notation is more pragmatic than theoretical, aiming primarily to clarify which tones are to be played.
Root ambiguity is most striking for chords or simultaneities that divide the 12-tone chromatic scale equally: the tritone interval, the augmented triad, the diminished seventh chord, and the whole-tone scale. But the root of BDFA, a non-symmetrical half-diminished seventh chord, is also clearly ambiguous. In that case, an important possible root is not sounded: G. The chord often functions as dominant ninth chord (G9) with missing root, which usually then resolves to the tonic (in C major: BDFA–CEG). Other possible roots are B (implying a Bø7 chord) and D (Dm6).
Three relevant papers
A theory of chord roots worthy of the label “scientific” should predict not only the most likely root of a chord, but also its ambiguity, and it should do that in a quantitative, testable, falsifiable fashion. How might such a theory be constructed? In 1982, three relevant academic publications appeared, all of which shed light on this problem—albeit in contrasting ways.
The first relevant publication came from the discipline of empirical psychology. Krumhansl and Kessler (1982) measured the relative stability of tones in a musical key by presenting a harmonic cadence (chord progression ending on the tonic) followed by a single tone and asking listeners, “how well, in a musical sense, each probe tone fit into or went with the musical element just heard” (p. 342). Participants were musicians: undergraduate students with 11 years of musical instruction and 12 years of music performance experience, on average, but no formal training in music theory.
The result was a hierarchically structured profile of stability as a function of scale step. In both major and minor keys, the peak of the profile was the tonic (diatonic scale step no. 1), followed by dominant (5) and mediant (3), followed by the other diatonic tones, followed by the remaining chromatic tones. The results can also be interpreted relative to the tonic triad: the peak was the root, followed by the fifth and the third, and finally, other tones that follow that chord well. Krumhansl’s 12-element profile was surprisingly robust when measured in different ways or with using different listeners (e.g., Lamont & Cross, 1994), with small exceptions (Quinn & White, 2017). The tone profiles could be interpreted as a measure of the clarity versus ambiguity of the tonal center: the greater the stability of the tonic relative to the other tones, the lesser the ambiguity; and the flatter the profile, the greater the ambiguity. The tonic was clearer, and the profile more peaked, for more diatonic or conventional triads and progressions, including those with falling fifths between successive roots (Parncutt & Bregman, 2000).
The second relevant publication was the pitch algorithm of Terhardt et al. (1982). The authors predicted the main perceived pitch of any steady-state sound, invoking only non-musical principles of pitch perception. Their model had the advantage of non-circularity (explaining aspects of music without prior assumptions about music) and falsifiability (the model could, in principle, make predictions that would contradict music theory). Beyond predictions about musical chord perception, the model accounted for diverse non-musical empirical findings in the area of pitch perception. When a musical chord was presented as input, several pitch candidates were predicted, not all of which corresponded to notes in the musical score. The predicted pitches differed from each other in perceptual salience. If the tones of a chord were equally loud or salient when heard alone, they were not equally salient when sounded simultaneously. The most salient pitch in the resultant chord, according to the algorithm, tended to correspond to the conventional root, especially in the bass region.
But the situation is more complex than that. Any tone in a chord can be made more salient simply by playing it louder—an effect that Terhardt’s algorithm also modeled. But in everyday common-sense music theory, the root of a chord does not change when one tone is played louder. To account for that, we need an octave-generalized model—one that ignores the octave register of the chord notes—in which tones are assumed to be equally loud or salient when heard alone.
The third 1982 paper addressed the issue of chord roots directly. Terhardt (1982) took the first 10 partials of a harmonic complex tone (i.e., the first 10 elements of the harmonic series) and octave-transposed them into one octave register, reducing the number of intervals to five: unison, perfect fifth, major third, minor seventh, and major second. The root of a chord was predicted by starting from the chord tones and transposing down through these five intervals, that is, looking for subharmonics of each chord tone. When a subharmonic of one tone coincided with a subharmonic of another tone, a root was predicted. This procedure was a simplified version of a more general subharmonic coincidence algorithm in the work of Terhardt et al. (1982).
Such subharmonics seldom coincide exactly, raising the question of how close two subharmonics need to be for a coincidence to be perceived. Moore et al. (1986) found that a partial within a harmonic complex tone can be mistuned by as much as a quarter tone (3% of frequency) or more without being consciously perceived, such that the complex tone as a whole is perceived to have two pitches rather than one (a virtual pitch at the fundamental and a spectral pitch at the mistuned partial). Moreover, the threshold of mistuning was almost independent of harmonic number. A possible reason for this seemingly wide range of tolerance is the physical existence of stretched harmonic series in the human environment (e.g., freely vibrating strings, such as piano strings). Another possible reason is psychoacoustic pitch shifts: the perceived pitch of a pure tone depends on its sound level and on the presence of masking sounds (e.g., de Cheveigné, 1999; Hartmann & Doty, 1996; Stoll, 1985; Verschuure & Van Meeteren, 1975; Walliser, 1969). The ear’s tolerance for mistuning of harmonics is consistent with experiments on categorical perception of musical pitch (e.g., Burns & Ward, 1978): a musical interval can be out of tune by up to a quarter tone and still be reliably identified by a musician.
These observations resolve long-standing issues about the relevance of different theoretical tuning systems. In a psychological theory of the perception of chord roots and musical intervals in musical contexts, exact tuning is not important. Predictions should be almost independent of theoretical tuning system—whether Pythagorean, just, equally tempered, or another musically acceptable approach.
The minor triad: A special case?
The predictions of Terhardt’s (1982) simple procedure were intuitively correct for most commonly heard triads and seventh chords, but not for the minor triad. That was nothing new—the minor triad had been bothering countless music theorists since Rameau. The problem begins with the major triad, which maps directly onto the harmonic series: Harmonics 4, 5, and 6 describe a major triad in root position, while 3, 4, and 5, are a major triad in second inversion. That clear mapping gave music theorists the idea that the harmonic series might represent the origin of all chords. The trouble is, the minor triad—the second most important chord in Western music—does not match the harmonic series in any transposition.
Addressing this problem, Terhardt (1982) sided with generations of previous music theorists, claiming that the minor triad is somehow less “natural” than the major and hence less suitable as a tonal reference. He sidestepped the minor triad problem by pointing out that the first four root candidates according to his simple algorithm were the same for two closely related chords: the C-major triad CEG and the A-minor triad ACE. In both cases, they were C, D, F, and A. That is at least consistent with the strong “parallel” or “relative” relationship between these two chords (Riemann, 1893).
The trouble with claiming that the minor triad is “unnatural” is that there is nothing particularly “natural” about the major triad—or any other chord, for that matter. Musical chords are cultural creations—results of cultural and historic processes, in which generations of composers and performers tried out different tone combinations in different musical contexts and gradually developed preferences for some over others. Western musicians of the past preferred chords that were more consonant for three main psychological reasons (Harrison & Pearce, 2020; Parncutt & Hair, 2018): harmonicity (Stumpf, 1911), smoothness (von Helmholtz, 1863), and familiarity (Cazden, 1945).
Inspired by the algorithm of Terhardt et al. (1982), Parncutt (1988) weighted Terhardt’s five intervals relative to each other (calling them root-support intervals). The unison was assigned weight 1, the perfect fifth 1/2, the major third 1/3, the minor seventh 1/4, and the major second 1/5. The idea was to reduce the importance of intervals associated with higher harmonics, given that they are less often audible than lower harmonics (because their sound level is lower and they mask each other more). Harmonics above the tenth were not considered because they are usually inaudible in everyday sounds, including speech (Terhardt, 1979, p. 160; 1998, p. 323; cf. Plomp & Mimpen, 1968). The revised algorithm correctly predicted the root of the minor triad. It also predicted (in agreement with Terhardt) that the minor triad’s root is more ambiguous than that of the major triad. That is, the difference in predicted salience between the first two root candidates is smaller for the minor triad than for the major. Specifically, the second most likely root of the minor triad is the minor third above the conventional root. These predictions are also consistent with rules of doubling and inversion as found in music theory texts. The model also predicted intuitively plausible 12-element tone profiles for any chord in the chromatic scale.
How can predictions of that kind be tested? Parncutt (1993) created a series of familiar musical chords from octave-complex tones—similar to the Shepard tones used by Krumhansl, but with flat spectral envelopes (the amplitude of individual partials being limited only by the limitations of sound preproduction equipment, and their audibility by hearing thresholds). These chords were presented to listeners, followed by single tones (also octave-complex). Each trial was randomly transposed around the chroma cycle, and trials were presented in a random order that was different for each listener. Listeners were asked to rate how well each tone went with the immediately preceding chord. The results did not clearly or consistently confirm the existence of chord roots, but a modeling procedure did suggest a slight change to the algorithm. The weight given to the less important root-support intervals was subsequently reduced further relative to the more important ones. The weight for the unison was changed to 10, for the fifth to 5, for the major third to 3, for the minor seventh to 2, and for the major second to 1 (Parncutt, 1997). Small whole numbers were chosen to simplify music-theoretic calculations. Given the noise in the data of experiments on pitch salience in musical chords, more accurate estimates are unlikely to significantly improve predictions.
Predictions of Parncutt (1988) with these revised root-support weights are shown in Figures 1 and 2. Musically interesting predictions include the following. All non-chord tones with relatively high predicted weights may be regarded as missing roots. A C-major triad has missing roots at A (nine semitones above the conventional root) and F (5). The root of a C-minor triad can be C or E♭ (three semitones); missing roots include F and A♭. The diminished triad on C has a strong missing root at A♭. The root of a C suspended fourth triad (CFG, Csus4) can be C or F; in the latter case (FGC), we have an F chord with suspended second. The tetrad with the least ambiguous root is the major–minor (“dominant”) seventh, which can explain why it is so prevalent despite the dissonant tritone between the third and seventh. The root of the C-minor seventh chord is ambiguous (C or E♭). The C-major seventh chord has a missing root at A. The most likely root of a C-half-diminished seventh chord is E♭, making it an E♭-minor sixth chord. The chord also has an important missing root: A♭. The C seventh with suspended fourth (C7sus4, CFGB♭) has two main roots, C and F—the latter making it into an Fsus2sus4 (FGB♭C).

Predictions for triads according to Parncutt (1988), with root-support weights adjusted as described in the text: (a) major triad, (b) minor triad, (c) diminished triad, and (d) suspended triad.

Predictions for tetrads according to Parncutt (1988): (a) major–minor seventh chord, (b) minor seventh chord, (c) major seventh chord, (d) half-diminished seventh chord, (e) diminished seventh chord, and (f) suspended seventh chord.
These predictions may make music-theoretic sense, but empirical confirmation is not so easy. Thomson (1993) pointed out that “empirical studies of interval perception have fallen short of confirming the phenomenal reality our concepts describe so confidently” (p. 385). To the authors’ knowledge, the first psychological experiment to clearly confirm the existence of chord roots was Experiment 3 of Parncutt et al. (2019). Here, listeners heard a chord (simultaneity of octave-complex tones) and actively chose the best-fitting chroma from 12 options (also octave-spaced). All tones were tuned to 12-tone equal temperament. For eight different trichords (major, minor, suspended, diminished, and four others; in semitones relative to the root, 047, 037, 027, 036, 015, 045, 025, 035), listeners clearly preferred the upper tone of a perfect fourth interval within the chord (or equivalently the lower tone of a perfect fifth). Supplementary data suggested that participants were only rarely able to recognize chords and respond on the basis of music-theoretical knowledge. Therefore, they were presumably responding to the most salient pitch.
Another approach to investigating the perception of chord roots is to statistically analyze a database of representative musical scores. Parncutt and Reisinger (2020) converted representative scores of vocal polyphony from the 13th to 19th centuries to chord progressions, automatically creating a new chord whenever there was a notated onset in any voice. They then studied profiles of tones that immediately followed specific chords, focusing on the same eight trichords as before. Again, the peaks of profiles of immediately preceding and following tones did not correspond consistently or clearly conventional roots. The shape of the 12-element profiles depended on instead on diatonic scales and fifth relationships, and to a lesser extent on missing fundamentals (roots) and completion tones (tones that complete familiar tetrachords).
The present study takes a new approach. We explore evidence for root perception in a symbolic database of jazz arrangements. Instrumental arrangements of jazz standards were downloaded from the internet in musical instrument digital interface (MIDI) format. They were automatically converted to chord progressions on the assumption that most of the tones that comprise chords in chord progressions begin on the tactus—for example, at or near one of the four beats in a 4/4 measure. On this basis, we explored—for the first time using automatic analysis of a large database—how often the root of a chord is voiced in the bass or doubled in different octave registers by comparison to other chord tones. Given that chords tend to happen more often in root position and chord roots tend to be doubled more often than other tones—at least according to some music theorists, and in our own experience—we assumed that new data on inversion and doubling in a relevant musical corpus would enable interesting insights into the relationship between chord tones and the root.
Our model enabled specific prior predictions to be made about chord tones that are most often voiced in the bass or doubled. These predictions were then compared with statistical analyses of chord voicings in a representative database. That procedure is scientific in the sense that the data could contradict the predictions, falsifying the theory. The analyses could in principle show that the predicted tones are not, as expected, voiced more often in the bass or doubled more often, or that quite different tones were voiced more often in the bass or doubled. If the data did not contradict predictions either way, the theory would be confirmed in an empirical scientific sense. It might nevertheless still be possible for a different or revised theory to better predict the data.
Method
Jazz arrangements were downloaded as MIDI files from two free internet pages: midkar.com and thejazzpage.de. We focused on arrangements of songs that had been composed in the 1930s or 1950s (listed in Appendix 1). These two different time periods were intended to make our music selection more representative; we were primarily interested in what these two periods have in common, for the purpose of modeling based on assumptions about perceptual universals.
The arrangements themselves were not dated. To our knowledge, the arrangers (“sequencers”) typically combine materials from different sources, including arrangements by other people, lead sheets, and recordings. They build up a MIDI file in stages by adding tracks using a MIDI keyboard and making manual adjustments—all relative to a constant beat track. The files are created for performance purposes, for example, as a backing track against which a soloist can improvise.
Criteria for inclusion in the databases were representativeness and quality. We listened to each arrangement and decided subjectively whether the music would normally be recognized as jazz by someone familiar with the style as a listener, and whether the arrangement was well written. Typical audible characteristics of jazz were understood to include constant dance rhythm (swing), relatively complex harmony (including “blue” dissonances), and quasi-improvised solos. A further criterion was the presence of a beat track in the MIDI file, which we assumed corresponded to the tactus.
According to these criteria, 50 files were accepted from each of two decades. The average duration of the files was 231 s for the 1930s and 195 s for the 1950s. The average number of notes per file was 3,037 for the 1930s and 2,531 for the 1950s. The tempo range of the beat track was 61–252 beats per minute (BPM) for the 1930s and 65–245 BPM for the 1950s. The mean BPM was 131 for the 1930s and 129 for the 1950s. If perceived tempo corresponds more closely to a logarithm of BPM, it may be more meaningful to convert tempos to logarithms, take the mean, and convert back (Parncutt, 1994); that calculation yielded an average tempo of 124 BPM for the 1930s and 123 BPM for the 1950s.
Each file was converted to a chord progression in the following way. A time window was created around each beat in the beat track. The duration of the window was 20% of the beat duration, starting 10% before the beat and ending 10% after it. The chord on that beat was specified by the pitch classes of all tones whose onset fell in that window. Sustained tones within the window were ignored. This procedure identified 28,564 chords in the 1930s sample and 29,061 chords in the 1950s.
Although these results differed from the chord progressions that a musician or music theorist would usually notate, the method did identify the main chords and did not, to our knowledge, introduce a bias that would affect our main findings. Nor could we find a more appropriate automatic procedure in the literature on retrieval of information from musical scores (symbolic data; cf. Lee & Slaney, 2008). Our method has the advantage of simplicity and replicability. More complex methods may be more arbitrary, and the results may be more difficult to interpret.
In advance of data collection and analysis, we predicted that:
conventional chord roots, and pitch classes predicted by Parncutt (1988) to have high perceptual salience, would be voiced more often in the bass and doubled more often, by comparison to other pitch classes;
if the root of a chord was predicted (either music-theoretically or according to the psychoacoustic model) to be ambiguous (e.g., the root of ACE may be A or C), the predicted ambiguity (including the specific pitch classes that compete for the status of root) would be confirmed by the data;
pitch classes of intermediate salience that do not correspond to notes (chord tones) would also emerge from the data (e.g., the pitches D, F, and A that are implied by the chord CEG, according to the psychoacoustic model); and
tones that are often added to familiar trichords to form tetrachords would be chosen to minimize psychoacoustic roughness or maximize harmonicity, and in that way to maximize consonance or minimize dissonance.
Regarding the difference between arrangements of songs composed in the 1930s and the 1950s, we made no predictions. Nor did we predict a difference between the algorithm’s ability to predict chord profiles in European classical music and African American jazz, despite the large stylistic differences. In all such cases we expected correlation coefficients across 12 the elements of a profile in the range .8 to .95 (cf. Parncutt, 2011).
Results and discussion
We first counted the number of pitch classes in each identified chord; the distribution of the results is shown in Figure 3. The figure suggests that most identified “chords” comprised only one pitch class. The figure shows further that the number of identified chords with two pitch classes was about the same as the number with three or four pitch classes. In fact, from informal listening, most chords had three or four pitch classes. Our method underestimated the number of pitch classes in a chord, because only onsets were counted and held tones were ignored.

Distribution of the number of pitch classes in each chord in the derived chord progressions.
Figure 4 shows how often each of the analyzed chords was identified, either alone or as part of a chord of higher cardinality. Major triads happened more often than minor (here and in the following, statistical tests are unnecessary, given the very large N), but the difference was small, and smaller than in the classical repertoire according to Eberlein (1994). A possible reason is that, unlike Eberlein, we counted major triads when they were part of tetrads, such as major and minor seventh chords. Minor seventh chords (037T, where T = 10; equivalent to a major added sixth chord 0479 in inversion) happened more often than major–minor sevenths (047T), which in turn were more prevalent than major seventh chords (047L, where L = 11). 1

Chord distribution: how often each of the tested chords occurred in the derived chord progressions. Chord labels are in semitones relative to the conventional root (T = 10, L = 11).
There were some interesting differences between arrangements of songs composed in the 1930s and songs composed in the 1950s. Diminished seventh chords (0369) and diminished triads (036) happened more often in 1930s songs, whereas suspended triads (057), major seventh chords (047L), and suspended seventh chords (057T) were more common in the 1950s songs. A possible reason is a move away from classical tonal clichés, such as dominant seventh to tonic progressions; there is still a tendency to avoid these recognizable classical markers in today’s pop (Ford, 2017). There was also an increasing tendency was to treat suspended triads (or chords created from stacking fourths rather than thirds) as chords in their own right (cf. Tagg, 2009), whereas in classical music suspensions almost always resolve to triads. Finally, the idea of major seventh chords as stable, consonant sonorities emerged in 20th-century popular music and jazz, again contradicting classical practice (McGowan, 2008).
Our analyses focused on 10 chord types: four triads (major, minor, diminished, and suspended) and six seventh chords (major–minor, minor, major, half-diminished, diminished, and suspended). Other chords were ignored; the number of rejected chords in the 1930s sample was 2,037 or 7%, and in the 1950s, 1,723 or 6%. We analyzed the chosen 10 chord types in two different ways. First, we focused on chords that included only the specified tones and no others. Second, we looked at all chords that included the chords in question. In the second case, a C-major seventh chord CEGB was counted three times: once as a C-major triad CEG (ignoring the B), once as an E-minor triad EGB (ignoring the C), and once as a C-major seventh chord CEGB. In each case, we counted how often each tone in the chromatic scale was voiced in the bass, relative to the chord’s conventional root, to test the prediction that roots were more often bass tones. We also counted how often each tone was doubled (voiced in more than one octave) to test the prediction that predicted roots are more often doubled.
Pearson’s correlation coefficients for comparison of data with predictions of Parncutt (1988, Fig. 1) are shown in Table 1. Pearson’s correlation coefficients are appropriate since the usual conditions (continuous, paired, independent, linear, normal, homoscedastic, and no outliers) are satisfied or nearly satisfied. Regarding continuity, note counts are large enough to be quasi-continuous. The Kolmogorov–Smirnov test did not reveal departures from normality for individual 12-element profiles. While the variation among correlation coefficients in the table is difficult to explain, coefficients did tend to be higher for more consonant or familiar chords. Data for 1930s and 1950s songs correlated strongly with each other, suggesting that the differences can often be disregarded.
Pearson’s Sentence case not title case, please: “Pearson’s correlation coefficients between predictions and findings for triads, calculated by comparing two vectors of 12 values each (10 degrees of freedom) (T=10, L=11).
Bass distributions for triads are shown in Figure 5. From visual inspection, the prediction that the root of a major triad 047 can be a major sixth (nine semitones) above the conventional root was confirmed (e.g., the root of CEG can be A). Also confirmed was the prediction that the root of the minor triad 037 is ambiguous; beyond the main two roots (0 and 3 semitones) the root can be a perfect fourth (five semitones) or a minor sixth (eight semitones) above the conventional root. The prediction that the root of a diminished triad 036 is often a minor sixth (eight semitones) above the conventional root was confirmed, as was the prediction that the root of the suspended triad often corresponds to the suspended tone (a perfect fourth above the conventional root).

Bass distributions for triads in the database. Data are shown for all chords that included the marked triads. For example, in part (a), the values for Pitch Class 9 (a major sixth interval) are for minor seventh chords in root position, considered as major triads with an added major sixth in the bass. (a) Major triad (047; N30 = 5,846; N50 = 5,939), (b) minor triad (037; N30 = 5,387; N50 = 5,448), (c) diminished triad (036; N30 = 4578; N50 = 3515), and (d) suspended triad (057; N30 = 3,272; N50 = 4,143).
Regarding bass distributions of tetrads (Figure 6), the prediction that the major–minor seventh chord (047T) has the least ambiguous root (0) is confirmed, as is the root-ambiguity of the minor seventh chord (the root of 037T can be 0 or 3). As predicted, the root of the major seventh chord (047L) is often a major sixth (nine semitones) above the conventional root, and the root of the half-diminished seventh (036T) is often a minor sixth (eight semitones) above the conventional root. The root of a diminished seventh chord (0369) is often a semitone below one of the tones (2, 5, 8, or L).

Bass distributions for tetrads: (a) major–minor seventh chord (047T; N30 = 1,735; N50 = 1,699), (b) minor seventh chord (037T; N30 = 2,523; N50 = 2,883), (c) major seventh chord (047L; N30 = 950; N50 = 1398), (d) half-diminished seventh chord (036T; N30 = 1,270; N50 = 1,252), (e) diminished seventh chord (0369; N30 = 1,832; N50 = 1,052), and (f) suspended seventh chord (057T; N30 = 1,171; N50 = 1,732).
In chord voicings, doublings refer to pitch classes that appear in more than one octave register. Figure 7 suggests firstly that doubling was generally less common in arrangements of songs from the 1950s than the 1930s. That is presumably because harmony had become more complex in the 1950s: the number of chords comprising five, six, or seven pitch classes had increased, as shown in Figure 3. For specific chords, our analysis shows that all tones were often doubled; the differences between doubling rates for different tones within a chord were smaller than expected from music theory texts. Regarding the major triad, music theory texts tend to recommend doubling the root, and our data are consistent with that principle in both the 1930s and the 1950s sample. In the minor triad, the root and the third were doubled almost equally often—consistent with model predictions, and also the theoretical idea that the third of the minor triad can be doubled more often than the third of the major. Results for the diminished and suspended triads were also similar to predictions.

Doubling distributions for triads: how often each pitch class in each triad was repeated in different octave registers, on average. The axis marking “1” means the pitch class appeared in only one octave register, even if that pitch was played by more than one instrument. The marking “1.5 repetitions” means the pitch class was doubled in 50% of cases. (a) major triad (047; N30 = 5,846; N50 = 5,939), (b) minor triad (037; N30 = 5,387; N50 = 5,448), (c) diminished triad (036; N30 = 4,578; N50 = 3,515), and (d) suspended triad (057; N30 = 3,272; N50 = 4,143).
Whereas these effects were generally small, all were consistent with predictions based on Figure 1. The same cannot be said for tetrads, as shown in Figure 8. An exception is the major–minor seventh chord (047T), whose profile had the clearest peak—consistent with the prediction that, of all tetrads, the major–minor seventh has the clearest root.

Doubling distributions for tetrads: (a) major–minor seventh chord (047T; N30 = 1,735; N50 = 1,699), (b) minor seventh chord (037T; N30 = 2,523; N50 = 2,883), (c) major seventh chord (047L; N30 = 950; N50 = 1,398), (d) half-diminished seventh chord (036T; N30 = 1,270; N50 = 1,252), (e) diminished seventh chord (0369; N30 = 1,832; N50 = 1,052), and (f) suspended seventh chord (057T; N30 = 1,171; N50 = 1,732).
Figure 9 shows which pitch classes were most commonly added to triads to increase their cardinality, that is, to create tetrads, pentads, and so on. When searching the chord progressions derived from the MIDI files, we identified each chord type (e.g., major triad 047) whenever it occurred within a chord, regardless of the chord’s cardinality (number of pitch classes). The graphs show distributions of the additional tones in each case. In general, the results suggest that a tone was most often added to a chord if it minimized the increase in dissonance—assuming that consonance has two main perceptual aspects, smoothness and harmonicity.

Pitch classes added to triads to create chords of higher cardinality: (a) major triad (047; N30 = 5,846; N50 = 5,939), (b) minor triad (037; N30 = 5,387; N50 = 5,448), (c) diminished triad (036; N30 = 4,578; N50 = 3,515), and (d) suspended triad (057; N30 = 3,272; N50 = 4,143).
The tone most often added to a major triad was a major sixth above the conventional root, presumably because that tone minimized roughness relative to other additional possible tones. Second most likely was a minor seventh; the dissonance in that case was greater due to the tritone interval between the third and seventh of the resultant major–minor seventh chord. The results for other triads can be explained similarly. The tone most commonly added to the diminished triad was a minor sixth above the conventional root, creating a major–minor seventh chord (high harmonicity). The pattern of extra tones suggests that this chord was perceived as more consonant than the less harmonic (and also less rough) diminished seventh chord, created when a diminished seventh interval (equivalent to a major sixth) was added.
Distributions of tones added to tetrads to create chords of higher cardinality are shown in Figure 10. Again, the results suggest that more consonant chords were preferred over less consonant, assuming that consonance is a combination of harmonicity and smoothness. The tones most commonly added to a major–minor seventh chord were the major second (creating a regular ninth chord) and the major sixth (sometimes called a 13th). For the minor seventh chord, the most common additional tone was the perfect fourth. That could happen in either the melody or the bass; the latter creates a familiar V11 (dominant 11th) or IV/V chord. For the major seventh chord, the major second (creating a maj9 chord) and major sixth (creating a min9 chord) were preferred. For the half-diminished seventh chord, the missing root at the minor sixth above the bass was the additional tone of choice. For the diminished seventh chord, a root could be created a semitone below any tone, and the same tones could happen melodically (creating a diminished scale). For the seventh chord with suspended fourth (7sus4), the most common additional tones were a major second or minor third, both of which avoided rough semitone intervals within the chord.

Pitch classes added to tetrads to create chords of higher cardinality: (a) major–minor seventh chord (047T; N30 = 1,735; N50 = 1,699), (b) minor seventh chord (037T; N30 = 2,523; N50 = 2,883), (c) major seventh chord (047L; N30 = 950; N50 = 1,398), (d) half-diminished seventh chord (036T; N30 = 1,270; N50 = 1,252), (e) diminished seventh chord (0369; N30 = 1,832; N50 = 1,052), and (f) suspended seventh chord (057T; N30 = 1,171; N50 = 1,732).
Conclusion
Our main finding is that, in chords with clear roots as predicted by a psychoacoustic model and in typical jazz arrangements, the conventional root is most often in the bass (i.e., the chord is most often in root position). That was true for a diverse set of 10 familiar chords (four triads and six tetrads). When the root was ambiguous, the psychoacoustic model successfully predicted both the degree of ambiguity and the specific alternative roots (tones that also appeared in the bass, but less often). The root was also the tone that was most often doubled, but that effect was weaker.
Interesting details include the following. The root of the minor triad was more ambiguous than that of the major (the alternative root being the third of the minor triad). Of the various seventh chords, the major–minor seventh had the clearest root. (Consistent with that observation, music theory texts advise that the function of major–minor seventh chords is not changed by inversion; consequently, they can appear in any inversion without changing the root, provided all tones are included.) When tones were added to existing chords, they tended to maximize consonance (harmonicity and smoothness).
The results confirm the psychological reality of chord roots more clearly than previous studies, perhaps because of the important role played by chord roots in jazz harmony. The results suggest that jazz musicians are sensitive to the relative stability or ambiguity of chord roots as predicted by Parncutt (1988) and confirmed by Parncutt et al. (2019). The composers of these songs, and the arrangers of these tracks, tend to voice roots more often in the bass. Moreover, they do so more consistently when the root is clearer or less ambiguous. In doing so, they are presumably guided by a combination of explicit knowledge of music theory and implicit understanding of how chords work in musical contexts. The latter process, which historically depends on the former, is consistent with the claim that chord roots are not only music-theoretic, but also psychological (or psychoacoustic) phenomena.
The term “psychological reality” can be interpreted in different ways. One refers to subjective experience. If something is demonstrably experienced by a conscious observer, it is psychologically real in that sense. To give an acoustical example: if a partial (pure tone component) within a complex sound is audible (i.e., not completely masked by other sounds), it can be individually perceived (as a spectral pitch) if attention is drawn to it. Even if it is not consciously perceived, it influences conscious perception in other ways (e.g., timbre; Stoll, 1982). That makes it psychologically real. Virtual pitches, if sufficiently salient, may also be considered psychologically real in that sense, whether or not they correspond to chord roots. “Psychological reality” also refers to a person’s explicit or intuitive music-theoretic knowledge, just as it refers to a speaker’s or a listener’s intuitive knowledge of grammar in psycholinguistics (Halle et al., 1978). Our analyses suggest that the musicians who arranged the jazz standards in our sample had explicit or implicit knowledge of chord roots that corresponded to our model predictions.
Results are broadly consistent with predictions based on Terhardt’s virtual pitch theory, which successfully predicts not only the main root of a chord and other root candidates, but also their relative salience—that is, a profile of all root candidates for each chord. The model not only accounts for the music-theoretical observation that the root of a chord is often ambiguous—it also predicts the degree of ambiguity of different chords, and the results are consistent with objective measures of chord ambiguity based on representative database analysis.
A theory of this kind has the potential to enrich music theory (including pitch-class set analysis) and jazz pedagogy. In an approach to tonal harmony based on pitch-class sets, students might learn that any simultaneous combination of tones in the chromatic scale has the potential to become a “chord,” that is, to be recognized by musicians and listeners as an acceptable way to combine tones, creating its own expectations for usage in specific contexts, voicing, resolution, and so on.
Each tone simultaneity implies roots that can be predicted from three assumptions. First, roots are salient pitches that act as perceptual references. Second, salient pitches usually correspond approximately to fundamentals of incomplete harmonic series. Third, those incomplete series usually comprise lower harmonics (Harmonic Number 10 or lower), because only lower harmonics are usually audible in speech sounds.
Our results offer an alternative account of chord “construction.” It is convenient to label chords according to the principle of stacked thirds, and it is instructive to imagine their tones representing the harmonic series. But our results suggest that neither stacked thirds nor the harmonic series (as normally understood) was the main principle behind the psychohistoric origin/emergence of familiar chords of three, four, or more pitch classes. Instead, the main principles were experimentation and consonance. Composers and performers experimented with harmony by adding tones to existing, familiar chords to create new ones (often following existing voice-leading conventions). Gradually, certain choices predominated over others. The tones were perceived to go together musically if the chord was psychoacoustically smooth, harmonic, and/or familiar. In the history of Western harmony—including later jazz harmony—processes of experimentation and familiarization with new sonorities took decades or centuries, and resulted in long-term preferences for certain tone combinations over others.
