Abstract
This study focuses on prosodic evolution in the French news announcer style, based on acoustic and perceptual analysis of French audiovisual archives. A 10-hour corpus covering six decades of broadcast news is investigated automatically. Two prosodic features, which may give an impression of emphatic style, are explored: word-initial stress and penultimate vowel lengthening, especially before a pause. Objective measurements suggest that the following features have decreased since the 40s: mean pitch, pitch rise associated with initial stress, vowel duration characterizing an emphatic initial stress, and prepausal penultimate lengthening. The onsets of stressed initial syllables have become longer while speech rate (measured at the phonemic level) has not changed. This puzzling outcome raises interesting questions for research on French prosody, suggesting that the durational correlates of word-initial stress have changed over time, in the French news announcer style.
Three perceptual experiments were conducted using prosody transplantation (copy of fundamental frequency and duration parameters on a synthetic voice), delexicalization and imitation. Rather than manipulating the parameters of, say, word-initial stress, we selected a subset of the corpus to represent the different decades under investigation. Results show that, among other factors, fundamental frequency and duration correlates of prosody contribute to distinguishing early recordings from more recent ones. The higher the pitch and the greater the pitch movements associated with word-initial stress, the more the speech samples are perceived as dating back to the 40s or 50s.
1 Introduction
Often we are able to recognize a recording made decades ago, distinguishing it from a contemporary one. Technical conditions (e.g., sound storage, type of microphone, distance of speaker from microphone) have evolved. They are in part responsible for a peculiar voice quality which is readily perceivable and can easily be caricatured. However, analysis of voice quality patterns (e.g., a pressed voice vs. a relaxed voice) is currently impossible when the signal does not have high-quality sound, and is even difficult with high-quality sound (d’Alessandro, 2006). Other linguistic features, in particular prosody, may also have undergone changes since the 1940s. Focusing on French, we investigate some of the parameters which may enable us to characterize a news announcer style of previous decades and contrast it with a current news announcer style. We now have archives spanning over 50 years of French broadcast news which make such an investigation possible.
A question we investigate is to what extent we are able to quantify and perceive variation in prosody between Gaumont-Pathé cinematographic news (named after a French film production company) dating back to World War II and that of radio or television news recorded half a century later. The communication situation, including recording conditions, has changed: radio and television have become commonplace, news is watched and/or listened to at home rather than in cinemas. Given these social and technological changes – the particular medium used, the location of both speakers and addressees, the topics approached, etc. – the style of speaking, defined as an adaptation to the communication situation (Bolinger, 1989; Eskénazi, 1993), should also have changed.
In the present study, the French broadcast news style prosody is addressed from a twofold perspective, through objective measurements and perceptual experiments. To our knowledge, few publications are devoted to the diachronic evolution of prosody, even though phonetic changes and their perception are of particular interest for phonostylistics and linguistics (Léon, 1993). Scholars rather examine the diaphasic dimension of prosodic variation across speaking styles (Fónagy & Fónagy, 1976; Fónagy, 1989; Vihanta, 1991, 1993; Léon, 1993; Astésano, 1998, 2001; Oakes, 2002; Goldman, Auchlin, Simon, & Avanzi, 2007). In particular, a tendency toward “barytonic” (or initial) stress has been observed in radio/television journalists, which would make their “jerky” speaking style recognizable even in filtered speech (Fónagy & Fónagy, 1976). For decades, this prosodic element, functioning like a socioprofessional marker, has been one of the most often cited characteristics of the broadcast news style. Before addressing this phenomenon in more detail and before relating it to possible stylistic changes, it is necessary to provide a brief overview of the French prosodic system.
Stress in French differs from lexical stress in the other Romance languages or in English, German, etc. Though French is traditionally said to possess a phrase-final stress, numerous theoretical, experimental and applied studies have highlighted the emergence of a word-initial stress in contemporary French (Carton, Marchal, Hirst, & Séguinot, 1977; Fónagy & Léon, 1980; Lucci, 1983; Martin, 1987; Pasdeloup, 1990; Fant, Kruckenberg, & Nord, 1991; Mertens, 1993; Di Cristo & Hirst, 1993; Delais, 1994). According to these studies, this initial stress is complementary to the final one, which remains a major property of the French accentual system, but is striking particularly in journalistic and didactic styles, in broadcast news, public conferences and classrooms where one has to be convincing (Lucci, 1983; Fónagy, 1989; Carton, 2000). It would date back to the late 19th century and even earlier (Di Cristo, 1999a), even though its origin is difficult to trace precisely. Consequently, recent accounts of French prosody have integrated the coexistence of primary (final) and secondary (initial) stresses into phonological models (Di Cristo, 1999a; Lacheret-Dujour & Beaugendre, 1999; Rossi, 1999; Jun & Fougeron, 2000, 2002; Post, 2000; Welby, 2006; Astésano, Bard, & Turk, 2007). They hypothesize a double marking of content words or phrases by an initial stress and a final stress, the former being essentially melodic (but also associated with onset lengthening) and the latter characterized by lengthening – as well as pitch rise when it is not utterance-final (Rossi, 1999; Astésano, 2001). Researchers may refer to this word-initial stress as a secondary accent or ictus (Pasdeloup, 1990; Rossi, 1999), an initial accent (Di Cristo, 1999b; Astésano et al., 2007), a word-initial or non-final accent (Post, 2000), an initial stress (Jun & Fougeron, 2002) or an early rise (Welby, 2006). The term “initial stress” will be used during this study, following Jun and Fougeron (2002). Its phonological status as a pitch accent (Post, 2000) or an edge tone (Jun & Fougeron, 2002; Welby, 2006), which constitutes an important part of the debate in French phonology in the framework of the autosegmental-metrical approach, is not discussed in this paper.
In some contexts, the underlying initial stress may be realized at the surface level as an emphatic stress (accent d’insistance or accent de focalisation), with more dynamic pitch patterns – it is usually realized by an abrupt pitch rise (Touati, 1987) – and additional lengthening associated with the prominent syllable onset and nucleus (Astésano, 2001). Whereas the emphatic stress has a proper pragmatic or paralinguistic function, consisting in giving a particular importance to certain words, the non-emphatic initial stress has a rhythmic function motivated by eurhythmic constraints favoring well-balanced stress alternations (Pasdeloup, 1990; Di Cristo & Hirst, 1993; Jankowski, Astésano, & Di Cristo, 1999): this device avoids over-long stretches of unstressed syllables. The non-emphatic initial stress may also have a demarcative function which is less well known: under certain conditions, it enables syntactically ambiguous sentences to be disambiguated (Astésano et al., 2007). However, with rare exceptions such as Astésano’s (2001) work, phonetic descriptions seldom propose clear acoustic correlates which would allow word-initial stressed syllables to be distinguished satisfactorily from their unstressed counterparts without resorting to interpretation. Admittedly, it may be listener- as well as speaker-dependent (Vaissière, 1983; Dahan & Bernard, 1996; Gendrot, 2006), and many influential factors call for further investigation. Even when an initial stress is unanimously perceived by a wide range of listeners, its functional interpretation is not always constant (Vaissière, 1997a).
Initial stress descriptions in the literature were essentially made through analyses of laboratory corpora (Jun & Fougeron, 2000, 2002; Welby, 2006; Astésano et al., 2007), sometimes under strong constraints upon the voiced nature of all speech segments (Welby, 2006) and always were done by hand-labeling. The study of diachronic variation requires the use of a large amount of data, namely archives which cannot be built up to fit specific requirements such as having continuously voiced segments. Large spoken corpora are now available, even though the workload their detailed manual analysis represents generally prevents prosodists from relying on such resources. Some exceptions can be found (Astésano, 2001; Mertens, 2004; Gendrot, 2006; Adda-Decker, 2006 inter alia) for the French language, with the most recent publications valuably making use of automatic speech processing tools. Because of the amount of data, the work presented in this paper also uses automatic speech processing tools, since the hours required for hand-labeling word-initial stress, for instance, seemed prohibitive.
Relevant criteria have to be found in order to study prosody-related variation cues in large-scale archives, and more particularly the presence and the putative evolution of word-initial stress. Following criteria proposed by ’t Hart, Collier, and Cohen (1991) especially, our own approach to initial stress is mainly concerned with clitic–nonclitic sequences such as un FAbricant de MAtériaux de CONstruction (‘a maker of building materials’, example borrowed from Di Cristo, 1998). Such a sequence of function–content words – a more precise definition will be given below – corresponds to the most frequent word chunk in French (Boula de Mareüil, d’Alessandro, Beaugendre, & Lacheret-Dujour, 2001), and constitutes a good candidate for initial stress on the polysyllabic nonclitic word (i.e., on the uppercase syllables in the previous example). The early rise on the first (or the second) syllable of the nonclitic (content) word is treated as a bitonal (LH-) phrase accent by Welby (2006, 2007), who also showed its use in speech segmentation. We considered clitic–nonclitic sequences in an extensive way, in an attempt to find out prosodic differences over the years. This empirical study is first intended to evaluate, through broadcast news archives dating from 1940 to 1997, whether or not initial stress is increasingly used in journalists’ style. On the one hand, an emphatic way of speaking – influenced by reading and therefore writing – is typical of old recordings (Méadel, 1994), suggesting a Decrease Hypothesis: word-initial stress loses ground. On the other hand, as a marker of professional identity, initial stress contributes to portraying today’s journalists and political figures by contrast to casual speech (e.g., Fónagy & Fónagy, 1976; Oakes, 2002), suggesting an Increase Hypothesis: word-initial stress gains ground. Our goal is to verify if a stylistic change is in progress (with an increasing or decreasing tendency to initial stress or other features) and to exhibit relevant pronunciation traits discriminating broadcast news speaking styles.
A second prosodic feature is also analyzed: the lengthening of the penultimate vowel preceding a pause. In that context, it may be particularly salient in some archaic ways of speaking (Carton et al., 1977; Léon, 1993), as it is in Quebec French (Thibault & Ouellet, 1996) and Belgium French (Hambye & Simon, 2004; Woehrling, Boula de Mareüil, Adda-Decker, & Lamel, 2008). It may be another ingredient of the emphatic style impression, even if it is not emblematic of the news announcer style. It is interesting to see if we can measure an evolutionary change in that respect, in particular as far as intrinsically long vowels (such as nasal vowels) are concerned. The Decrease Hypothesis is expected to be validated for this feature, which is now obsolete in “standard” French.
Whether or not we are capable of measuring a given phenomenon, listeners may be able to perceive differences to some extent. Perception, whose link with production and acoustics is known to be complex, is the focus of the second half of this paper. Experiments were conducted, based on prosody transplantation (copy synthesis of pitch and duration parameters), varying the content and overall prosody of the utterances.
Description of the corpus and the method is given in the next section (section 2). Section 3 deals with prosody-related acoustic measurements: (a) word-initial stress through pitch rise and nucleus/onset lengthening, and (b) penultimate lengthening. Section 4 explores the perception of the evolution of the French broadcast news style. The experimental setup using synthetic speech, delexicalization and imitation is presented along with acoustic analyses of the listening test corpus, and perceptual results are reported. Section 5 is the conclusion.
2 Corpus and method for the acoustic analyses
The corpus consists of about 10 hours of speech collected within the framework of the
In order to make sure that the data under investigation present comparable contexts over the years, the entire corpus was listened to by the authors. Speakers are mainly adult male journalists, matched in age and dialect background across the different decades. Therefore, despite obvious content changes, the data seem to consistently portray the targeted journalistic speaking style. Moreover, since we have at our disposal the video of the audiovisual archives, we wondered if the fact that journalists appear on screen might affect their way of speaking. But it turns out that, apart from rare exceptions (such as 1964, 1982 or 1985), the news announcer’s face shows up in few sequences of the corpus. Most often the video is in “voice-over” and does not allow us to draw conclusions in that matter.
Documents were orthographically transcribed by hand and segmented into phonemes by automatic alignment using extensively-trained context-independent acoustic models with Gaussian mixtures (256 Gaussians per state, for each phoneme) and a specifically-tuned pronunciation lexicon – the corpus contains over 9500 different words. The method illustrated in Figure 1, described in Adda-Decker Boula de Mareüil, Adda, and Lamel (2005), has been validated in several publications since then (Gendrot & Adda-Decker, 2005; Adda-Decker, 2006; Woehrling et al., 2008). Pitch values were then assigned to each phoneme by averaging fundamental frequency (F0) measurements taken every 10 ms by the PRAAT software (Boersma, 2001) with standard settings – in which F0 values below 75 Hz appear as undefined. This led to a rather raw but readily readable representation of prosody in which each phoneme is defined by its duration, mean pitch, the word it belongs to, and other information for further processing. In particular, formant estimation was performed, by using PRAAT as in Gendrot and Adda-Decker (2005). In addition to the local tonal configurations, different melody stylizations are possible for a closer look at the shape of the F0 curve (’t Hart et al., 1991; Mertens, 2004), but they were not applied here. In an analogous prosodic approach to regional French accents (Woehrling et al., 2008), F0 measurements taken every 10 ms were split into two, for the first half and the second half of each vowel. Very few occurrences of within-phoneme pitch variation greater than 1 or 2 semitones were observed. Therefore, a representation based on static pitch targets was thought to be a good approximation.

Block diagram of the alignment procedure: from a speech signal and its orthographic transcription, given acoustic models as well as a pronunciation dictionary, the decoder provides the most likely sequence of phonemes.
In order to have balanced subsets of data in terms of time chunks, four periods are considered: 1940–1959, 1960–1969, 1970–1979 and 1980–1997. The 60s and the 70s are more represented than the other decades in our data. Table 1 shows the duration of each subset and the average phoneme duration which appears to be comparable between the retained periods. This does not prevent small differences (even differences of 2 ms) from being statistically significant, owing to the quantity of data manipulated. Mean values along with standard deviations are reported in Table 1 (and subsequent tables); however, detailed analyses of variance (ANOVAs) are only reported for the data in connection with the perceptual test results (section 4).
Total duration of the data, phoneme duration, standard deviation of the phoneme duration distribution, percentage of octave jumps between consecutive vowels, percentage of unvoiced vowels and mean pitch (for males).
In order to validate the quality of the prosodic measurements, we also computed the percentages of octave jumps between consecutive vowels and the percentages of vowels that were detected as unvoiced (both may be linked to mistracking of the F0), to estimate the potential impact of background noise on F0 measurements. These percentages are reassuringly low: they are reported in Table 1, as well as mean pitch. The latter is restricted to males owing to the under-representation of females in the archives (only once in the 40s and once in the 50s). Males’ pitch was higher in the 40s and 50s, and has regularly decreased since then. Our subjective impression on listening to the corpus is that the mean F0 difference does not correspond to a change in the aim of broadcast news (conveying sensationalism or sounding reliable), as could be accounted for by the so-called frequency code (Ohala, 1984).
On the other hand, the 1940–1959 vowel triangle is larger than the other ones (see Figure 2). The link between pitch and vowel triangle size is corroborated by previous work (Calliope, 1989; Benolken & Swanson, 1990). An interpretation is that speakers produced hyper-articulated speech in the 40s and 50s, with more vocal effort which resulted in higher pitch, more extreme formant frequencies, and changes in voice quality. The more reduced vowel space observed later on might result from a more hypo-articulated speech (de Jong, 1995; Lindblom, 1999; Hermes, Becker, Mücke, Baumann, & Grice, 2008). Both higher pitch and hyper-articulation may be due to the recording conditions which compelled news announcers to speak clearer and louder, when standing in front of their microphones. It does not necessarily mean that males used to speak with a higher pitch. The original recording support (record or tape) may impact parameters such as mean pitch. Also, it is interesting that this high pitch is found in imitations of that period of newscasts (see section 4.3). Moreover, local features such as word-initial stress and penultimate lengthening should not be affected by recording condition and technical support changes.

Vowel triangles of males for the 1940–1959 (straight line), 1960–1969 (dotted line), 1970–1979 (dashed line) and 1980–1997 (dashed-dotted line) periods.
3 Prosody-related acoustic measurements
3.1 Word-initial stress
As explained in the introduction, clitic–nonclitic word contexts (like un fabricant ‘a maker’) are good candidates to receive initial rises on nonclitic words. Such contexts are considered in studies through laboratory corpora on word-initial stress in French (Di Cristo, 1999a, 1999b, 2000; Jun & Fougeron, 2000, 2002; Welby, 2006; Astésano et al., 2007). From the study of a 10-minute corpus of non-laboratory speech, Astésano (2001) measured pitch rise and onset lengthening as acoustic correlates of initial stress and additionally nucleus lengthening as indications of emphatic initial stress. We here measure these cues in clitic–nonclitic word contexts, in our 10-hour corpus.
3.1.1 Pitch rise
A problem arises with the detection of F0 rises. Most studies use some degree of subjective hand-labeling to decide if a high tone is realized. If a 10% pitch threshold is chosen as in Astésano et al. (2007), another problem arises as to the location of the initial stress rise. According to the literature, the high peak appears on either the first or the second syllable of the nonclitic word. On the basis of continuously voiced consonant–vowel syllabic structures, Welby (2006) reports that the low tone is aligned more or less with the left edge of the accentual phrase, before the left edge of the content word (i.e., minimal F0 values are on the vowel of the clitic word preceding the nonclitic word), while the high tone appears somewhere around the boundary between the first and second syllables of the nonclitic word. In her study, almost all the high tones are observed either during the last 60 ms of the first syllable (i.e., on the first vowel of the nonclitic word) or during the first 80 ms of the second syllable (i.e., mainly on the consonant of the second syllable). The first vowel of the nonclitic word therefore seems to be a good choice for measuring a possible F0 initial rise if one cannot assume the presence of a following voiced consonant. Furthermore, in an archive corpus, precise measurements of “early L elbow” or high tone peak locations – as in Welby’s (2006) work – cannot be reliably done due to the presence of unvoiced segments and the difficulty of measuring reliably continuous pitch contours. As a consequence, here, we concentrate on the fundamental frequency difference (ΔF0) between the mean pitch of the first vowel of a nonclitic word and the mean pitch of the last vowel of the preceding clitic word. This difference is expressed in semitones in order to correspond with pitch perception. This robust measurement allows us to examine word-initial pitch rises for various thresholds, as exemplified in Table 2 and Figure 3.
Number of clitic–polysyllabic word contexts with the percentage of cases in which the pitch difference between the polysyllabic word-initial vowel and the clitic vowel is greater than 1, 2, 3 or 4 semitones.

ΔF0 distribution between polysyllabic word-initial vowels and clitic vowels (in semitones). The percentages of contexts for which the ΔF0 value ranges from -1 and 0 ST, 0 and 1 ST, etc. are provided.
Another question arises as to the definition of “clitic”. Generally speaking, a clitic is a word which is unstressed and bound with a content word (Mertens, 1993). Based on prior knowledge (Boula de Mareüil, d’Alessandro, et al., 2001), a first set of about 300 function words was established, including forms of auxiliary verbs such as être (‘to be’) and avoir (‘to have’). Based on the most frequent words of our corpus, we built up a second set of 30 function words: le, la, les (‘the’), un, une (‘a’), du, des (‘some’), de (‘of’), à, pour (‘to’), en, dans (‘in’), et (‘and’), que, qui (‘that’), est (‘is’), a (‘has’), il (‘he’), on, nous (‘we’), etc. The negations pas (‘not’) and plus (‘[no] more’), which are usually not considered as clitics (Fónagy & Fónagy, 1976; Mertens, 1993; Oakes, 2002), were excluded. The resulting number of clitic–nonclitic sequences was 23,000 with the first set, and 21,000 with the second set. Because the latter is more controlled and displays a wide coverage in terms of clitic–nonclitic word contexts, we used it. The drawback is that contiguous sequences such as ne plus (‘no more’) then fall into clitic–nonclitic sequences. It was therefore important to focus on clitic–polysyllabic word sequences, which additionally avoids the merger of initial and final stress in monosyllabic words. Table 2 reports the number of corresponding contexts, and the percentages of cases in which the pitch difference between the polysyllabic word-initial vowel and the clitic vowel is greater than 1, 2, 3 or 4 semitones (ST). The distribution with non-cumulated percentages (including ΔF0 values less than 1 ST) is depicted in Figure 3. A higher difference in absolute value leads to a lower percentage. However, for each F0 interval (e.g., 1–2 ST, 2–3 ST, 3–4 ST), the ranking between the four periods is the same. In Figure 3, it is noticeable that the more recent the document is, the lower the corresponding curve is on the right (positive) side and the higher it is on the left (negative) side.
In Table 2, the proportion of pitch differences between polysyllabic word-initial vowels and clitic vowels that are greater than 3 ST, for example, is over 25% in the 40s and 50s, and decreases with the later decades. Following ’t Hart et al. (1991) and more recent studies on French (Goldman, Avanzi, Simon, Lacheret, & Auchlin, 2007), the 3 ST threshold is assumed to be a good estimate of the acoustic correlates of prosodic prominence. According to this interpretation, more than one clitic–polysyllabic word sequence out of four gives rise to a prosodic prominence in the 40s and 50s, compared to less than one out of five in the 80s and 90s. Astésano et al. (2007) used a 10% (i.e., 1.6 ST) threshold between the peak at the beginning of the word and the preceding valley for determining the presence of an initial stress; Astésano (2001) found that the pitch rise is twice as large for the emphatic stress as compared to the non-emphatic stress. Applying these criteria to our data, the later in time, the fewer pitch patterns relating to both emphatic and non-emphatic initial stresses occur.
According to previous work (Jun & Fougeron, 2000; Welby, 2006; Astésano et al., 2007), the percentage of word-initial rise occurrences increases with the number of syllables of the nonclitic word, and this increase is mainly visible between disyllabic words and longer words. The difference between 3-syllable words and at least 4-syllable words concerns the peak position on the first or the second syllable (Jun & Fougeron, 2000) rather than the percentage of occurrences.
As expected, in our corpus, the proportion of ΔF0 greater than 3 ST slightly increases between disyllabic words and longer words (see the first columns of Table 3). This proportion decreases with passing decades independently of word length, while the proportions of disyllabic, trisyllabic and tetrasyllabic words have remained almost unchanged. This tendency of decreasing ΔF0 between a polysyllabic word-initial vowel and the preceding clitic vowel is also clear in the three graphs presented in Figure 4. This figure represents mean pitch contours: hence, it is not surprising that none of the curves exceed a 3 ST threshold. It shows that the mean initial rise on disyllabic words regularly decreases as time goes by, and reaches a lower level than the final syllable. This could be expected, since in the corpus there are fewer utterance-final contours than continuations. For words of more than two syllables, where the second syllable is not the final one, we observe a maximum rise for the first vowel (except for the longest words in the 1940–1959 period). This observation supports theoretical considerations as well as Welby’s (2006) results on the F0 peak location.
Percentages of disyllabic, trisyllabic and tetrasyllabic words as well as proper names (

Mean pitch contour (in semitones) from a clitic vowel (equated to zero to rule out mean pitch differences) to the second vowel of the following nonclitic word. Results are plotted for (a) 2-syllable words, (b) 3-syllable words and (c) at least 4-syllable words.
Proper names have been examined in further detail because they provide a good control for nonclitic words, in terms of whether they are considered as particularly marked by initial stress in the news announcer style (Oakes, 2002) or not (Fónagy & Fónagy, 1976). Their proportion among polysyllabic words in the studied contexts decreases as time goes by (see the % PN column of Table 3). More interestingly, proper names seem to receive more prosodic prominence in the news clips of the 40s and 50s than in the following decades (see the rightmost column of Table 3), even though they do not attract initial stress more than other words do. These results, with systematic tendencies concerning pitch differences, speak for a decrease of word-initial stress over time.
There is no clear tendency as to the lengthening of polysyllabic word-initial vowels with respect to the preceding clitic vowels. For instance, the percentages of duration differences that are greater than 20 ms – the just noticeable difference (JND) according to Bartkova and Sorin (1987) – are respectively 34%, 28%, 31% and 30% for the 1940–1959, 1960–1969, 1970–1979 and 1980–1997 periods. Nevertheless, other duration-related patterns may have evolved since World War II. This issue is addressed in the remainder of this section.
3.1.2 Stressed syllable nucleus/onset lengthening
The duration of polysyllabic word-initial onsets and vowels preceded by clitics was calculated in the same contexts as above, by applying syllabification rules proposed by Adda-Decker et al. (2005) and assessed by Bigi, Meunier, Nesterenko, and Bertrand (2010). A dozen rules have been implemented: they place a syllable boundary between two vowels, before an intervocalic consonant, before a consonant in a vowel–consonant–glide–vowel sequence, etc. For example, in de puissants ([dǝ pɥisɑ̃] ‘of powerful’), the onset is [pɥ], the vowel nuclei are [ǝ] and [i]; in les artistes ([le zaʁtist] ‘the artists’), the onset is the liaison consonant [z] which surfaces before the word-initial [a], the vowel nuclei are [e] and [a]. A syllable level was accordingly added to the automatic labeling, and the onset duration, reported to be a characteristic element of initial stress (e.g., Astésano, 2001), was measured.
In addition to raw duration measurements, normalized durations were computed following Campbell (1992, 1993). The z-normalization procedure which was applied allows for controlling the intrinsic and contextual variability of both vowels and consonants. It makes use of mean values and standard deviations which are calculated for each period under investigation with respect to 12 phoneme macro-classes, as proposed by Astésano (2001), inspired by Di Cristo (1985): unvoiced stops, voiced stops, nasal consonants, unvoiced fricatives, voiced fricatives, glides, liquids, closed vowels, mid-closed vowels, mid-open vowels, open vowels and nasal vowels. However, no a priori distinction is made between stressed and unstressed syllables, since this piece of information has to emerge from the data.
For the sake of readability and simplicity, since similar trends can be observed whether raw or normalized durations are taken into account, results are only reported for raw durations in Table 4. For all onsets and simple onsets (composed of only one consonant), mean values and standard deviations are tabulated, together with the numbers of contexts considered, which are lower than those in Table 2 because empty onsets have been disregarded. The onset lengthening is also expressed in the form of percent of onsets which are at least 20% longer than the average onset duration (here, 79 ms). Given the average durations measured, this 20% threshold (Rossi, 1972; Klatt, 1976; Bochner, Snell, & MacKenzie, 1988) is lower than the just noticeable difference considered to be 20 ms by Bartkova and Sorin (1987), but it is the one used by Astésano (2001). Referred to as % >JND, the proportion of onsets longer than 79 × 1.2 = 95 ms by applying the same reasoning was thus calculated, as well as the proportion of simple onsets longer than 88 ms.
Onset duration of polysyllabic words preceded by clitics: numbers of contexts, mean values of raw durations, standard deviations, and percentages of occurrences exceeding a given threshold (% > JND) corresponding to the Just Noticeable Difference (respectively 95 ms and 88 ms for all onsets and simple onsets).
By and large, Table 4 reveals that onset duration increases over the decades. The 10 ms increase and the increase of the % >JND value are even more regular when the analysis is restricted to simple onsets. As in section 2, all the differences are highly significant according to ANOVAs.
This onset lengthening over the years contradicts the tendency suggested in section 3.1.1. An opposite tendency should be posited if onset lengthening were a correlate of initial stress (Mertens, 1993; Jankowski et al., 1999; Astésano, 2001; Astésano et al., 2007). An alternative interpretation is that the relative importance of initial stress correlates may have changed for half a century. Indeed, a parallel evolution can be observed, with an increase of onset duration over time, when we only consider contexts in which the nonclitic initial vowel is at least 3 ST higher than the preceding clitic vowel. The % >JND values combining the >3 ST criterion also increase for both all onsets and simple onsets – in the latter case from 26% in the 40s and 50s to 50% in the 80s and 90s. The phonetic manifestations may have changed, as well as the communicative functions (Kohler & Niebuhr, 2007). Astésano (2001) proposed that an onset lengthening of 80% would characterize an emphatic type of initial stress. Yet, in our data, this criterion would result in only 3–6% of emphatic stress. According to Astésano (2001), vowel lengthening is also characteristic of emphatic stress. In our data, vowel duration and lengthening with respect to a 20% JND were measured in the post-clitic initial syllable of polysyllabic words. Results for each period under study are reported in Table 5.
Vowel duration in the initial syllable of polysyllabic words preceded by clitics: numbers of contexts, mean values of raw durations, standard deviations and percentages of occurrences exceeding a given threshold (% > JND) corresponding to the Just Noticeable Difference (86 ms).
Vowel duration in the initial syllable of polysyllabic words preceded by clitics decreases from 1940 to 1997 (see Table 5). The % >JND percentage of vowels longer than 1.2 times the mean value in that context (i.e., 86 ms) also decreases, even though it is a little higher in the 70s than in the 60s. Again, the effect of the period (1940–1959, 1960–1969, 1970–1979 or 1980–1997) is highly significant according to ANOVAs.
The increase of onset duration and the decrease of nucleus duration over the decades make the duration of stressed syllables stable. Several interpretations are thus possible. Our interpretation is that the initial stress used to be more prevalent in the 40s and 50s than in the later decades, supporting the Decrease Hypothesis. In particular, emphatic stress, which is characterized by lengthening of the syllabic nucleus according to Astésano (2001), has decreased. This might be explained by the poor equipment quality in the earlier decades, which speakers tried to compensate for by producing a larger vocal effort in order to convey their message. However, our data do not show a shift from an emphatic to a rhythmic-demarcative initial stress. Also, these two types of initial stress are difficult to differentiate functionally (Vaissière, 1997a; Oakes, 2002). Based on our acoustic analysis of pitch contours, both of them seem to have decreased over time. We shall return to this in the final discussion. Previous studies, without distinguishing emphatic from non-emphatic initial stress, have also shown percentages of initial stress which are comparable to ours (based on pitch): 33% in the journalistic style of the 70s (Fónagy & Fónagy, 1976), 29% in the current reading style (Woehrling et al., 2008).
3.2 Penultimate vowel lengthening
Scripts were written to compare the duration of the last two vowels or syllables of polysyllabic words (and the last three vowels or syllables of at least trisyllabic words). In particular, the percentages of penultimate vowels that are longer than final vowels were calculated. The final schwa was excluded because of the controversy about attaching the word-final syllable containing a pronounced schwa to the preceding syllable containing a full vowel (Durand & Eychenne, 2004). For example, if the mute e was pronounced in a word like pneumatique ([pnømatik(ǝ)] ‘tyre’), this word was not considered. This way, 13% of all occurrences were filtered out – in a manner balanced according to the different periods.
In a word like amitié ([amitje] ‘friendship’), for example, the duration of the vowel [i] was compared to that of the vowel [e]. At first glance, the distributions of penultimate–final vowel duration differences are very similar across the periods under investigation: percentages of positive differences are in the 5% range. Nonetheless, when the analysis is restricted to prepausal positions, the 1940–1959 patterns separate from the other ones. This leaves a large number of contexts, as shown in Tables 6 and 7, and this position triggers an impressionistically more salient effect. The inter-pause interval is 2.11 s in the 40s and 50s, 1.72 s in the 60s, 1.68 s in the 70s, and 1.67 s in the 80s and 90s.
Number of polysyllabic words preceding a pause, mean duration of penultimate vowels, standard deviation of the corresponding duration distribution and percentage of occurrences in which the penultimate vowel is longer than the final one – the right part of the table presents results for penultimate nasal vowels only.
Number of at least trisyllabic words preceding a pause, mean duration of antepenultimate, penultimate and final vowels, standard deviation of the corresponding duration distributions and percentage of occurrences in which the penultimate vowel is longer than the antepenultimate one.
Table 6 presents the mean duration of penultimate vowels, the standard deviation of the duration distribution and the percentage of words in which the penultimate vowel is longer than the final one, for each period: for all vowels in the left part and for penultimate nasal vowels in the right part. In the left part of this table, the average percentage is strikingly stable in the most recent recordings (18% from the 60s), but it goes up to 25% in the oldest recordings. A z-normalization of durations keeps these figures (almost) unchanged. Also, it is well known that French nasal vowels are intrinsically longer than are oral vowels, and we see this in our data (121 ms for nasal vowels vs. 87 ms for oral vowels on average). There is no quantity contrast within nasal vowels but within oral vowels there used to be phonological oppositions such as mettre (/mϥtʁ/ ‘to put’) vs. maître (/mϥːtʁ/ ‘master’). Such distinctions have become obsolete in more recent days (for the benefit of short vowels), which could partly account for the lengthening decrease. To investigate whether a more general prosodic change is in progress, we looked at penultimate nasal vowels in further detail. The right part of Table 6 (restricted to penultimate nasal vowels) shows fewer contexts and higher percentages than the left part (for all vowels). More importantly, the gap widens between the different periods. In the 40s and 50s, more than half of penultimate nasal vowels are extra-long – longer than final vowels in spite of prepausal lengthening. The decrease of mean duration (from 140 ms to 113 ms) is also noticeable. As described above for word-initial stress correlates, ANOVAs show a significant effect of the periods studied. Syllable-based patterns obtained by applying syllabification rules proposed by Adda-Decker et al. (2005) are similar.
No obvious tendency of decreasing or increasing duration of final vowels before pauses over the decades is observed (see the results for words of at least three syllables in Table 7). By contrast, the duration ratio between final and penultimate vowels has increased: 1.8 in the 40s and 50s, 2 in the 60s, 2.2 since then. On average, Delattre (1965, 1966a, 1966b) found a 1.8 duration ratio between stressed (i.e., final) and unstressed (i.e., non-final) syllables. The increase we observe here seems to be due to the decrease of the penultimate vowel duration throughout the decades.
We wondered about the behavior of antepenultimate vowels of at least trisyllabic words, even though there are too few prepausal contexts to break down the results for penultimate nasal vowels. Table 7 shows that the lengthening of penultimate vowels with respect to antepenultimate vowels has not decreased over time, because both penultimate and antepenultimate vowels have become shorter since the 40s, and again we see a significant effect of the periods studied according to ANOVAs. The latter finding is consistent with the decrease of post-clitic word-initial vowel duration discussed above (see Table 5). In most cases, the antepenultimate vowels are also the initial vowels of words with at least three syllables. This decrease of the prepausal penultimate vowel duration over time is also reflected by JND-based results obtained as in section 3.1.2 for initial stress durational correlates. To summarize, acoustic analysis suggests that the Decrease Hypothesis applies to both initial stress and penultimate lengthening.
4 Perception of the evolution of broadcast news style
The corpus-based study reported so far enabled us to quantify changes in the style of French broadcasting over the decades as decreases in mean pitch, word-initial stress (in a clitic–polysyllabic word context) and penultimate lengthening (before a pause). The purpose of this section is to ascertain whether the prosodic differences, as well as voice quality changes and other factors, are perceptible. In order to complete this task, three perceptual experiments using prosody transplantation were designed. This paradigm enables the F0 and duration correlates of prosody to be separated from recording conditions and voice quality effects. This method, even given its shortcomings, has been used to investigate prosody in various languages and situations (Jilka, 2000; Boula de Mareüil & Vieru-Dimulescu, 2006).
We selected a subset of the corpus utterances to represent each decade, and used prosody transplantation on a synthetic voice. A professional journalist was also asked to record a set of sentences from the oldest period (the 40s and 50s) in his contemporary style and imitating what he thought a newscaster production of this era could have been. On this basis, three experiments aimed to rate the relative importance of different speech dimensions (recording and voice quality, lexical content and prosody) in perceiving the changes in newscasters’ speaking style. We did not manipulate the parameters of word-initial stress, even though Jankowski et al. (1999) demonstrated that onset lengthening gave rise to perception of initial stress. Prosody transplantation was used in Experiment 1 to copy F0 and duration parameters of the original speech to the synthesized speech. Experiment 2 in addition used synthetic speech, generated by a text-to-speech (TTS) system, to vary the content and utterance prosody. Specifically, a procedure was used to “mask” the lexical content, which we refer to as “delexicalization”. For Experiment 3, prosody transplantation was done for broadcast news archives and the journalist we recorded. The methods for synthesizing the various speech sounds are described in more detail below.
4.1 Experiment 1
In Experiment 1 (as in Experiment 2), subjects were asked to assign a date (between 1940 and 1999) to the speech excerpt they listened to. The experiment consisted of two blocks: (1) synthetic stimuli, the prosody of which was copied from audio archive utterances, and (2) the original stimuli. With the prosody-copied stimuli, listeners had access to both lexical and prosodic information of the original stimuli, but not to the recording and voice quality characteristic of the original stimuli.
4.1.1 Corpus
For Experiment 1 (and Experiment 2), 30 utterances were selected among male speakers from the corpus described in section 2: 5 utterances per decade, 10-seconds long on average (see list in the Appendix). The utterances selected were those produced with the least number of lexical indices for the period, such as cultural references which could bias the results. To determine initial stress in this subcorpus, experts in prosody were asked to pinpoint prominent syllables, but no consensus emerged – which was unsurprising because many phoneticians made similar observations about the difficulty of agreeing about syllable prominence in French (Fónagy & Léon, 1980; Vaissière, 1997b). Thus, clitic–polysyllabic word sequences were considered, based on F0 measurements provided by PRAAT as described above. The pitch difference between the polysyllabic word-initial and the clitic vowels was computed, and the percentage of occurrences in which this difference is greater than 3 ST was assumed to be a good estimate of word-initial stress acoustic correlates. Comparative results for the experiment corpus (173 clitic–polysyllabic word contexts) and the whole corpus (12,158 contexts) are shown in Table 8. In both cases, a decrease of what can be interpreted as word-initial stress is observed. A similar decrease of mean pitch is noteworthy for both the experiment and the whole corpus: roughly from 170 Hz in the 40s and 50s to 140 Hz in the 80s and 90s. There were too few prepausal contexts to investigate penultimate lengthening before a pause.
Percentage of clitic–polysyllabic contexts in which the pitch rise is greater than 3 semitones.
The prosody transplantation method and the diphone speech synthesis system which were used with these stimuli are described in Boula de Mareüil, Célérier, et al. (2001), and Boula de Mareüil and Vieru-Dimulescu (2006). Given a speech file (the original), the transcription of what is said is utilized to build the sequence of diphones the original corresponds to. A diphone database is used, derived from a male speaker whose speech units are prestored (recorded for TTS purposes independently of the present study) and concatenated. Prosodic parameters are extracted from the original stimuli and grafted onto the corresponding diphone string. The resulting synthetic speech is time-aligned with the original signal using a Dynamic Time Warping (DTW) algorithm, as in Malfrère and Dutoit (1997). The F0 and duration parameters are then modified with the aid of the TD-PSOLA (Time Domain Pitch Synchronous Overlap and Add) algorithm (Moulines & Charpentier, 1990). Energy is not processed; the normalized level of the diphone database is retained.
4.1.2 Participants and task
Twenty-six subjects (18 males, 8 females, aged 34 on average) took part in Experiment 1. French was their mother tongue and they had no known hearing impairments. Prior to the test, they were requested to rate their ability to distinguish old from recent recordings on a 1–5 scale (from most likely unable to most likely able). The 26 subjects gave themselves an average rating of 3. They were not asked to explain their ratings.
After a practice test with some sample utterances other than those of the actual test, the subjects listened to 30 synthetic stimuli and then to the 30 original stimuli. In each block (synthetic or original stimuli), the stimuli were presented in random order (different for each subject). Participants could listen to each stimulus as often as they wished through a Web-based interface, but it was not possible to correct previous responses once a new stimulus was displayed. A slider allowed them to assign a date between 1940 and 1999 to each stimulus. The participants had to use a mouse to move this slider from the default position which was 1940.
4.1.3 Results
To analyze the results, we first considered the responses per decade. Partly owing to the difficulty of the task (judge the perceived date of a recording) and partly owing to the different sources of the stimuli, listeners’ answers show variability which cannot be accounted for if the results are described stimulus by stimulus. However, robust tendencies appear when stimuli are grouped. For each decade, a vector was created by computing the number of stimuli recognized by listeners as dating back to the 40s, the 50s, the 60s, the 70s, the 80s or the 90s. A hierarchical agglomerative clustering (using the Euclidean distance between vectors and the ward grouping method) was performed on the confusion matrices obtained for both the synthetic and the original stimuli. Results are presented in Figure 5. For the original stimuli (on the left), the 40s and 50s separate from the other decades, and the stimuli recorded during the 60s and 70s are regrouped. The 90s receive good recognition scores (perceived and real decades well-matched), while we see confusion with the 80s. In the prosody transplantation condition (on the right of Figure 5), the 40s and 50s also depart from the other decades, which are not as well-perceived. All in all, listeners seem to categorize the proposed utterances in three categories of 20 years each (hereafter referred to as epochs).

Hierarchical clustering resulting from the responses obtained per decade with the original stimuli on the left and prosody transplantation on the right (see text).
The following results were pooled for the 40s and 50s, the 60s and 70s, the 80s and 90s. In Figure 6, the x-axis stands for the averaged real dates and the y-axis stands for the averaged perceived dates. With prosody transplantation (synthesized speech), the perceived dates of the 40s and 50s are overestimated as compared to the original stimuli. In other words, the old-fashioned nature of these stimuli is better perceived when the characteristic voice quality and recording quality are heard.

Results of (a) Experiment 1 and (b) Experiment 2. The x-axis represents the real date (averaged for the 40s and 50s, the 60s and 70s, the 80s and 90s). The y-axis represents the averaged perceived dates for the original stimuli and prosody transplantation (Experiment 1), delexicalization and TTS prosody (Experiment 2).
To further compare the results for original stimuli and synthesized ones, an ANOVA (completely randomized two-factorial design) was carried out on the listeners’ answers (that is, the perceived date of each stimulus, expressed on a continuous 1940–1999 scale). The two fixed factors were (1) the 20 year Epoch of the recording (3 levels: 40s–50s, 60s–70s and 80s–90s) and (2) the Type of stimulus presented (2 levels corresponding to prosody transplantation and original). The significance level was set to 0.01.
Both factors have a significant effect: the perceived dates significantly increase with the Epoch, F(2, 1554) = 721.13, p < 0.01, and the mean judgments significantly differ between the two Types of stimuli, F(1, 1554) = 13.23, p < 0.01, mainly due to the more accurate rating of the oldest original stimuli which led to a lower perceived date for the original stimuli than for the ones with prosody transplantation. A significant interaction between the Epoch and the Type of stimulus was found, F(2, 1554) = 47.56, p < 0.01. The curves corresponding to the original stimuli and the synthesized stimuli show that the main difference between the two types of stimuli across the three epochs falls in the 40s–50s (see Figure 6a). For World War II or the post-war era, the original stimuli are well classified, whereas for the synthesized stimuli, the mean perceived date is 1960.
In order to assess whether the Epoch factor was significantly relevant for any Type of stimulus, a simple main effect analysis was performed separately for both the original and the synthesized stimuli. For the two kinds of stimuli, the Epoch factor has a significant effect at the 0.01 level, F(2, 777) = 837.8, p < 0.01 for the original stimuli, and F(2, 777) = 150.8, p < 0.01 for the synthesized stimuli. In both conditions, response patterns increase. Nevertheless, perception may have been biased by the informational content. In spite of efforts to select lexically-unmarked sequences, especially using reports on everyday life, news items often refer to topics which give indications to the listeners (e.g., war vs. agriculture or females’ status). Experiment 2 was intended to rule out this drawback.
4.2 Experiment 2
In a second experiment, the same 30 sentences as those of Experiment 1 were used, but modified using two different methods. The first one was based on “delexicalized” speech, rendered unintelligible as explained below; the second one was based on text-to-speech output, whose rule-based prosody presented similar variations for all sentences. In the delexicalized stimuli, listeners had access to the original prosody but not to the meaning of the sentences, while in the TTS stimuli, it was the opposite. This experiment therefore separates the prosodic and lexical information mixed in Experiment 1.
4.2.1 Corpus
The first set of stimuli was made up of the delexicalized versions of the prosody transplantation speech from Experiment 1, which retained the archive prosody. They were delexicalized in the following way: each vowel was replaced by an [a], each plosive by a [t], each fricative by an [s], each nasal by an [n], each glide by a [j] and each liquid by an [ʁ]. The latter was preferred over the [l] initially proposed by Ramus and Mehler (1999), who inspired this sartanaj mapping, because [tʁ] is a much more natural and frequent sequence than [tl]: there are 20 times as many [tʁ]s as [tl]s in our 10-hour corpus. An example of a sentence produced by such a technique is [ta tata atsasjaʁ ta nasa tʁatʁa] for the French du côté accessoires de nouveaux progrès (‘turning now to the topic of accessories, new advances’). This way of ensuring unintelligibility was also preferred over older methods like the reiteration of [ma ma ma] syllables because it is somewhat more ecological and preserves phonotactic rhythm patterns.
The second set of stimuli resulted from the output of a TTS system which synthesized the transcripts of the 30 original sentences with the same diphone voice as previously and a “neutral” prosody generated by phonotactic-syntactic rules (Boula de Mareüil, Célérier, et al., 2001). This TTS synthesis system produces a reading style prosody judged as a natural (“modern”) reading style – as natural as a human reading (Prudon, d’Alessandro, & Boula de Mareüil, 2004).
4.2.2 Participants and task
Twenty-nine subjects (22 males, 7 females, aged 31 on average, different from those of Experiment 1) took part in Experiment 2. All were native speakers of French with no known hearing problems. Their ability to distinguish old from recent recordings was self-estimated at 3 on a 1–5 scale, as in Experiment 1.
After a familiarization phase, the subjects listened to the 30 delexicalized stimuli and then to the 30 TTS stimuli. In each block (delexicalization and TTS stimuli), the stimuli were presented in random order. The protocol of Experiment 2 was the same as in Experiment 1: the task consisted of assigning a date to speech excerpts by using a slider.
4.2.3 Results
The results of Experiment 2 are shown in Figure 6b. The task was difficult with the delexicalized samples – this is a well-known issue with such tests (Ramus & Mehler, 1999). It was more difficult than the task in which subjects had access to the lexical content with the TTS prosody and the same diphone voice. These results suggest that content is more reliable than prosody alone to date a speech excerpt, even though for both conditions an upward slope is observed in Figure 6b. To analyze its statistical significance, results were analyzed as was done for Experiment 1.
An ANOVA (completely randomized two-factorial design) was performed on the listeners’ answers with two fixed factors: (1) the 20-year Epoch of the recording (3 levels: 40s–50s, 60s–70s and 80s–90s) and (2) the Type of stimulus presented (2 levels corresponding to delexicalization and TTS prosody). The significance level was set to 0.01.
The Epoch factor was found to have a significant effect: the perceived dates significantly increase with the Epoch, F(2, 1734) = 95.77, p < 0.01, whereas the factor Type of stimulus was not significant, F(1, 1734) = 0.78, Observed Power = 0.046. A significant interaction between the Epoch and the Type of stimulus was found, F(2, 1734) = 18.66, p < 0.01. This is mainly due to the fact that the recording date is better perceived on the basis of the lexical and thematic information than on the basis of the delexicalized prosody (see Figure 6b). In the former case, the slope of the perception curve is steeper from the 40s–50s to the 80s–90s than for the prosody-only stimuli.
To assess whether the Epoch factor was significant for any Type of stimulus, a simple main effect analysis was conducted separately for both the delexicalized and the TTS stimuli. For both, the Epoch factor has a significant effect at the 0.01 level, F(2, 867) = 13.1, p < 0.01 for the delexicalizations, and F(2, 867) = 120.8, p < 0.01 for the TTS stimuli. We may thus state that listeners do perceive time-related differences even on delexicalized stimuli. Even if the mean ratings do not reflect real dates, the prosodic information allows listeners to perceive significant changes across the epochs.
Comparing the results of Experiments 1 and 2, we can see that, with respect to the original stimuli, the perceived dates of prosody transplantations are slightly closer than are the perceived dates of the TTS prosody stimuli (e.g., 1981 with prosody transplantation vs. 1979 with TTS prosody) for the period spanning the 1980s and 1990s, while the average perceived date is 1984 for the original stimuli. This holds for each epoch, even if for the one which covers the 1960s and 1970s the perception of the TTS prosody stimuli (average date = 1971) is closer to reality (average date = 1968). A kind of default response (1970) may explain this behavior. Participants were not the same in Experiments 1 and 2, but this qualitative comparison of results further supports the contribution of prosody to the perception of stylistic changes.
The observations made about Experiments 1 and 2 show that the date perceived by the listeners consistently rises across the different epochs and that each type of information contributes to the listeners’ decision about the recording date: the recording and voice quality (presented together) are the most effective cues, lexical-informational content and then prosody are secondary cues. Despite the comparatively weak importance of prosody among the different parameters under investigation, it seems to play a role in the perception of the evolution of the newscaster style. However, the cognitive load required to process delexicalized speech may have biased the results. We therefore designed a third experiment to assess listeners’ abilities to distinguish a recent recording from an old recording solely on the basis of prosody.
4.3 Experiment 3
Experiment 3 concentrated on the prosody of the 40s and 50s, comparing it with the prosody of a contemporary journalist who read the same sentences (1) in his own style and (2) imitating the 40s–50s style. Impersonation has been used in various studies to investigate what is stable across speakers in intonation, which features are salient and which habits are hard to imitate (Zetterholm, 2003). In the present study, the term “imitation” is not used as a direct convergence between speakers (Holt & Nguyen, 2012) but in the following meaning: an imitator or an actor may speak an utterance by imitating a person, an accent or a style without having heard the utterance (just) before. In the reported experiment, the journalist was also asked to replicate a particular, prototypical way of speaking without being presented first with the audio stimuli. Prosody transplantation was used in order to remove the voice quality of the journalist, as well as the sound quality of the original stimuli. Therefore, only prosody (pitch and duration) differentiated the three types of stimuli.
A pairwise test was set up with the synthesized speech to compare original, imitation and modern styles: the sentences of each pair had the same semantic content but a different prosody (3 possibilities). Listeners were instructed to indicate if the presentation order was older–newer or newer–older. Could they be misled by an (amateur) imitator, who on the sole basis of prosody would succeed in sounding “older” than the originals? Since imitators are often prone to caricature and exaggerate some cues, a positive answer to that question would yield insights into what is important in the perception of an “old-fashioned prosody”.
4.3.1 Corpus
For this third experiment, 10 stimuli used for Experiments 1 and 2 were kept (the 5 from the 40s and the 5 from the 50s), and another 5 stimuli (marked with a star in the Appendix) were added from war archives. As explained above, a contemporary journalist was asked to read the 15 orthographic transcripts in his modern style and then to read them again, imitating the 1940–1950 style, not listening to the original excerpts but relying on his own representations of the Gaumont-Pathé style. Thus, the corpus was composed of 15 sentences × 3 styles: archive original (O), contemporary (C) and imitation (I). The prosody of these 45 stimuli was copied and imposed on a synthetic voice (the same diphone voice as in Experiments 1 and 2), because it would have been too easy for listeners to identify the voice quality of the actual recordings. From the resulting copy synthesis speech, 45 pairs of stimuli were constructed, concatenating O, C and I sentences: the order of appearance could be OC or CO, OI or IO, IC or CI, with a 1-second pause in between. (This inter-stimulus interval was felt sufficient because the content was the same for the two stimuli of each pair.) The average duration of a stimulus pair was 18 seconds.
An acoustic analysis of the imitation corpus was performed. Interestingly, the journalist spoke louder while imitating the Gaumont-Pathé style, but energy is not taken into account in the prosody transplantation method, contrary to pitch and duration. The journalist’s mean pitch rises from 142 Hz in his contemporary style up to 173 Hz in imitation, while the mean pitch of the original stimuli is 171 Hz. The percentage of clitic–polysyllabic word contexts exhibiting a pitch rise greater than 3 ST is also higher in imitation (37%) than in the contemporary style (28%). This percentage is 47% on the 15 original stimuli selected in the 40s and 50s. These figures are consistent with the decrease of initial stress over the years observed in Table 8. Table 9 provides comparative examples such as the word sequence du tournoi (‘of the tournament’), where the prominence is more marked in the I-style than in the C-style, with a difference greater than 3 ST.
Examples of clitic–polysyllabic word contexts with the associated ΔF0 in the I-style, the C-style and the O-style.
A closer look at accentual patterns enables us to better understand these mean and local pitch differences. Before a pause especially, initial stresses are much more marked in the imitation than in the contemporary style. An illustration is given in Figure 7, for the utterance-final phrase un vif succès (‘a great success’), where pitch rises up to 250 Hz in the imitation style, whereas it does not exceed 190 Hz in the contemporary style. It is reminiscent of Fónagy and Fónagy’s (1976) observation according to which, at the end of utterances especially, the primary stress gives the impression to shift from the last syllable to the initial syllable of the final phrase.

F0 curve of the utterance-final phrase un vif succès (‘a great success’) in the contemporary style (bottom) and the imitation (top).
A stylization of pitch in which each vowel is defined by an initial target, a final target, and possibly an intermediate target is extended to other utterance-final accentual phrases in Figure 8. This figure matches vowel durations and discards consonants (in particular discontinuities due to unvoiced consonants). Thus, it should not be viewed as a faithful F0 contour but as a more readily readable visualization of the style effect.

Stylized melody of utterance-final vowels in the C-style and the I-style. The y-axis reports F0 values in Hz.
4.3.2 Participants and task
Twenty-six subjects (12 males, 14 females, aged 32 on average, different from those of the first experiment) took part in Experiment 3. They were all native speakers of French with normal hearing. Their ability to distinguish old from recent recordings was self-estimated at 3.5 on a 1–5 scale – somewhat higher than in Experiments 1 and 2, even though in terms of background, the subjects had similar profiles: a third to a half were speech scientists.
The interface to listen to stimuli was the same as in the first experiments, and the pair order was also randomized, in order to balance the order of presentation. Yet, unlike previously, listeners had to make a binary forced choice, to indicate for each pair if the utterance order of appearance was older–newer or newer–older. Each experiment lasted 20–30 minutes.
4.3.3 Results
For each type of pair involving the prosody transplantation of original, contemporary or imitation excerpts, results were rated 1 if the first sentence of the pair was judged the older, and 0 otherwise. Results were thus expressed in terms of percentages of “older–newer” answers. As the pairs were presented in both ways, they were all flipped so as to have the type of pair (OC, IC or OI) plus a parameter indicating the order of presentation. For example, an OC pair was coded “OC order 1”, a CO pair was coded “OC order 2”, and the results were expressed as the percentage of original stimuli perceived as older than contemporary stimuli. Results are presented in Figure 9.

Results of Experiment 3: the percentage of “older–newer” answers is given for OC, IC and OI pairs, and the percentage of “newer–older” answers is given for CO, CI and IO pairs (O = original, C = contemporary, I = imitation).
An ANOVA (completely randomized two-factorial design) was conducted on the answers of Experiment 3, coded as explained just above (with the percentages of “older–newer” responses). The two fixed factors were the type of Pair presented to the listeners (3 levels: OC, IC, OI) and the Order of presentation of the pair (2 levels). Again, the significance level was set to 0.01.
The type of Pair had a significant effect on listeners’ answers, F(2, 1164) = 73.02, p < 0.01. The pair Order of presentation was not significant, F(1, 1164) = 0.08, Observed Power = 0.013, neither was the interaction between Pair and Order, F(2, 1164) = 3.03, Observed Power = 0.35: listeners answered “older–newer” as often as “newer–older”.
Subsequent post-hoc comparisons (Tukey’s HSD tests) with an α level of 0.01, performed on the Pair factor, show that OC and IC pairs form a homogeneous subset, but differ significantly from OI pairs. The prosody of imitations and original stimuli is perceived as older than that of contemporary stimuli in over 80% of cases. For comparisons between imitations and original stimuli, responses are more balanced. The original prosody is found slightly older when compared with the imitation prosody, which is in keeping with the tendency for a decrease in word-initial stress reported in 4.3.1: greater in the O-style than in the I-style and greater in the I-style than in the C-style. However, this tendency is to be taken into account with caution, since a chi-square test shows no significant difference between the results obtained by the listeners on OI pairs and a random distribution. In other words, the journalist succeeded quite well in imitating the Gaumont-Pathé prosody: even though he had not listened to the stimuli, he captured the prototypical speaking style.
In their comments (expressed in non-specialist terms), some listeners of Experiment 3 indicated that they relied on pitch range and height, sentence endings, syllable duration and pauses. A higher pitch and more pitch movements, especially at the end of utterances, are associated with old recordings. Inversely, a more monotonous intonation is judged as more recent, which is in accordance with our acoustic analyses (mean pitch measurements, ΔF0 especially in clitic–polysyllabic word contexts and rhythmic patterns before a pause).
A more in-depth analysis of the listeners’ answers was carried out for each pair of stimuli. While mean results for OC comparisons show a clear tendency to judge the originals as older than the contemporary stimuli, for two samples, listeners could not decide which one of the C-style or the O-style was the older. Moreover, with the same two original stimuli in OI comparisons as well as in a third OI pair, the imitations were clearly perceived as older than the originals (in over 70% of answers, whereas mean results for these comparisons are close to chance). These three original stimuli are interesting because they do not convey very marked prosodic cues, contrary to the corresponding imitations.
In order to match these subjective ratings and acoustic parameters, additional prosodic analyses of all the stimuli of Experiment 3 were performed. Since there are too few occurrences to count clitic–nonclitic pitch differences and prepausal contexts for each utterance, prosodic variation was quantified in terms of mean pitch, mean syllable duration and 9th decile of all F0 values (extracted with PRAAT). For example, sentence 02 (in the Appendix), for which the OC comparison is near random and the imitation was perceived as older than the original, mean pitch is 127 Hz in the O-style, 148 Hz in the C-style and 169 Hz in the I-style, while the 9th F0 decile is 156 Hz in the O-style, 193 Hz in the C-style and 250 Hz in the I-style. Syllable duration is almost equal.
Correlations between perception results and overall prosodic measurements were computed by calculating the difference of these prosodic measurements in each stimulus pair (e.g., O-style mean F0 – I-style mean F0, in semitones), for each OI pair in its order of presentation. Correlations with perception are r = 0.89 for mean pitch difference, r = 0.38 for mean syllable duration difference, and r = 0.77 for the 9th F0 decile difference. These results point out that, whereas the correlation with syllable duration is weak, the higher the pitch of an utterance (both on average and in terms of F0 peaks), the more old-fashioned its perception. The perceptual salience of high F0 peaks may be related to the presence of more word-initial stress in old-fashioned sentences. These final findings should be taken with care, as they concern only a few speech samples, but they corroborate the prosodic analyses performed on the whole corpus.
5 Discussion and future work
Speech processing (automatic alignment and pitch extraction) allowed a data-driven approach to stylistic changes, in particular prosodic changes in the French broadcast news style. From this well-defined context, it enabled the identification of epoch-specific patterns and a quantitative study that goes beyond usual impressionistic descriptions. For the French journalistic style, corpus-based results showed a decrease of mean pitch, word-initial stress (whether emphatic or non-emphatic, at least as far as its melodic correlates are concerned) and prepausal penultimate lengthening which was more marked in the 40s and 50s, in particular for nasal vowels. These two decades are the most different from the other ones, as confirmed by the acoustic measurements made on the more limited but more controlled material used in the perceptual tests.
Several issues deserve discussion. In particular, did we really measure a decrease of word-initial stress? Do our measurements reveal a stylistic change or something else? If they correspond to a genuine stylistic change, where does it come from? Finally, what may be the impact of such a data-driven study on the prosody and speech processing communities?
While speech rate has not changed, word-initial stress height and the associated nucleus lengthening – which may distinguish an emphatic stress from a non-emphatic stress according to Astésano (2001) – have decreased over the years. On the other hand, an increase over the years of the onset lengthening associated with what may be considered as initial stress was found. This intriguing result raises challenging questions about French prosody, suggesting that the acoustic correlates of initial stress have changed since the 40s. Other statistical analyses are necessary to understand the apparently paradoxical discrepancy between onset duration-related and pitch-related correlates of stressed syllables. With respect to a 20% just noticeable difference, we measured an overall increase of onset lengthening cases from 19% in the 40s and 50s to 28% in the 80s and 90s. A lengthening of over 20% is not negligible, but it certainly results in fewer linguistically relevant cases than pitch differences of over 3 ST, which are assumed to have a phonological role (’t Hart et al., 1991). By applying the same 3 ST threshold to clitic–nonclitic word contexts on a large corpus of contemporary standard French, Woehrling et al. (2008) found initial stress rates of 29% in text reading and 12% in spontaneous speech. Based on perceptual criteria, Fónagy and Fónagy (1976) reported initial stress rates of 33% in the broadcast news style of the 70s and 21% in conversational speech. The decrease we measured from 28% in the 40s and 50s to 18% in the 80s and 90s in our archive corpus falls within that range.
In our study, speech processing enabled us to select audible samples of the studied phenomena (e.g., words such as nation or présent, where the initial or penultimate vowel is longer than the final nasal vowel). Listening tests based on prosody transplantation were conducted with a controlled subset from the 10-hour archive corpus. The phonetic differences we found validated the observations made on the whole corpus, leading us to speak of a stylistic change occurring over the decades, not merely an epiphenomenon of technical, situational and content-related conditions. Our first perception experiment (Experiment 1) showed that war reports dating back to the 40s and broadcast news of the 50s were perceived somewhat similarly. These two decades were pooled in the analysis of the results of Experiment 2 (as they were in the acoustic analysis of the whole corpus, broken down in four periods), and were perceived as significantly different from the later decades. Finally, Experiment 3, based on the imitation of 10 excerpts taken from war archives, showed that prosody could differentiate a current journalist style from an old-fashioned style with the same content. Even though complex issues remain, these findings strongly suggest a prosodic evolution in news announcers’ style.
Some of the differences in prosody are attributed to decreases of mean pitch and word-initial stress, thus corroborating the Decrease Hypothesis put forward in the introduction. Whether this prosodic change applies to styles other than the newscaster style is not explored. However, since word-initial stress is often associated with the French journalistic style (Fónagy & Fónagy, 1976; Léon, 1993; Oakes, 2002; Gendrot, 2006; Goldman, Auchlin, et al., 2007), a decrease in that matter is interesting. The current study illustrates a methodology for automating analyses of large corpora including the use of perception/imitation experiments, which can be applied to other types of data and other language varieties, and may shed light on the phonetics/phonology of French prosody. Additional perceptual experiments are needed to advance our understanding of the role of prosody in the characterization of old-fashioned speaking styles. We hope they will allow us to tease apart the contribution of onset lengthening and pitch rise to the perception of initial stress in French.
Acoustically, especially prosodically, and perceptually the 40s and 50s departed from the subsequent decades. This tendency was confirmed by the imitation-based experiment which showed good correlations between perception and prosodic parameters. The 60s were marked by no dramatic breakthroughs in sound capture techniques but rather by advances in recording restitution and broadcasting techniques (Montagné, 1995; D’Almeida & Delporte, 2003). The advent of the transistor, the massive use of audio tape, the popularization of television (as later on the liberalization of the FM band in the 80s) have modified listening practices which have become more democratic. News announcers no longer speak to movie theatres or living-rooms but to particular (especially young) listeners. They no longer talk to an audience in the solemn way of learned teachers or like village fair advertisers. Most often, they still read news items (Vihanta, 1991, 1993), in a manner to keep the listeners’ attention by marking the information structure, focusing on important or dramatic events, and underlining topic changes (Carton, 2000). Also, news announcers more often speak in a conversational way as they discuss topics on the air with fellow journalists or other peers (Wenk & Wioland, 1984). This is interesting in the light of Astésano (2001), who points out that, in many respects, the journalistic style stands prosodically between that of reading and of spontaneous speech. This conversational style is facilitated by close-talking microphones. The technique impacts the sociology of the audience, the audience’s image of the news announcer, and in turn, the announcers’ speaking style. This adaptation is also interesting with regard to communication models such as Bell’s (1984) audience design: the way of speaking is flexible and may be adjusted to the listeners’ profile, even when there is no exchange as is the case in public speech. Taking the H&H theory terminology (Lindblom, 1999), this change can be regarded as a listener-oriented shift from hyper- to hypo-articulated speech. Familiarity is now prioritized over intelligibility (Eskénazi, 1993). For automatic speech processing, it is noteworthy that archive documents do not entail that many more recognition errors than recent newscasts. This counter-intuitive result (Barras et al., 2002) may be explained by the fact that in the 40s and 50s French news announcers made more effort to be understood, and this was in part via a more emphatic prosody.
In this study, we did not examine linguistic functions of prosody, for example that used in signaling important words and/or ideas, positive/negative intensification (Kohler & Niebuhr, 2007). Instead, we highlighted another relevant function in speech: the indexical function of prosody which is used to signal an old-fashioned speaking style. More thorough analysis is needed to explore the different functions performed by prosody and specifically time-related changes in prosody. Examination of broadcast news of the early 21st century is currently being planned. Will the trends observed continue or reverse, especially as far as the emblematic French initial stress is concerned? Has this feature lost its value on the linguistic market (Bourdieu, 1982), as if it had propagated too much? Linguistic changes are not linear (Labov, 1994); the answers will also depend on a number of social factors such as news announcers’ need to distinguish themselves, their inclination for sensationalism or seriousness, the prestige or, on the contrary, the stigmatization of the journalistic style.
Footnotes
Appendix
List of sentences used during the perception tests. The sentences marked with a star are only used for the imitation corpus.
| Date | Text |
|---|---|
| 1940* | Parallèlement, la visite à Berlin de Monsieur Molotov, que nous voyons ici accueilli à sa descente de train par Monsieur von Ribbentrop, montre clairement que la Russie suit avec sympathie l’édification de l’Union Européenne. |
| 1941 | Nous voici maintenant au sud du gigantesque front où les troupes roumaines attaquent la ligne de puissants fortins, protégeant la région située à l’est du Dniepr. |
| 1941 | Paris. Les artistes et l’orchestre au complet du grand opéra de Berlin arrivent à Paris, où, on le sait, ils ont donné plusieurs représentations qui remportèrent un vif succès. |
| 1941* | La légion des volontaires français contre le bolchévisme s’est constituée : sa première réunion a lieu au vélodrome d’hiver. |
| 1941* | L’ex-champion d’Europe, l’Italien Locatelli, est opposé au champion de France poids moyen Hassan Diouf. |
| 1942* | Ici, au centre du front, c’est par un violent blizzard que cette équipe de téléphonistes va construire une nouvelle ligne. |
| 1943 | Sur le central du Roland Garros se joue la finale du tournoi de masse. |
| 1944* | C’est pour se pencher sur ceux qui souffrent que le Maréchal est venu en Lorraine. Nous voici à Nancy, Place Stanislas. |
| 1945 | Dans cette maison de campagne, toute fleurie, s’est déroulée la conférence qui fut la plus secrète du monde. |
| 1948 | C’est à midi que Monsieur Robert Schuman est monté à la tribune, sous les applaudissements de tous les délégués présents. |
| 1953 | La presse parisienne s’est tue. Un bulletin des informations officielles la supplée, sans la remplacer. |
| 1955 | Du côté accessoires de nouveaux progrès ont étés accomplis, aussi bien en ce qui concerne l’équipement électrique, que l’équipement pneumatique. |
| 1956 | Les hommes-grenouilles de la Marine française et les scaphandriers de la Royal Navy sont à l’œuvre pour reconnaître les épaves et tenter de dégager le port. |
| 1957 | Malgré une pluie diluvienne qui est venue quelque peu gâter les manifestations extérieures, c’est une imposante cérémonie qui vient de se dérouler au capitole, où a été signé tout à l’heure, avec tout l’appareil des grandes rencontres historiques, l’acte de naissance des deux nouvelles communautés européennes. |
| 1959 | Sur le trajet, la foule était nombreuse. Plus de curiosité que d’enthousiasme, mais le dégel ne faisait que commencer. |
| 1960 | N’hésitons pas à employer les grands mots, la marée montante de la délinquance juvénile, ça n’est pas seulement du cinéma ou de la littérature. C’est un fait qu’enregistrent les froides statistiques officielles de l’éducation surveillée. Quelles en sont les causes? Que fait-on contre elle? Que faudrait-il faire? C’est ce que l’on va tenter de déterminer dans cette enquête, au long de laquelle Maître Hélène Falconetti, avocat à la cour, spécialiste de ces problèmes, va nous conduire. |
| 1962 | L’agriculture de groupe, c’est la formule qu’ont choisie quatre jeunes ménages qui, ici, à Kergounan Ploumoguer, à l’extrême pointe du Finistère, ont créé un GAEC, un groupement agricole d’exploitation en commun. |
| 1963 | Immédiatement, les ministres étaient cernés par les journalistes, voici leurs toutes premières réactions. Le français: les négociations ne sont pas rompues elles sont suspendues. |
| 1964 | À Bruxelles, la coopération entre l’Europe et l’Afrique a pris un bon départ aujourd’hui. C’était en effet la première réunion de travail entre les ministres des six pays du marché commun et les représentants de dix-huit pays africains qui sont maintenant associés au marché commun par la convention de Yaoundé. |
| 1965 | Mais euh ce courrier est intéressant parce qu’il montre une jeunesse qui est finalement assez différente, de celle qu’on imagine, une jeunesse qui a soif de vérité et une jeunesse qui s’intéresse à tout, de la politique à l’amour. |
| 1970 | Un million quinze mille six cents demandeurs d’emploi fin octobre. Mais qui sont-ils? Qui sont exactement les chômeurs? À quelles branches d’activité appartiennent-ils? Dans quelles régions vivent-ils? Regardez bien cette carte: c’est la carte de France du chômage. |
| 1971 | Nous allons faire un tour des capitales européennes en commençant par Bruxelles où Monsieur Malfatti, le Président des communautés Européennes, a parlé de dimensions nouvelles à Noël de Winter. |
| 1972 | Comme des dizaines de milliers de femmes, elle a choisi le travail à temps partiel, le travail à mi-temps. Ecoutez-la. |
| 1973 | Porte d’Asnières porte Dauphine ce sont les trois kilomètres qui manquaient au périphérique afin que l’anneau de béton entoure complètement la capitale. |
| 1975 | Pourquoi la fête à Lomé? En quel honneur ce déploiement de faste? Pour un événement qui dépasse, et de loin, le folklore local: la signature de la convention qui associe le marché commun à l’Afrique et quelques pays des Caraïbes et du Pacifique. |
| 1981 | Voici le moment tant attendu par tous ceux qui souhaitaient que l’Académie française consacre un grand écrivain, même si elle est une femme. |
| 1982 | Revenons maintenant à des affaires plus graves. Savez-vous combien d’appartements en France sont dépourvus des éléments de confort les plus élémentaires? |
| 1983 | Les nouveaux accords communautaires vont redistribuer les cartes équitablement. |
| 1985 | Il s’agit notamment de la création d’un grand marché qui abolira les frontières européennes en 1992, d’un rapprochement des politiques économiques, du renforcement de la cohésion monétaire, de l’extension des pouvoirs du parlement européen. |
| 1987 | Quinze mille agriculteurs selon la police, trente-cinq mille selon les organisateurs ont donc manifesté hier à Angers leur mécontentement. |
| 1991 | Ils s’attendaient à plancher toute la nuit et ils n’ont pas été déçus. |
| 1991 | Pas de diplôme, pas de qualification, un horizon bouché. |
| 1994 | Après le très médiatique envahissement de l’immeuble de la rue du Dragon hier, ce lundi devait être exclusivement consacré à la logistique. |
| 1995 | Le dossier sera donc à nouveau évoqué lors du prochain conseil européen au printemps et entre temps les parlementaires de Strasbourg seront saisis de l’affaire lors de la session du dix au quatorze décembre. |
| 1997 | En quarante-sept jours les manifestants auront déridé plus d’une fois les forces de l’ordre ; jusqu’à présent les leaders de l’opposition tiennent leur pari: éviter les affrontements. |
Acknowledgements
This work was partially financed by OSEO under the Quæro program and ANR within the framework of the PADE project. We are indebted to Laurent Vinet and the French Institut National de l’Audiovisuel (INA) for the audiovisual corpus and its transcription which were made available for the
