Abstract
This article reports a study examining whether foreign language (FL) word learning can be improved with reduction in cognitive load. Cognitive load theory has received substantial supports in various fields of learning but never in FL word learning. Due to the defined poverty in exposure to the FL, hence deprived cognitive pre-requisites for natural FL development, cognitive load could be critical to FL learning success. Thus while word learning may be a simple attempt of associating word forms with their meanings for L1 children, for FL learners, the cognitive load is multiplied by the additional task of taming the often intractable phonological forms (both perceptive and productive) at the same time they are making the association. In light of cognitive burden reduction, FL learners could thus benefit from learning phonological forms first as their L1 counterparts do. The present study examined whether beginning learners of English as a foreign language (EFL) learn English novel names better if first familiarized with the phonological rimes of target names whose referents are taught only later. Chinese-speaking first graders were assigned to one of three teaching conditions: rime familiarization, which familiarized children with rimes through rhyming activities without any meanings involved; spoken vocabulary, which taught words in rhyming groups together with their referents; and semantic control, which focused on word use. As the results showed, the rime familiarization group outperformed the other two by an improvement score several times greater, suggesting the critical role of cognitive load in FL word learning success.
I Introduction
Word learning is fundamentally a process of associating concepts with their corresponding word forms (i.e. sound strings and their articulatory routines) or vice versa (e.g. Zamuner et al., 2014). It can be easy if both meanings and forms are readily available for making the association. Even for infants, for example, given the hundreds of receptive phonological forms (which has been shown to be critical to development of productive phonology, i.e. articulatory routines; for a review, see Stoel-Gammon, 2011) that they are already familiar with when they start to utter their very first words, ‘vocabulary building is a matter of linking existing phonological forms with their denotations’ (Swingley, 2007: 454; italics added). Word learning becomes even easier when they produce their first 50 words when, according to Swingley’s estimate, they could have been exposed to – in less than a year – thousands of word types and millions of word tokens. The abundant exposure would have equipped them with sufficient practice opportunities for productive routines and thus reduce word learning to a simple task of form–meaning mapping, a cognitive load affordable to most typically developing L1 children.
In foreign language (FL) teaching, words have often been taught in a similar fashion, i.e. via paired-associate learning (e.g. Grade 1–9 Curriculum Guidelines: English Subject; Ministry of Education, 2011). While this may pose no special challenge to most native-speaking children given their well-established phonological repertoire, to FL beginners, it could. School age children starting to learn an FL are often poorly equipped phonologically in the FL, and thus the seemingly same act of word learning actually imposes double cognitive burden on FL learners, with the additional task of coping with strange FL word forms. This difference in prior phonological knowledge carries critical cognitive implications, as the dual task could easily exceed the limited capacity of the young FL learner’s working memory, which for average adults has been argued to be seven plus or minus 2 items by some (e.g. Miller, 1956) and much fewer by others (e.g. four chunks according to Cowan, 2001).
According to the cognitive load theory of Sweller and colleagues (e.g. Paas et al., 2003; Sweller, 2010), indeed, cognitive load is crucial to learning success simply because one’s working memory is limited in both the amount and duration of the information one can hold and process at a time. And one way to go around the inborn limitation is through the help of long-term memory, in the form of creating schemas that integrate lower-level info into higher-level units (Paas et al., 2003), hence a reduced number of info items within the working memory’s capacity. One critical type of cognitive load according to Sweller and colleagues is the intrinsic cognitive load, which is measured by the number of elements to be simultaneously processed and the extent of interactions among them, a property also known as element interactivity. Thus word learning could be low in cognitive load if it involves only one interaction between two elements, as when a child is learning a new word by making association (interaction) between its two elements, i.e. sounds and meanings. If, however, either element has not been earlier acquired (i.e. overlearned to become one single automatic schema in the long-term memory), the numbers of both the elements and their interactions to be concurrently processed will increase, hence a much heavier cognitive load. One major source of word learning difficulties for FL learners thus lies in the extra burden of learning at the same time also an intractable number of new phonological elements including, but not limited to, the FL’s phonetic features, phoneme inventory, phonotactic patterns, stress patterns and, most critically, their interactions. All these would potentially exhaust the FL leaner’s working memory, leaving little or no room for the critical association task of word learning to take place.
Empirically, even for children acquiring their mother tongue, cognitive load has been shown to be critical to their word learning success. Stager and Werker (1997) tested 14-month-olds for their performance in discriminating similar-sounding phonemes (e.g. /p/ vs. /b/) in minimal pairs (Experiment 4), a sensitivity children as young as 4–6 months have been shown to have developed, and in learning two word (form)-object pairings that differed only in one distinctive feature (e.g. ‘dih’ vs. ‘bih’; Experiment 1). Note that the two tasks differed mainly in the amount of cognitive load they carried: single tasking for the discrimination task but multi-tasking – association as well as sound identification – for the other. The results showed that while the novice word learners were, as expected, able to distinguish phonetic details, they failed to apply this sensitivity to word learning. The differences appeared to result from the greater computational demands of the latter task. More specifically, the difficulty a ‘novice word learner faces is one of being able to detect and encode the detail that is perceptually available at the same time that he or she is attempting to link a newly heard word with an unnamed object’ (Werker et al., 2002: 25; italics added).
Challenges of the same nature can only be more severe to FL learners, whose poverty in exposure opportunities to, hence phonological representations of, the FL has indeed led to their observed reliance on other inappropriate yet readily accessible cognitive resources such as L1 phonology or orthography in learning FL words (e.g. Chen, 2011b). Given the rich literature in the critical role of word forms in word learning (e.g. Swingley, 2005, 2007) as well as related literature in cognitive load (e.g. Paas et al., 2003; Sweller, 2010) and L1 lexical development (see for a comprehensive review Stoel-Gammon, 2011), and the common cognitive mechanisms L1 and FL learners must share, there has been surprisingly scant research taking advantages of findings from the L1 cognitive development literature to FL learning. The reported study was an attempt to bridge the gap by examining whether the critical role of prior knowledge of phonological forms observed in L1 acquisition would be equally critical to word learning in English as a foreign language (EFL). Specifically, it asks whether learning sequence matters, that is, whether word learning is facilitated if the learner has been first familiarized with the target words’ productive forms.
It was hypothesized that one main obstacle in FL word learning is the extra cognitive load brought about by the unfamiliar FL phonology or the low availability of phonological forms due to limited exposure. If instead FL word learning is held back until the crucial phonological forms have been acquired, the cognitive load involved should be reduced and/or the phonological forms should have now become better accessible. Either way, word learning should be expedited. It is further assumed that training in productive forms should be more economical, as motoric practice has been shown to facilitate not only lexical development (e.g. McCune and Vihman, 2001; Stoel-Gammon, 1998; Stoel-Gammon and Cooper, 1984) and word learning (e.g. Leonard et al., 1981; Schwartz and Leonard, 1982) but lexical perception (DePaolis et al., 2011, 2013; Majorano et al., 2014; Marchman et al., 2010; Velleman and Vihman, 2002; Vihman, 1993; Vihman and Nakai, 2003) and phonological memory (e.g. Keren-Portnoy et al., 2010; Munson et al., 2005) as well. It has even been suggested that children tend to learn only words whose spoken forms they are already able to produce (e.g. Leonard et al., 1981; Schwartz and Leonard, 1982), an observation which Ferguson and Farwell (1975) terms lexical selection.
Various theories have been proposed to explain such phonological selectivity behavior in children. Some researchers view it as a tendency to avoid difficult-sounding words (e.g. Schwartz and Leonard, 1982). For others, it is a result of a feedback loop linking acoustic signals to the speaker’s own articulatory movement (Stoel-Gammon, 1998), gradually leading to an automatic, top down process matching similar patterns in the ambient language to those in the child’s newly constructed phonological repertoire, a process Vihman (1993) called articulatory filter. As a result, only sounds or sound patterns similar to those the child has acquired become perceptually salient for selection from the adult input. Whether an avoidance behavior, perceptual filter, or both, vocal practice of phonological units or patterns appears desirable for successful word learning. Such solid benefits are unlikely confined to natural language acquisition, as the same underlying cognitive mechanisms appear to apply to both L1 and FL learning. That is, sufficient vocal practice would result in automatized planning and execution of practiced speech sounds (e.g. the practiced syllables in Vihman, 1992), which is then ready (as a long-term memory schema) for mapping to corresponding meanings in word leaning as well as for perceiving similar sounds in the ambient language when available. Given the a priori poverty in the ambient FL exposure on the one hand and productive phonology’s demonstrated benefits to both perception and production on the other, FL learners should thus find direct productive training – i.e. skipping perceptual training – especially instrumental in their FL word learning.
A practical yet crucial issue involved in direct training is what units should be taught for the hypothesized benefits. Among the candidate pool of words, syllables, phonemes, and subsyllabic units, the first two cannot be good choices because there are too many syllables as well as words to learn. Neither can phonemes, as awareness of phonemes has been shown to occur only with the learning of an alphabet (for a comprehensive review, see Castles and Coltheart, 2004). Since introducing orthography and its mapping to sounds requires additional cognitive loads, teaching phonemes to preliterate beginning learners cannot be ideal. These leave subsyllabic units as more reasonable units of teaching, given its reduced number of units to learn, hence lightened cognitive load. Indeed, subsyllabic units such as rimes have been shown to have developed prior to phonemes and to enjoy an especially salient role in English as well as many other alphabetic languages (for a review, see Ziegler and Goswami, 2005). However, it has been argued that an even more natural linguistic unit for children is core syllable (i.e. syllables without codas and with optional singleton onset; see Chen, 2011a). As core syllable is a special type of open syllable, it is small in quantity, hence also cognitively affordable.
While core syllable may be a universally available and the earliest developed syllable type (e.g. Levelt et al., 2000), however, it is important mainly for its role in early linguistic development for alphabetic languages. Later in their language development, children (as well as adults) speaking alphabetic languages such as English and Dutch have been found to base their phonological processing on rimes (e.g. Treiman, 1986). Such preference in older children as well as adults has been attributed to the statistical distribution properties of (adult, hence target) English as well as many other languages. Frequency studies of phoneme collocation, for example, suggest that the rime is a more coherent subsyllabic unit as it has, compared to other rival combinations (such as, given a CVC syllable, lead, i.e. CV, or consonantal frame, i.e. C_C), the most phonological neighbors (e.g. De Cara and Goswami, 2002; Ziegler and Goswami, 2005). The difference in preference (for core syllable for younger children and for rimes for older children) has been suggested to result from syllable restructuring due to developmental pressure (e.g. Chen, 2011a). With increased vocabulary size, that is, restructuring could take place for learners to attend to finer distinction among sounds (Metsala and Walley, 1998), which is especially important to discriminate words in the dense neighborhood, i.e. words with many phonological neighbors (Storkel, 2002).
Syllable restructuring however occurs only to speakers of languages with complex syllable structure such as English for the said acquisition economy. For languages with simple syllable structure such as Chinese, whose permissible codas are limited to singleton nasals only, hence no need for restructuring, their speakers tend to retain core syllable as an intact unit (Chen, 2011a). Such tendency works well with their mother tongues but becomes an obstacle when they are learning a rime-based FL such as English. Training in productive rimes could thus help Chinese EFL learners for the same efficiency benefits their native speaking peers have enjoyed with syllable restructuring from, given a CVC syllable, more intact CV to more cohesive VC. As an example of the cognitive economy, there are only 400 rimes for the most frequent 3,000 monosyllabic English words (Ziegler and Goswami, 2005: 19), an advantage especially valuable given the limited time allotted to FL learning time. In Taiwan, for instance, the officially recommended EFL vocabulary size is 300. If the same rime/word proportion for L1 English (400/3000) applies, then it takes only 40 rimes to learn the 300 words. Pedagogically, moreover, it would be much more boring to repeat the same whole words or, by the same logic, whole syllables than to practice rhyming words and nonwords that vary in onsets. In light of both learning economy and practical teaching considerations, rime thus appears a better unit for vocal practice than whole syllables or words.
The present study examined whether vocal practice of English rimes could benefit EFL word learning as a result of increased phonological familiarity. Given the documented benefits of existing productive phonology on word learning and the salience and developmental appropriateness of rimes, it was hypothesized that EFL word learning could be facilitated through first familiarizing its learners with the target rimes of the words to be learned. More specifically, productive familiarity with rimes as a result of intensive vocal practice prior to actually learning the target words should enable EFL learners to focus their cognitive resources on linking word forms to their signification, without attention resources divided for processing word forms not yet automatically available. If the theory is tenable, one would predict that EFL learners with prior training in only spoken rimes should perform better in EFL word learning than learners trained in traditional ways, i.e. with both word forms and meanings (or use) introduced simultaneously.
II Method
1 Participants
Eighty-seven first graders (mean = 81.12 months; SD = 3.73 months) from three classes of an urban elementary school in southern Taiwan were recruited. English was formally taught beginning in grade two, and the participants’ English was thus basic. Two participants failed to show up at the post-test, leaving the total number of participants 85 for the initial statistical analyses reported below.
2 Materials
a Outcome measure
The participants were individually tested twice, in the pretest and the posttest, on their ability to learn new heard words by a nonnative novel name learning task adapted from Hu and Schuele (2005). In their study, disyllabic novel names were used for their older, third-grader participants. Given the age differences, simpler, monosyllabic novel names, Ked, Dook, and Pite, were used for the first graders in the present study. In the first (training) trial, the children were, for each of the three names, first shown a cartoon figurine on a laptop monitor and played its name; they were then required to repeat the name without feedback given. Starting from the second trial, they were asked to name the picture when shown, without having been first played the name. The name was played to them if they did not respond or responded incorrectly. For each trial, the pictures were shown in orders different from the previous one, while the first picture was never the same as the last of the immediately preceding trial; the same orders were given to all the children. The task stopped when the child correctly responded to all three names for two consecutive trials or at the end of the tenth trial. The responses were scored as well as phonetically transcribed and audio-recorded by trained undergraduate testers. They were instructed to score in a strict manner so that any produced sounds that were confusable with similar phonemes were judged to be a failed attempt. A trained graduate assistant later verified the transcription and scoring based on the recorded materials. The researcher was consulted when differences occurred. Given the participants’ beginning level English proficiency, the traditional scoring based on the number of trials to criterion or termination was not used. Instead, as in Hu (2003), a total score (maximum = 27) was calculated by counting the number of individual correct responses.
b Control measures
Two groups of control measures as well as age (in months) were taken of the participants, cognitive abilities and English proficiency levels. Three dimensions of their cognitive abilities were measured: nonverbal intelligence, phonological memory, and phonemic awareness; and two aspects of their English abilities were tested: letter name knowledge and auditory vocabulary.
Nonverbal intelligence: The Pattern Completion subtest of the Matrix Analogies Test-Expanded Form (MAT; Naglieri, 1985) was taken for the participants’ nonverbal intelligence. The participants were shown a picture with a missing piece along with five small pictures the size of the missing piece. They were required to select the one that corresponded to the missing part in terms of overall patterns of the big picture. The maximal score was 15, and the Cronbach’s alpha was .70.
Phonological memory: There has been substantial research pointing to the significant role of phonological short-term memory in spoken word learning in both L1 (e.g. Gathercole and Baddeley, 1990; Gathercole et al., 1997) and L2 (e.g. Hu, 2003; Service and Kohonen, 1995). However, as it has been shown that phonological memory is subject to wordlikeness, hence the influence of long-term memory (e.g. Snowling et al., 1991), Chinese instead of English pseudowords were used (for a discussion, see Hu, 2003) to give a purer gauge of phonological memory with its overlapping with lexical knowledge removed, which was tapped here by a separate task (PPVT; see below). In the present study, phonological memory was measured by the pseudoword repetition from Hu and Schuele (2005). The participants were asked to repeat three two-syllable Chinese pseudowords, in the exact order they had been heard, for each of the six trials. Scores were calculated as the number of the correctly recalled syllables in their heard positions (maximum = 36, given the 6 syllables in each of the 6 trials). The Cronbach’s alpha was .82 for this task.
Phonemic awareness (PA): This was assessed by both phoneme deletion and phoneme isolation tasks. In the former, the participants were played a monosyllabic nonword (e.g. kes) and asked to repeat it and then to respond to the request of removing a sound (e.g. /k/) from it. In the latter, the participants were again played a monosyllabic nonword (e.g. kes) and asked to repeat it but were then required to produce either the initial or the final sound. Eight nonwords for the former and ten for the latter, all with the simple CVC (C for consonant and V, vowel) structure, were created. Items in each task were balanced in the position of the target sounds, so that there were equal numbers of (C)VC items, which required the removal or isolation of the initial sound, and CV(C) items, which required the removal or isolation of the final sound. Unlike earlier studies, two sub-measures were computed for the construct: onset PA (PA measured by (C)VC items) and coda PA (PA by CV(C) items). Earlier literature has suggested that phonological awareness contributes not only to reading but also to word learning (e.g. Hu, 2003). On the other hand, it has also been shown that, at least for Chinese readers, not all types of PA play the same significant role in reading success. Chen (2011a), for example, showed that only performance on onset PA items, but not that on coda PA items, contributed to its second-grade participants’ beginning literacy. It would be therefore interesting to see if the different correlations of PA to reading acquisition occur to word learning as well. The reliability for onset PA was .83 and that for coda PA was .78.
Letter knowledge: Orthography has been consistently shown to influence both spoken word recognition (e.g. Chereau et al., 2007; Taft et al., 2008) and auditory word learning (Hu, 2008; Nelson et al., 2005). The present study followed Muter et al. (2004) in its measurement of letter knowledge. The letters were given in lowercase in a random order on a laptop. Responses with the letters’ names or sounds were accepted as correct (maximum = 26). All participants received items in the same random order. The Cronbach’s alpha was .95 for this task.
English auditory vocabulary: Existent lexical knowledge has been shown to contribute to word learning abilities (e.g. Gathercole et al., 1997). The Peabody Picture Vocabulary Test, 3rd edition (PPVT-III henceforth; Dunn and Dunn, 1997), was used to test the children’s English auditory vocabulary. The children were shown four pictures and asked to decide which best represented the word they had just heard. The PPVT-III measures receptive vocabulary in Standard American English based on a national norm from 2 and a half to 90 years of age and consists of 17 sets, each with 12 items, ordered in levels of difficulty. Given the limited proficiency of the participants and based on prior research (e.g. Chen, 2011b), only the first two sets (24 items) were used. The Cronbach’s alpha obtained for this sample was .65.
3 Training programs
To test the cognitive load hypothesis, one experimental group and two control groups were selected with their pretest and control measures taken in the beginning of the study. The three groups differed in their instructional emphases. The experimental group stressed the sole training focus on phonological rimes without any semantic components involved; that is, only sound patterns were taught. The two control groups, on the other hand, shared the similar core focus on actual word learning, i.e. simultaneous learning of sounds, meanings, and their associations, but differed in which side of the association, sounds or meanings, was reinforced. The spoken vocabulary control further stressed the phonological patterns among the learned words whereas the semantic control reinforced the form–meaning association through semantic richness. Thus the major difference between the experimental and the control groups lay mainly in whether cognitive load is minimal, as for the experimental group, or multiple, as for the two control groups, which further differed from one another in which side of the coin, sounds or meanings, were further emphasized.
a Rime familiarization training
Children in the rime familiarization group were aurally/orally familiarized with rhyming words or nonwords without any meanings or visual aids attached to them. The three target rimes were recycled every three weeks, with each week focusing on just one target rime in addition to a non-target rime. Given only three target rimes for a total of 12 weeks, the non-target rimes were included to avoid monotony and were introduced in the same manner and number as their target counterparts. That is, for each week, the participants were familiarized with two rimes, one target and the other non-target. The target nonwords appearing in the pretest and the posttest were not used during any part of the training to avoid additional exposures than those experienced in the pretest, which were supposedly similar to all three groups. The non-target rimes, like the targets, were from high-frequency words and they were ‘-eed’ as in ‘need,’ ‘-ake’ as in ‘cake,’ and ‘-air’ as in ‘chair.’
The twelve week training was divided into four 3-week phases: introduction phase, receptive phase, production phase, and general practice phase, each with different instructional emphases while repeating training in the same six targets and non-target rimes. For each target or non-target rime in the first phase, the participants were asked to repeat after the teacher words or nonwords ending in the target (or non-target) rime, excluding the tested novel names. Each rhyming non/word was first repeated alone, and then sequences of rhyming words were repeated, with the length of the sequences gradually increasing, starting from a two-item rhyming sequence. In the receptive phase, auditory familiarization was practiced with games like Simon Says, which required the children to respond in a certain, usually playful, manner like thumping on the floor or crying out ‘wow,’ when they heard a word or nonword containing the taught rime. In the production phase, the instructor lead the participants in playing productive games like Relay Race. The children were divided into several groups and a child in each group was told a non/word with the target rime and s/he had to add a new non/word rhyming with the given one and said the two-item sequence to the second member, who in turn had to add yet another new item to the existent sequence. The final phase was a blending of the two earlier practices and concluded the training and, together with the earlier phases, were meant to provide maximal exposure opportunities.
b Spoken vocabulary training
Children in the spoken vocabulary training group were taught to learn words ending in the same target- or non-target rimes as those for the rime familiarization group. They were taught to associate the words’ concepts with their spoken forms in English. The words were then practiced in, sometimes nonsensical, rhyming pairs, such as ‘red bed,’ ‘white kite,’ or ‘look cook,’ without explicit instruction to their overlapping properties. Three sets of words with target rimes and three sets of words with non-target rimes were, like the other experiment group, recycled every three weeks. That is, for each week, the participants received instruction in six new words, including three ending in one target rime and the other three ending in one non-target rime. The three sets of target-rhyming words were ‘bed, red, head, book, look, cook, kite, white, and right.’ Note that, similar to the other two groups, no written forms were used. Distribution of the four 3-week phases was similar to that of the rime familiarization group. In the introduction phase, the children were introduced to the names together with the associated concepts, with one new rime embedded in three new words given each week for both targets and non-targets. In the following two weeks, words learned in the previous week(s) were reviewed while new words were introduced. In phase 2, the receptive phase, the participants were familiarized with the learned words’ auditory forms by responding to heard English words with their Chinese translations or pointing to pictures representing them. In the third phase, production was stressed by requiring the children to provide the English words when prompted with their Chinese counterparts or pictures of the referents. The final phase provided a chance to practice both receptive and productive forms of the learned words with or without associated referents and, like the other phases, games were included as part of the curriculum to avoid boredom.
c Semantic training
Children in this group were taught the same words in the same schedule as those for the spoken vocabulary group, but differed from the latter in its instructional emphasis on semantic richness. The meaning of each word was first introduced but the word was practiced in its use in simple but various contexts. The children were encouraged to combine learned words in simple but meaningful phrasal or sentential templates. Some examples were ‘a white (red) kite (bed)’ for the former and ‘Look at the ____’ (where the blank could be any words they had learned) for the latter. Similar to the spoken vocabulary group, the first phase saw the introduction to the nine target and nine non-target words. Phrasal and sentential combinations of learned words were practiced respectively during the second and third phases, and were blended during the final phase. Again, games were part of the learning activities in all phases and no written words were involved during the training. Unlike the spoken vocabulary group, however, instructional emphasis in word combination practices were placed on semantic possibilities but not shared rimes, which could incidentally occur though.
4 Procedure
The participants in their original intact groups were randomly assigned to either the semantic control (receiving semantic training), the spoken vocabulary control (receiving spoken vocabulary training), or the experimental condition (receiving rime familiarization training). The participating children were tested twice, with pretest given in late December and early January of the following year (i.e. the end of the fall semester) and posttest in May (i.e. the end of the spring semester). In-between the pretest and the beginning of training, there was a winter vacation for a month. The inclusion of the winter vacation was to ensure there would be sufficient time (twelve weeks) for training, as it took about a month to collect data for either the pretest or the posttest, the latter of which began in the 13th week of the spring semester. Since the Chinese lunar new year was celebrated during the winter break for around half a month, when there was no teaching in and out of the school, any confounding factors thus introduced, if at all, were likely to be insignificant.
The instruction for each group was given in two 15-minute sessions a week (in the first period in the morning) for a total of 12 weeks. The participating instructor was a certified full-time English teacher at the participating school. Based on teaching content, schedule, and covered materials provided by the researcher, the instructor came up with feasible lesson plans for each of the twelve weeks for each group, including the objectives, materials, and procedures. The lesson plans were then discussed with the researcher, revised, and carried out from the very beginning of the spring semester, which begins in late February, for a period of twelve weeks.
III Results
An initial analysis of score distribution with the two dependent variables, pretest novel name learning and posttest novel name learning, identified five outliers based on a cut-off criterion of 2.5 standard deviations above or below the means for either of the two name learning variables, leaving a final sample size of eighty participants. Of the five outliers removed from later analyses, one was from the semantic control group (n = 28 = 29 – 1), three from the rhyming training group (n = 26 = 29 – 3), and one from the spoken vocabulary group (n = 26 = 27 – 1). Table 1 gives the descriptive statistics of the pretest measures by group and the accompanying F-scores and p-values of the one-way ANOVAs comparing the three groups (Ns = 28, 26, 26 for respectively the semantic control, the rime familiarization group, and the spoken vocabulary group). No statistically significant differences were found among the three groups for the pretest or control measures except for letter name knowledge, F(2,77) = 3.57, p < .05. Post hoc comparisons showed that significant difference occurred only between the two control groups (p < .05, with the vocabulary training group, Mean = 20.38, SE = 1.16, enjoying greater letter name knowledge than the semantic control, Mean = 15.36, SE = 1.63), whereas the rime familiarization group did not differ significantly from either of the other two (ps > .05). This suggested letter knowledge as a likely confound for training effect on novel name learning performance. To detect potential confounds, correlation coefficients among the potential confounds and the name learning performance were computed and given in Table 2. Among the confound candidates themselves, as the table shows, spoken vocabulary was substantially correlated with both measures of phoneme awareness (onset deletion and coda deletion) as well as letter name knowledge; nonverbal intelligence, on the other hand, correlated only with age but not with other verbal measures, both consistent with their association patterns in earlier literature. Posttest novel word naming performance measures, in contrast, were correlated with none of the pretest measures. This justified the use of ANOVA without inclusion of the potential control measures as covariates, as follows.
Means, standard errors (in parentheses), and group differences F-scores a for pre-test measures (Ns = 28, 26, and 26 respectively).
Notes. aOnly letter name knowledge differed among the three groups (p < .05). The difference occurred only between the two control groups. bOnly phoneme awareness measured by (C)VC items is reported, as that measured by CV(C) items was ceiling for both groups.
Correlations among the outcome and predictor variables.
Notes. aPhoneme awareness (PA) measured by (C)VC items. bPA measured by CV(C) items. *p < .05, two-tailed; **p < .01, two-tailed.
To examine the relative training effects of the three groups, a mixed 2 (time of test) x 3 (condition) ANOVA was performed with time of test (pretest vs. posttest) as a within-participant factor and training condition as a between-participant factor. The means and standard errors of the three groups’ novel name learning performance taken at the pre- and post-tests are given in Table 3. The results showed a main effect of time, F(1, 77) = 11.47, p = .001, and a main effect of training condition, F(2, 77) = 5.12, p < .01. The effect of time of test occurred because of the better performance in the posttest than in the pretest. The training condition effect occurred because participants in the rime familiarization condition outperformed their peers in the semantic (but not spoken vocabulary) control (p < .05), whereas the spoken vocabulary group did not differ from either of the other two (ps >.05).
Pretest and posttest group means and standard errors (in parentheses).
Most critically, the main effects were qualified by a significant interaction between time and training condition, F(2, 77) = 3.25, p < .05. The interaction was a result of significant improvement between the pre- and post-test only for the rime familiarization group, t(27) = 3.56, p < .002, η2 = .34 – a large effect (> .14) according to Cohen’s (1988) magnitude scales of effect size – but not for the other two groups (both ps > .10, η2 < .1). This is reflected in the comparison of performance improvement across the three groups, with the mean improvement between the two tests for the rime familiarization group (mean difference = 3.88) substantially greater than those of the other two. The improvement scores for the rime familiarization group were actually nine times greater than that for the spoken vocabulary control (mean difference = .42).
IV Discussion
Productive phonology has been consistently shown to be critical to lexical development in first language acquisition, as it provides a phonological foundation upon which word learning could be made easier with reduced cognitive demands. The prerequisite phonological knowledge necessary for efficient word learning however has seldom been practiced in an FL context. The present study investigated whether FL words are more easily learned when learners have been first vocally trained in their rimes (rime familiarity group) than when their word forms are introduced simultaneously with their referents (spoken vocabulary group) and/or their use (semantic control). The results suggested a superiority of form-first approach to the other two traditional approaches. The much larger improvement score of the rime familiarization group, which is several times greater than those of the other two speaks favorably of the form-first approach as more efficient as well as successful in EFL word learning, at least for learners whose L1 phonology is much simpler than that of the target FL, as in the case of Chinese learners of FL English. The superiority of rime familiarity training to traditional spoken vocabulary teaching, in the latter of which the same rimes were taught, indicates that, given the input poverty, prior, separate sessions of vocal practice appear desirable in order to sufficiently acquaint EFL beginners with phonological patterns that could then come in handy when learning new words containing them.
While the results attest to the importance of learning sequence, the study remains exploratory as the mechanism behind the obtained effect is still subject to further research. One possible explanation of the obtained improvement is the cognitive load theory (CLT) of Sweller and colleagues. Based on CLT, learning is most effective when the learning task is cognitively affordable. Of the two main elements (phonological form and referent) and one interaction (association between them) in FL word learning, phonological form was apparently the main obstacle to the tested children, whose other cognitive abilities showed no significant differences across the three groups (see Table 1). As discussed earlier, many Chinese EFL learners have been shown to rely on L1 orthography or phonology when learning English words. If the new phonology indeed lies at the core of EFL word learning difficulty, phonological learning prior to word learning should enhance the learners’ phonological ability, which would in turn reduce their cognitive load when learning new English words. Thus the improvement is a result of a chain causality: while the training was assumed to result in better productive phonology in the learner, it was the resulting reduction in cognitive load (due to now better readiness of the word forms) that facilitated FL word learning. In light of word learning success, cognitive load is thus the proximal cause, whereas rime training is the distal cause. It should be noted that cognitive load can also play a part in phonological training itself. That is, focused learning of only phonological patterns could lead to better phonological representations critical to word learning success than the simultaneous learning of form and meaning in the two controls. Thus, cognitive load could work at both training and testing (which is also a form of training) phases. While a CLT interpretation is entirely consistent with the obtained results, other possibilities, such as phonological similarity, exist.
That is, the improvement could be alternatively interpreted as a direct result of greater accessibility to the target forms due to repeated rhyming practice. The greater improvement of the experimental group was accordingly attributable to the greater exposure opportunities it enjoyed. Such frequency account, however, is unlikely. The two controls differed in the type of extra reinforcement each received, with the spoken vocabulary control by design being exposed to the targeted phonological patterns (i.e. rimes) much more often than the semantic control. Yet, the two controls did not differ in the extent of their improvement. Exposure frequency is thus unlikely as a straightforward explanation of the results, as it seems some threshold has to be crossed before phonological familiarity’s influence could be felt. While this alternative seems refutable, the cognitive load account remains nevertheless tentative, as the study had not been designed to test the CLT itself. That said, the results do accord with predictions based on CLT, making it a viable, despite tentative, explanation of the data.
The similar performance between the spoken vocabulary group and the semantic group, both of which failed to improve their word learning performance, was unexpected, as the former also enjoyed practice of the target rimes and was thus expected to benefit as well, though presumably to a lesser degree, and to perform at least better than the semantic group. One possible explanation is the division of the limited attentional resources to both phonological factors and the form–meaning association attempt, leaving little room for further processing of phonological information, even though the target rimes had also been taught. This possibility was suggested in Stager and Werker (1997), where the toddler word learners were able to distinguish minimal pairs perceptually but failed to apply the same ability to word learning, a more cognitively demanding task. Such failure however seemed to vanish with a threshold spoken vocabulary of 25, as suggested in a latter work (Werker et al., 2002). The traditional spoken vocabulary training in the present study could have failed for similar reasons. That is, the dual demands of learning the word form and the mapping at the same time may divide the learners’ attentional resources so to retard word learning. Cognitive resources allocated to phonological rime practice could accordingly be diluted. Taken together, the results suggest that simultaneous introduction to both word forms and their signification could be disadvantageous on both quantitative (i.e. insufficient in practice) and qualitative (i.e. divided, hence limited, in allotted cognitive resources) grounds. Given the substantial improvement in word learning of the rime familiarization group on the one hand and the failure of the other traditional programs to improve, prior vocal practice of phonological rimes thus appears a more desirable teaching approach in terms of both sufficiency and efficiency.
One may wonder, given the results, what their practical meanings are when contextualized in a real FL classroom. After all, an improvement of 3.79, though three times greater than that of the other two groups, seemed relatively small in size given the 27 opportunities to recall the words after the first exposure. The ‘seemingly’ small effect is, to begin with, actually not small, given the obtained eta square (effect size) of .28 which, according to Cohen (1988), is a large effect. It is nevertheless interesting to note that, given the same amount of training time, the other two groups did not appear to improve. Most notably, the vocabulary training group had also been exposed to the same rimes and yet showed no improvement. As simultaneous processing of multiple elements was likely too much for beginning learners to handle, more time may be needed to make up for the high demands for cognitive resources that are unfortunately limited. Practically, it means that much more than 15 minutes a week and/or more than 12 weeks of training would be needed to become acquainted with the taught rimes. The same amount of exposure, in contrast, appeared relatively sufficient for the rime training group to master and apply in the learning of new words, as training in phonological rimes, though difficult in and of itself for beginning FL learners, is nevertheless cognitively affordable given its single tasking, hence its relatively low demands on cognitive resource.
A related issue concerns the practical benefits of the various training programs. It is true that while the rime training group did outperform the vocabulary control in learning new words, the latter might have learned the trained ‘real’ words whereas the former have learned nothing but the sound patterns. Even if the vocabulary training group had indeed learned the trained words, however, the benefit is limited to the taught words only, as the training apparently did not help in learning new words. The rime training group, on the other hand, had acquired phonological representations transferable to learning of new words, as they had never been exposed to the target words during the training, except for the pretest. That is, while we could expect children thus trained to find any new words containing trained rimes easily acquired, the same is not true of children required to learn both meaning and sounds at the same time.
V Conclusions
The results of the present study carry both theoretical and pedagogical implications. Theoretically, they pointed out similar relationships between productive phonology and word learning in both L1 and FL word learning. That is, pre-existing productive phonology appears desirable also to FL learners to ensure learning efficiency. While this pre-requisite knowledge develops naturally in L1 acquisition, in FL separate sessions focusing solely on sound familiarization seem advisable before words are formally introduced, at least for EFL learners whose mother tongue lacks the complex rime structures that play an important role in English vocabulary building. The results of the present study suggest that sufficient familiarity could be achieved through oral practice, which would automatize in learners both perception and motoric routines of the target words so as to free working memory from the extra burden of phonological coding of unfamiliar FL sound patterning. Pedagogically, FL curriculum for beginners would better start with vocal training of English sound patterns without meanings attached. In the case of EFL learning, given the limited exposure to the target language, hence the unlikely occurrence of lexical restructuring, rimes appear to be a more desirable practice unit for its demonstrated preference among native English speakers.
As an initial step in translating findings from L1 speech acquisition to FL teaching, the current study necessarily leaves some issues to be addressed in future research. It would be interesting to examine, for one, whether the hypothesized perceptual salience of the practiced sounds resulting from motoric practice, as suggested by the articulatory filter theory (Vihman, 1993), could be obtained in FL lexical development. A positive result would certainly add to the benefits of productive phonology training. Given the observed reluctance to speak out in many EFL learners, it would be also of great significance to investigate whether the said training could prove beneficial also as part of EFL remedial courses whose students are far behind in English proficiency. As many of them seem unwilling to talk in English and, when they do, they have apparent problems in pronunciation, it would be important to explore if they are similar to late talkers, who are simply late in development, and if vocal training could help them by laying their phonological foundation.
Finally, it would be of great interest to examine to what extent the obtained effect could apply to various types of FL word learning, as some FLs are more different phonologically from a particular L1, whereas others are more similar, and it is conceivable that phonological familiarity issue could be lighter in some but more serious in others. While, for example, Chinese-speaking learners of English may find English phonological rimes a contributor to cognitive burden given the limited syllable types in Chinese, English learners of L2 Chinese may not, given the much greater syllable variety of English. Even when the same languages are involved, moreover, the cognitive load issue could be different so that for instance English learners of Chinese could find Chinese syllables easier to tackle but its four tones harder. In any case, the cost of the limited exposure opportunities that define FL learning could be compensated for pedagogically, if the source and nature of the learning challenges are duly understood and the resulting difficulties are well addressed.
Footnotes
Declaration of Conflicting Interest
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The reported study was funded by National Science Council, Taiwan (grant number: 98-2410-H-006-072). The author wishes to thank the participating teachers and students from Tainan Municipal Wunyuan Elementary School, especially Dr Feng-Tzu She, the principal, Joyce Hsieh, and Jasmine Liu, for their generous participation and assistance.
