Abstract
This article clarifies some ideas presented in this issue’s keynote article (Amaral and Roeper, this issue) and discusses several issues raised by the contributors’ comments on the nature of the Multiple Grammars (MG) theory. One of the key goals of the article is to unequivocally state that MG is not a parametric theory and that its current version follows very explicit minimalist assumptions. We also refine the notion of ‘minimal multiple rules’ to make their theoretical status more precise. Overall, we would like to acknowledge the important contribution of all the articles in this issue to the evolution of the theory.
To any evolving theory, criticism is more than welcome. We were pleased to see that overall the nature of the comments by this issue’s contributors acknowledge some uncontroversial assumptions about the existence of multiple grammar rules in L2 representation. 1 A number of the comments helped us to make the refinements in Multiple Grammars (MG) theory we shall discuss. In particular, we take the opportunity to clarify the following main aspects of our proposal: the abstract, representational nature of the rules of Multiple Grammars and their operation, the notion of productivity from a MG perspective, and the limited role of processing in our account at this stage.
A major claim of ours is that rules must be simple, following the spirit of minimalism. Optionality is an observable behavior of second language (L2) and first language (L1), but if rules must be simple, then optionality reflects two rules, not subcases of one rule. Thus we differentiate ‘observable optionality’ in the form of a specific speaker behavior from ‘descriptive or grammatical optionality’ in the notation of an abstract rule.
Second, the MG proposal points at the critical notion of ‘productivity’: when does a rule become productive for a child and for an adult? Productivity is marked on a rule when it reflects abstract categories. It emerges in the shifts from lexical item to lexical class, or critically to a phrasal class (VP or NP, for example). In effect, some rules are marked [+productive] when their definition allows them to apply to a general category that would include novel lexical items.
Third, a critical feature of our analysis is a recognition of the difference between production and comprehension. MG may apply differently to comprehension and production. Production applies to what a speaker says, while comprehension is forced by what a person hears. Pérez-Leroux’s comments provide a welcome confirmation of this perspective in her work on the Overt Pronoun Constraint (OPC) that supports exactly our results. Recent work by Rankin (2013) shows that even advanced German speakers of English are susceptible of analysing SVO as OVS in cases like ‘the mouse chased the cat’ as if the cat is the subject. 2 We will explore the V2 rule in detail, but first we want to clear up some confusion.
First and foremost, it is important to clarify that MG is not a parametric theory per se. 3 The simple rules that exist in the theory to handle specific language properties are not really ‘contradictory’ or ‘conflicting’. There is nothing truly ‘conflicting’ about the fact that English has an extremely productive rule to license sentences with overt subjects and a very idiosyncratic and lexically-based one to license a small subset of null subject matrix sentences, such as in ‘looks good’ or ‘seems interesting’. We used the term ‘conflicting rules’ as an explanatory device to illustrate the issue of optionality as an observable behavior from a parametric point of view. The graphs that illustrate ‘parameter setting’ are also an explanatory device that connects the classical notion of parameter with the notion of productivity in a proposal that allows for multiple rules to coexist in any given grammar. MG is a theory influenced by the notion of parameters in a sense that it does not deny the possibility of certain linguistic properties being able to trigger larger grammatical generalizations made available by UG. However, there is a clear difference between the traditional notion of parameters and the way MG represents grammatical knowledge, since the theory we presented is truly minimalist in this respect.
If we use ‘minimal multiple rules’ without descriptive optionality in them to be the core representational mechanism behind human grammars, productivity then becomes the central notion that defines both L1 and L2. The notion of productivity we advocate is based on the minimalist view where there are general language principles and a universal computational system, and where the differences between languages lie in the lexicon. By embracing minimalism, MG supports the idea that the speaker of a language L will essentially have to decide how to handle the rules that exist in UG:
They can assign a ‘productive status’ to a given rule and link it to a general category.
They can assign it a ‘less productive status’ and decide which subset of lexical items this rule can use to license possible constructions; or
They can decide that it does not apply to any lexical items in language L, which is the same as saying that the rule does not apply to L (or exist in L).
The process in which this decision occurs is an empirical question that should be investigated, and we suspect that option 1 may not be the first step in language development. Westergaard (this issue) brings up an extremely important contrast between L1 and L2, which our initial article does not adequately articulate. Let’s take the case of wh-movement. L1 acquisition seems to proceed with highly lexical definitions of rules that are sensitive to micro-cues. The set of wh-words is small and therefore it could be represented as a lexical list: who, where, why, when, how, how come, etc. Although movement triggers subject–auxiliary inversion, there is a good reason to represent the rule lexically because there is an exception: how come (1).
(1) a. How come John went home? b. * How come did John go home?
If the rule is always lexical, then children will simply learn the cases one-by-one, and in fact there is evidence that inversion is acquired one-by-one in English L1 (see de Villiers, 1991). Now suppose that L2 speakers, who have developed a general category for wh-words, allow the very same inversion rule to require the general wh-category, instead of individual lexical items. In MG terms, these L2 speakers would be assigning a more productive connection between the inversion rule and their lexical representations. In this case, the theory would predict, correctly we think, that L2 learners would produce (1b).
The idea that L2 speakers potentially assign a more productive use to certain rules by primarily searching for lexical categories instead of lexical items can be used to address Unsworth’s (this issue) rightful concern about the differences between adult and child bilingual acquisition and some of Serratrici’s
4
(this issue) comments about MG and bilingual acquisition. According to Unsworth: if it really is the case that it is the language-dependent nature of linguistic cues that makes L2 acquisition difficult, there must also be some additional factor(s), specific to adult L2 acquisition, which contribute to this difficulty, because the evidence from bilingual language acquisition suggests that the co-existence of two grammars is not a sufficient condition for problems to arise.
Indeed, as we argue below, simultaneous bilingualism could show differences from L2 acquisition. In particular, the application of over-general rules should be more restricted among simultaneous bilinguals.
The fact that MG takes minimal multiple rules to be its primary descriptive apparatus allows the theory to propose a unified description to all types of language acquisition realities, from multilingual adults to monolingual children. Liceras (this issue) in her comments points to a series of proposals for L2 representation that dealt with (observable) optionality from a parametric perspective. As she highlights, many of those approaches were successful in explaining the differences between L1 and L2 by allowing multiple parameter setting in L2. It is important to remember that MG was not originally proposed to explain L2, but rather to describe the apparent optionality that could be observed in any given language (Kroch and Taylor, 1997; Roeper, 1999), particularly during monolingual L1 acquisition. The extension of MG proposed in this issue is not meant to explain why L2 speakers behave differently from L1, but rather to approximate the representational mechanism used by both groups. From a MG perspective, bilingual or L2 grammars are not deficient or incomplete versions of L1 grammars. In fact, they could be better described as monolingual grammars on steroids.
Slabakova (this issue), Westergaard (this issue), and Lardiere (this issue) brought up V2 questions and have led us now to a more detailed discussion, illustrating MG through simple rules, what we could call a ‘representational approach to performance’. We argue that there is one V2 rule (2) which is available from UG, but which may appear slightly different in various languages through its interaction with lexical restrictions in the domain of CP (Discourse and movement to IP are different). This rule has the XP in the Spec-of-CP and moves the V into C. Two parts of the rule exhibit apparent variation across grammars: the possible content of XP and the lexical categories that could be present in V.
(2) V2 rule: XP Z V → XP V Z (where Z is any intervening string within a clause)
In English, XP applies only to quotation and stylistic fronting of PPs, both highly restricted. Therefore, for the English native speaking child to avoid overgeneralizations he or she needs to enter into the V2 rule those possible XPs, quotation and stylistic inversion, one at time as he or she hears them, which presumably is very rarely. For quotation, the rule must require that the verb be interpretable as a speaking verb, for instance ‘carry on’ as in ‘No’ carried on John insistently.
In German, the child who does not know if he or she is learning a restricted XP language, like English, also adds each XP type as he or she hears it, until at some critical moment he or she simply substitutes XP for the whole class. It is important to grasp the full range to see that experience and frequency of each category in German will vary dramatically.
(3) Typical forms of V2 a. Subject NP: He eats meat. b. LOC: There sings he. c. ADV: Quickly moves he. d. DO: Meat eats he. (4) Less discussed forms of V2 a. Quotation: ‘Welcome’ said Mr Anders. b. VP fronting: To me alone come wants he not. ‘He does not want to come alone to me.’ c. Empty Topic in Discourse: Wo ist das Fleisch? Where is the meat? __ ate John already. d. Conjunction: Hanns spielt oft, Hanns plays often, so can he without difficulty us help. ‘Hanns plays often, so he can help us without difficulty.’
The discourse case involves a linking between an NP in one sentence and an empty subject in another. The factors involved may engage what is called ‘information structure’: how topic and focus implicate intonation and, as in this case, potential null subjects, which V2 can operate upon. The discourse case is a special challenge to an empiricist theory that tries to induce patterns from simple strings. It appears as if it is a V1 structure, which would then seriously confuse any further development of a coherent grammar that assumes initial subjects. Thus the VP-fronting case, if looked at in terms of surface strings, would complicate any simple definition of XP, since it allows a huge range of XP-category sequences inside it, and therefore must be analysed as a VP unit by the child or the acquisition task would be extraordinary. These are the cases any usage-based theory must confront directly. The discourse case requires prior analysis to know that the sentence has an empty Topic-XP and is not V-initial. These XP-versions may come in late, though we have not researched the question. Notably what is excluded without evidence is a Head non-XP element such as a particle and there are no reported examples of them (5):
5
(5) * Aus schreit er. Out yelled he. ‘He yelled out.’
At some point the German child links the abstract category XP (but not X) in the V2 rule to a whole range of phrasal projections. In MG terms, the child makes the German V2 rule very productive. As Yang (2002) observes the frequency of input matches the child acquisition point and could predict that the object–verb–subject form arises later in children. If that is the point where XP is adopted rather than any specific subcases, then that is the point where general productivity is fixed. Since before XP is fixed and a child hears VP as a series of XPs (‘zu mir [XP] allein [XP] kommen [XP]’), it is a natural hypothesis that acquisition of VP-fronting would not occur until the general XP in the rule had been posited, hence later.
The English child, by contrast, hears only two environments and therefore never adopts the full XP domain but lexically marks the quotation in terms of speaking verbs of any type like (6a) and stylistic inversion cases in some pragmatically obscure way 6 . Therefore, full XP productivity never occurs while X is always excluded without positive or negative evidence for the child (6b).
(6) a. ‘never’ continued Bill. b. * up ran Bill suddenly.
The other acquisition challenge in the V2 rule presented in (2) is what lexical categories can replace V in the rule. Again the MG notion of productivity will address this issue. The element marked [+V] shows variation at the lexical level. In German it can be auxiliaries and main verbs; besides, particles are separated from their verb in the lexicon (simplifying somewhat). Thus auxiliaries can move with V2, and verbs move without their particles. In English the auxiliary modals form separate lexical categories and the particle is incorporated into the lexical description of verbs. 7 So the English child can say (7a) but not (7b), while the German child says only the latter form. The English child also can carry out quotation as if it were simple topicalization in a different rule, such as in (7c). One can even predict that the topic-quotation form precedes the other quotation forms given that children exhibit topicalization early in other environments.
(7) a. ‘Nothing’ yelled out Bill. b. * ‘Nothing’ yelled Bill out. c. ‘Nothing’ Bill said.
Thus we argue that they do not simply have ‘different’ rules, but use the broad rules with different degrees of categorial and lexical generality. Our focus upon the productivity of rules allows us to throw L1 controversies into a different light. Wexler (2011) argues that children have V2 very early, while Yang (2002) argues that they do not get it until later because the object–verb–subject examples emerge later. Our approach shows that both positions are half right: children acquire the V-movement part of the rule very early, but the XP generalization later.
Westergaard’s (this issue) comments lead us to underscore the distinction between L1 and L2 and articulate a suggestion on how L2 studies make a unique contribution to understanding UG. The L1 speaker may construct the rule with cautious sensitivity (implicitly not wanting to overgeneralize English into German if English is being acquired), thus following micro-cues as she suggests (which captures language change as well). What does the L2 speaker who has productive V2 do when learning a grammar with more restricted use? Here we could predict that the MG simple rules lead to the opposite possibility: the L2 speaker applies a general category to the V2 rule above and can be led into incorrect comprehension and production, as the evidence from many commentators confirms (see also Rankin, 2013). A conceivable prediction is that L2 deviation will be toward the full generality of the categories associated with XP and V in the rule (2) above. This would lead to an overly liberal, rather than overly restricted ‘transfer’, i.e. a more productive assignment of the specific lexical categories of each language. These variations in lexical description and rule assignments are responsible for all sorts of negative transfer phenomena. For example, while German speakers learning English will erroneously leave the particle behind, as in (8a), English speakers learning German will erroneously carry the particle along, as in (8b). Both errors could be explained by the way in which particles are treated in the lexicon, and not by differences in syntactic rules from each language. Notice that these are quite testable predictions.
(8) a. * ‘Nothing’ yelled Bill out. b. * ‘Nichts’ rief aus Fritz. ‘‘Nothing’ called out Fritz.’
The idea that the abstract rule exists can actually be obscure in L1 because of many intervening historical and idiomatic factors. If the rules are cross-linguistically applied at the abstract level we have defined them in, then it is a unique piece of evidence that the level of abstraction is appropriate. This is particularly relevant to current Minimalist discussions where the nature of Labeling is important. Does the V+particle or modal auxiliary carry a simple V label, or a more differentiated AUX or V-part label? Again, just as with lexical intrusion into L2, we would argue that the labels of lexical items can also intrude, and their behavior in movement operations then provides a window on what the label is.
Muysken has reminded us that our questions address only the technical side of L2 acquisition, leaving the profound social variables (age, class, attitude, environment) untouched. Our account has little, but not nothing, to say. The point at which a speaker makes a larger generalization can engage those attitude factors. One eager and confident L2 learner may be able to observe the lexical information more accurately and avoid assigning general phrasal categories to the abstract XP category in the V2 rule described above. In that sense, representations may allow open pragmatic and social variables to be added to them in informative ways that allow a teacher to avoid thinking that a learner’s attitude will affect every part of learning the same way: your social background may allow you to approach phonology, lexicon, syntax, and pragmatics each differently.
We turn now to the explanatory role of processing in L2. We, in no way, deny the possibility of parsing problems and the importance of processing research to understand the behavior of L2 speakers; much on the contrary, we have been pursuing collaborations in this regard with some of our research partners (see Faber et al., in progress; Lawall et al., 2012). However it is important to look at the relationship between parsing and grammars from the symbolic perspective we adopt in order to understand why our data cannot be explained by any problems in parsing. Parsing is a process that assigns a specific syntactic representation to input strings according to the grammar. In a symbolic model, a parsing algorithm needs to make an explicit reference to the grammar if it is to succeed. If the parsing algorithm and the grammar rules are incompatible, or if the necessary grammatical or lexical information is missing, that the parsing fails and no representation is generated. By this view, problems in the parsing algorithm cannot yield constant ‘non-target-like representations’. There must be something in the grammar that allows for that representation to exist in the first place.
Amaral and Leandro (2013) ran a sequence of experiments with recursive constructions in Wapichana, an indigenous language spoken in Guyana and Brazil. 8 Among the constructions tested, they looked into how adults and children in Guyana interpreted embedded genitive constructions in that language. Participants saw a picture where four different people and their respective dogs had balls and flowers of different colors. There were 10 different colors in total. Then they heard a story establishing the relationship between the characters. After that, they were asked about the color of the objects each participant had. The target questions had between 1 and 4 possessives, such as in (9).
(9) Xa’apauram Cedrick daduku minhayda’y uza bala- What Cedrick sister friend dog ball- ‘What color is Cedrick’s sister’s friend’s dog’s ball?’
Three-year-olds and four-year-olds, when faced with two possessives (‘Cedrick’s sister’s ball’), preferred the recursive (adult-like) interpretation, 87% and 75% respectively. With three (‘Cedrick’s sister’s friend’s ball’) and four possessives (such as in (9)), their rate of recursive interpretations was much lower: three-year-olds giving the recursive reading 50% of the time in both cases, and four-year-olds with 42% for three possessives and 33% for four possessives. The interesting cases for our argument comes from the non-target interpretations that varied from totally nonsensical (just a random color) to conjunctive readings (where Cedrick’s sister’s ball is interpreted as Cedrick’s and his sister’s ball). Limbach and Adone (2010) had already shown that English-speaking three-year-olds tend to drop genitive DPs in embedded constructions more often than four- and five-year-olds who tend to provide a conjunctive reading in the same situations. They also tested adult L2 speakers, and their rate of dropped embedded genitive DPs (22%) was much higher than those of native speakers (12%).
In what concerns our argument, dropping a constituent or providing a random response could be interpreted as a processing problem, i.e. the parser (for whatever reason) was not able to assign a full syntactic construction to the complete input string. On the other hand, providing a consistent conjunctive reading means that the parser was able to assign a representation to the full string, which consequently indicates that the grammar has suitable rules that allow this representation to exist. Roeper (2011) and Arsenijevic and Hinzen (2012) argue that conjunctive interpretations are a default form that reflects direct recursion, hence a specific representation is chosen. In other words, a parsing failure generates lack of consistent representations and consequently random (or no) answers, while a parsing success generates a constant and consistent representation, which may or may not be compatible with that of the monolingual adult speakers of the language. The data we present in the keynote article suggests constant interpretations of OPC responses in two different tasks; therefore we argue that the explanation for such phenomenon cannot be purely a result of processing, and that it is necessary to have an appropriate representation in the grammar that is able to license both constructions.
Having made clear that when we operate under our current assumptions about symbolic parsing it is necessary to develop a representational explanation for the facts stated above, we want to state that we fully agree with Hopp’s (this issue) and Truscott’s (this issue) comments that indicate the need for research on processing to shed some light on which conditions make learners prefer one particular sub-grammar over the other. We welcome contributions in processing research that establish clear connections to specific grammatical models under generative assumptions. In this regard, we agree with Slabakova’s (this issue) observation that the theory as presented in the keynote article needs further ‘elaboration of the psychological mechanisms in L2 acquisition’ by requiring formulable links to representations.
Slabakova also emphasizes the problem that a MG model could potentially predict a massive proliferation of rules with unwanted and not empirically plausible consequences. This is quite true, and like in the history of generative grammar, when transformations were proposed it became equally clear that they had to be constrained in precise terms, which determined the research agenda. MG is the same. We agree with Slabokova that a mechanism is necessary to restrict this potentially uncontrollable proliferation, and that some of the solutions suggested by the previous models cited by her provide interesting ideas on how such a mechanism could operate within a MG representation. At this point, we are very interested in exploring possible solutions that could further advance the MG notion of productivity. There is a lot to be said about the way in which syntactic rules operate in regards to lexical selection, and we believe that a MG model that would explain why L1 grammars seem to be more stable than L2 ones could be framed, at least in part, around this idea.
Sorace (this issue) points out that our core claims are not new, with which we agree, but we do think they constitute significant refinements of previous ideas on the interaction of grammars. Linguistic theory has not proceeded by abandoning the initial crude notion of transformation that Chomsky introduced in 1956 (Chomsky, 1956), but rather the core idea has undergone radical reformulations over 50 years to the point where traces, invisible movement, and feature-checking have made the original theory much more robust. We believe that extending the philosophy of minimalism to L2 leads to a promising research agenda where new, highly refined, predictions about L2 effects can be made.
In conclusion, we thank the commentators for their stimulating remarks and hope that, coupled with our response, we have collectively made progress in the field. We would like to re-emphasize the singular role of L2 in our argumentation. It is precisely through the unique nature of L2 that we can more easily observe the generality of simple rules. The lexical restriction for how come in English might lead one to say that there is no general lexical rule, a position that appears to be advocated by some usage-based theorists. However its generalization into other grammars indicates that the abstract wh-rule must exist on its own and must be separable from its lexical environment. The default forms revealed by bilingual speakers of Wapichana similarly indicate the necessity of the direct/indirect contrast in the formulation of recursion. In that sense, work in L2 illuminates the simple ‘minimal’ character of ‘minimalism’ in a way that is critical and vital to the minimalist enterprise. One might say, in fact, that the logic of minimalism predicts exactly this theory for which L2 work provides evidence.
Footnotes
Declaration of conflicting interest
The author declares that there is no conflict of interest.
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
