Abstract

Over the past twenty years or so, a fairly large number of introductory books on corpus linguistics have been written (e.g., Biber, Conrad & Reppen 1998; Kennedy 1998; Hunston 2002; McEnery, Xiao & Tono 2006; Teubert & Čermáková 2007; Lüdeling & Kytö 2009; O’Keeffe & McCarthy 2010; McEnery & Hardie 2011; Lu 2014), with some unavoidable overlaps but also markedly different foci and target readership. Lindquist and Levin’s (2018) Corpus Linguistics and the Description of English, the second edition of Lindquist’s (2009) book with the same title, occupies a unique niche among these books, with a primary focus on illustrating how existing corpora can be and/or have been used to examine and describe diverse types of linguistic phenomena in seven specific areas of English. This focus makes the book especially appropriate for its main target readership, i.e., university students of English language and literature with some background in the description of English but no or limited experience in working with corpora.
In the first, background chapter, the authors start by introducing the emergence of corpus linguistics as a methodology for analyzing and describing language, one that is frequently associated with a usage-based view of language use and change. The authors then present concordances and frequency figures as the primary types of results from corpora. While this classification appears to be an oversimplification, it aligns well with the ways in which analytical results are presented in the numerous examples in subsequent chapters. A brief but highly informative discussion of the criticism of corpus linguistics from armchair linguists and corpus linguists’ response to that criticism follows. The bulk of the chapter then explains the various types of corpora that exist, along with brief descriptions of some of the best known corpora for those types.
The second chapter further primes the reader with the basics of corpus linguistics methodology. Following a brief discussion of the differences and connections between qualitative and quantitative methods, the authors dive into the concept of frequency, using analyses of the most frequent words, lemmas, and lemmas of different lexical categories in the British National Corpus (BNC) to illustrate the usefulness of word frequency information. For example, the top fifty most frequent nouns, verbs, and adjectives in the BNC are listed in separate tables, and some remarks are made on the types of words that appear on these lists, such as body-part nouns, auxiliary verbs, and the word other as, perhaps unintuitively, the most frequent adjective. Such frequency information is argued to be useful for teachers and material developers when they decide what words to present to learners. Several important methodological issues in analyzing frequency information are covered, including the use of chi-square to compare frequencies, the examination of frequency distribution across different components of a corpus, and the need to normalize frequencies when comparing results from corpora of different sizes. The authors then offer a brief treatment of the notion of representativity as a relative concept, cautioning that results from a corpus should only be generalized to the language of which the corpus is representative. The chapter ends with a section on corpus annotation, including a fairly thorough discussion of part-of-speech tagging followed by a rather succinct mention of syntactic parsing and other types of annotation.
Chapters 3-9 focus on specific areas of English, in the following order: looking for lexis, checking collocations, finding phrases, metaphor and metonymy, grammar, male and female (i.e., language and gender; title original), and language change. Each chapter starts with a brief discussion of the conceptualization of the area and/or its role in language or in linguistics research, followed by reports of published research or original case studies done by the authors on a combination of methodological issues (e.g., how to identify collocates of a keyword within a window) and research topics (e.g., Dickens’s recurrent long phrases) within the area. The authors take a descriptive approach to presenting and discussing the research results, with a large number of tables and figures, making the discussion highly accessible.
A final chapter shows how the web can be used as a source for linguistic investigations and for compiling new corpora. The advantages and drawbacks of the web as corpus are both discussed. Multiple examples in the areas of phraseology, grammar, and dialectal variation are used to illustrate how the web can be exploited for corpus linguistic research. For instance, to examine variation in the verb in the phrase to outstay one’s welcome, a researcher may type in to * one’s welcome in a commercial search engine and subsequently analyze the hits returned (199-200). There is detailed discussion of several corpora compiled from the web to represent new web genres, text types, and internet registers, and the methodology for and ethical issues involved in collecting and analyzing such corpora. This chapter offers a very useful look at a highly important developing trend.
Each chapter is accompanied with a summary, four to six study questions, a set of hands-on exercises hosted on the book’s companion website, and an annotated list of further reading. The summary provides a succinct and precise synopsis of the main topics covered in the chapter. Most of the study questions are on relatively broad conceptual, theoretical, or methodological issues that help guide the readers’ thinking about the concepts, approaches, and research questions discussed in the chapter. The hands-on exercises are designed to encourage the readers to explore and analyze freely available corpora online to extract specific types of information and answer questions relevant to the topics discussed in the chapter. The annotated list of further reading offers highly useful details on relevant further literature pertaining to the topics covered in the chapter.
There are a few areas in which the book could be improved. First, in terms of organization, most of the seven areas covered in chapters 3-9 are well motivated and align with the broad areas of research in language description; the chapters on “metaphor and metonymy” and “male and female,” however, appear to sit at a lower level of generality than the other areas and could have been replaced with chapters on, for example, “analyzing meaning” and “language variation.” Second, on some topics, a specific perspective or method is introduced, sometimes without reference to other perspectives or methods of which the readers should ideally be made aware. For example, the decision to not explicitly engage in the debate on the theoretical status of corpus linguistics when defining it in the first chapter may be a sensible one, given the methodological emphasis of the book, but a brief reference to this debate (e.g., Taylor 2008) could have been included. The introduction of the chi-square test in the second chapter as the most frequently used measure for significance testing without qualifying the scenarios of such testing and with no mention of other commonly used measures is problematic. Third, while the largely descriptive approach to presenting the research results makes the book highly accessible, the readers could have also been gently exposed to the range of inferential statistical analysis commonly adopted in the published studies reported. Finally, the amount of methodological details included for the studies reported varies substantially, and the book’s reference to the computational aspects of the corpus linguistics methodology is sparse, even in the annotated lists of further reading. Consistent inclusion of more transparent and explicit descriptions of the actual steps taken in the published or original case studies or consistent reference to further reading that covers the necessary computational steps to derive the corpus results would have further increased the benefit to the readers.
These areas for potential improvement notwithstanding, the updated edition of Corpus Linguistics and the Description of English makes an outstanding introductory book on corpus linguistics. It covers the most important basics of corpus linguistics methodology in a highly accessible way, and illustrates the applications of corpus linguistic analysis in some of the most prominent areas of English language description in a nontechnical manner. The study questions, hands-on exercises, and further reading provided at the end of each chapter dramatically improve its suitability as a textbook for a first course on corpus linguistics for undergraduate and master’s students in English language and literature. Readers of the first edition will also find the new edition decidedly superior, with its new coverage of social media, DIY corpora, and ethical concerns, new materials on metonymy and pragmatics, comprehensive information on new corpora, and numerous updated example investigations, exercises, and study questions.
