Abstract
This study created a medical word list (MWL) to bridge the gap between non-technical and technical vocabulary. The researcher compiled a corpus containing 155 textbooks across 31 medical subject areas from e-book databases (totaling 15 million running words) and examined the range and frequency of words outside the most frequent 3,000-word families along the British National Corpus scale. To reach 98% lexical coverage for adequate comprehension of medical texts, 595 of the most frequently-occurring word families in the corpus were ultimately chosen and formed the MWL, and these accounted for 10.72% of the tokens in the medical textbooks under study. Excluding highly-specialized medical terms of Greek/Latin sources, the MWL encompasses various sub-technical and lay-technical vocabularies. It is suggested that with the help of free online concordancers, medical teachers can raise their students’ awareness of the commonly-used medical words reported in this study by incorporating concordance data into teaching materials, thereby consolidating the vocabulary knowledge acquired from the MWL. For medical novices, the present MWL provides a window to the medical register.
I Introduction
English is not an official language in Taiwan. In 2003, Taiwan’s Ministry of Education published a basic English word list of 2,000 most commonly-used words as a curricular standard for the 3-year English course design for junior high schools. Since then, the 2,000 lexical items have been presumed to be the target vocabulary that junior high school graduates entering senior high schools should master.
For admission to various universities and colleges, Taiwanese senior high school graduates must take college entrance exams (General Scholastic Ability Test or Department Required Test). A high score is desired as an admission criterion to prestigious institutions. Both of the college entrance exams include English tests, which involve a vocabulary of more than 4,000 most frequent word families (College Entrance Examination Center, 2011). Medical colleges/universities in Taiwan rank top due to intense competition. Those who are accepted to a medical school always perform the best (full marks or nearly) in the tests of various academic subjects, including the English subject. Therefore, after three years of senior high school education, matriculating medical students are supposed to have an English vocabulary of 3,000 words at the minimum.
During the 7-year medical program with the final two years being hands-on training at a teaching hospital, medical undergraduates need to heavily rely on their abilities to read specialist textbooks in the field, mostly in English. The first step to access medical English is to learn medical vocabulary. The course Medical Terminology is a required course for medical students to meet the demands of their future jobs, such as writing diagnosis and case reports. Medical terminology here refers to highly technical words of Latin and/or Greek origin that are used almost exclusively in medical contexts. For instance, kardia (Greek)-cardi means heart while brachium (Latin)-brachi means arm. Approximately 75% of the medical terms are either borrowings derived from Latin or Greek, or neologisms with Latin or Greek elements (Salager, 1985). They involve descriptive anatomy (e.g. ileum and endothelium), epidemiology and semiology (e.g. nephritis and epilepsy) as well as chemical compounds (e.g. potassium and thallium uptake). According to Hutton (2006), most of the fully-technical medical words are made up of word roots and affixes (i.e. prefix and suffix). Any single medical term has at least one root/stem determining its meaning and one or more prefixes or suffixes to change the part of speech or change the meaning of the word. For example, it may not be difficult to guess the meaning of esophagoenterostomy – a surgical formation of a direct communication between the esophagus and intestine – if one knows the word root esophagus from Latin êsophagus, the prefix entero- from Greek enteron and the suffix -stomy from Greek stoma, which in turn means the gullet, intestine and an opening (a surgical operation in which an artificial opening is made into a specified organ or part). As can be seen from the above, medical terminology – fully-technical terms – has constancy of meaning unmodified by association and has only denotations. It is more easily manipulated than other kinds of vocabulary, as Hoffmann (1981) stated. As for learning medical terminology, Fang (1985) identified two successful strategies for Taiwanese students: (1) using affixes and roots to engage learners to analyse word structures and (2) finding the relationship between pronunciation and spelling of medical terminology.
In view of the constancy of denotation for fully-technical vocabulary, the present study is more concerned with sub-technical vocabulary (e.g. component, syndrome, synthesis and mechanism) and lay-technical vocabulary (e.g. immune, enzyme and metabolism) beyond the general English vocabulary that medical freshmen have learned in secondary education. Sub-technical words are the common words which occur with special meanings across different disciplines, and their specialized meanings may not be apparent. Lay-technical words are those that, although usually found in a specialized text, can be easily understood by the layperson.
As mentioned above, Taiwanese matriculating medical students may have a greater base vocabulary (i.e. over 2,000 words). When students’ vocabulary reaches the level of, for instance, the top 3,000 English words, their concern may be shifted to less frequent but urgently-needed words in their fields of study. It is worth mentioning here that some students may even have a vocabulary of over 4,000 words while others may not. If we set the 4,000-word level as a starting point for the selection of medical words, we would eliminate some medical words which may occur between the 3,000–4,000 word levels and may not be known to the students with a vocabulary size of only 3,000 words. On the other hand, if we sift frequent medical words from the top 2,000 general English words, the medical word list would include many words that our students have already acquired. As such, adopting the 3,000-word level as a point of departure seems to more fit the current EFL medical context, compared with other corpus-driven studies in the literature using a vocabulary size of 2,000 words as an investigation basis.
Excluding the general vocabulary at EFL learners’ command as well as ubiquitous, esoteric terms in medical texts, which would be taught in an adjunct course Medical Terminology, this study aimed to establish a medical word list, which may provide a direct access to the most frequently-used medical vocabulary for medical novices. In other words, the more restricted, specialized words with high frequency of occurrences may be the next set of vocabulary for medical undergraduates to learn after the top 3,000-word level, and in parallel to medical terminology learning.
II Literature review
1 Vocabulary size
According to Nation (1990), there are around 54,000 word families in English. A well-educated native speaker of English has a vocabulary around 20,000 words, excluding proper names and transparently derived forms (Nation, 2006). Vocabulary may be a good predictor of reading comprehension (Hu & Nation, 2000; Qian, 2002). A rich vocabulary makes a reading task easier to perform and limited vocabulary may be a major impediment to fluent reading comprehension. EFL learners like ours with limited English vocabulary may face quite an amount of dictionary work when reading a professional text. If Nation’s (2001) estimate that native speakers read about 10–12 books per year to acquire 1,000 words is correct, then a gradual increase of vocabulary due to less English exposure in the EFL context may not easily cover the lexical gap between 3,000 words and 20,000 words in a short time. Fortunately, not all English words are equally important in different phases of language learning. It has been proposed (e.g. Coxhead, 2000; Nation, 2001; Nation & Waring, 1997) that for different purposes or in different stages of learning, some words deserve more attention and effort than others. Consequently, targeting a more restricted vocabulary with relatively high frequency of occurrences in order to raise the lexical coverage may be more practical in this regard.
2 Vocabulary types
Nation (2001) divided vocabulary into four categories: (1) high-frequency or general service vocabulary, (2) academic/sub-technical vocabulary, (3) technical vocabulary and (4) low-frequency vocabulary. High-frequency words refer to those basic general service English words that constitute the majority of all the running words in all types of writing. Academic vocabulary with medium-frequency of occurrence – also called sub-technical vocabulary (Cowan, 1974) or semi-technical vocabulary (Farrell, 1990) – is a class of words between technical and non-technical words and usually has technical and non-technical implications. It covers lexical items that are neither specific to a certain field of knowledge nor general in the sense of being everyday words. In contrast to sub-technical vocabulary, technical words are the ones used in a specialized field and are considerably different from subject to subject. They are context-bound, topic-dependent frequently-occurring words in a given field. Low-frequency words are rarely used lexis.
Among the four types of vocabulary, sub-technical vocabulary, which lies in the middle area between highly-specialized and general categories, may be elusive and confusing for many learners and is itself made up of several kinds of vocabulary, which may require different teaching techniques. Baker (1988) identified six types of sub-technical vocabulary, each of which plays a rhetorical/organizational role in structuring the writer’s arguments and serves as clues by which the reader can interpret the writer’s intentions and evaluations. They are the items (1) which express notions general to all or several specialized disciplines, (2) which have a specialized meaning in one or more disciplines, (3) which are not used in general language but which have different meanings in several specialized disciplines, (4) which are traditionally viewed as general vocabulary but which have restricted meanings in certain specialized disciplines, (5) which are used to describe and comment on technical processes and functions, and (6) which are used in specialized texts to perform specific rhetorical functions.
Adopting a simpler categorization scheme, Fraser (2003) classified sub-technical vocabulary into two kinds: discourse-organizing words and cryptotechnical words. The term ‘cryptotechnical’ was coined by Howard (1991) to define words which, in addition to their general meanings, have a specialized meaning in a particular discipline. Fraser (2003) further distinguished lay-technical vocabulary from technical terms.
Using an anatomy text of 450,000 running words, Chung and Nation (2003, 2004) designed a 4-step rating scale to measure the strength of the relationship of a word to a particular specialized field. Different from the divisions between the GSL (General Service List), AWL (Academic Word List), and technical vocabulary, words were classified as being technical and non-technical on a four-point scale. They distinguished two kinds of technical vocabulary: common technical words (at step 3) that are closely related to a specialized field, may occur in non-specialized usage and may be familiar to people without specialist knowledge of the field (i.e. words that occur in general use with little change in meaning), and specific technical words (at step 4) that are largely unique to a particular specialized field and are not likely to be known in general language. Therefore, technical vocabulary may include words that come from the high frequency words or the AWL with a technical sense. Namely, specialized medical lexis is not only made up of esoteric terms but also words taken from the ordinary speech and given a new technical dress for a new use. Their data results showed that technical words (at steps 3 and 4) accounted for as high as 31.2% lexical coverage of the anatomy text, as some crypto-technical GSL and AWL words were also included. Chung and Nation (2003) concluded that technical words should be much more frequent in the technical corpus.
In general, the identification of a class of words between non-technical and technical words aims to help learners to understand the different ways in which the message is being developed and raise their awareness of some hidden technical meanings extended from general meanings.
3 Medical discipline-specific word list
Coxhead (2000) compiled a corpus of around 3.5 million running words from university textbooks and materials from four different academic areas (law, arts and commerce as well as science), and identified 570 academic word families and formed a list – the so-called Academic Word List (AWL) – which was claimed to cover 10% of the total words in a general academic text. Her research suggested that for learners with academic goals, the academic word list contains the next set of vocabulary to learn after the top 2,000-word level. To put it concretely, greater lexical coverage is gained by moving on to learning 570 academic words (10% coverage on the average) than by continuing to learn the next 1,000 words (‘3–5%’ coverage for the 3rd 1,000 words; Nation, 2006, p. 79) after the top 2,000 word families on a frequency list.
Coxhead’s (2000) interdisciplinary AWL has stimulated the establishment of all sorts of word lists specific to a certain field. In spite of its important coverage (about 10% as a whole), the advocacy of Coxhead’s (2000) AWL for the establishment of academic vocabulary in specific purpose courses has been questioned recently. Hyland and Tse (2007) raised a doubt about the widely held assumption that students need a single core of high-frequency words for academic study because they are common in an English academic register. They argued that each subject discipline has its own ways of explaining experience and its forms of argumentation with lexical variations that merit research within disciplines.
Using Coxhead’s (2000) AWL, Chen and Ge (2007) confirmed that academic vocabulary had a high lexical coverage and dispersion throughout medical research articles (RAs) and served some important rhetorical functions, but they concluded that the AWL was far from complete in representing the frequently used medical academic vocabulary in medical RAs and proposed a medical academic word list. Following Chen and Ge’s (2007) research, Wang, Liang and Ge (2008) built a corpus comprising 1.09 million running words found in medical research articles from online resources in a bid to establish a medical academic word list (MAWL) after the GSL. The 623-word MAWL (including 342 words from the 570-word AWL) accounted for 12.24% of the tokens in the medical RAs. The high frequency and the wide text coverage of MAWL throughout medical RAs substantiated the idea that a more restricted, discipline-based lexical repertoire as a teaching instrument plays an important role in ESP contexts.
Despite that the MAWL contained 281 medical-related academic words outside the AWL (623 − 342 = 281), the outcome that more than a half of the 623-word MAWL overlapped with the 570-word AWL implies that the MAWL still offers a too general academic vocabulary, which may result in medical English learners’ lack of exposure to the discipline-specific academic vocabulary that they will need. Moreover, 384 of the 570 words in the AWL are within the range of the top 3,000 most frequent words, which EFL medical students may have already known. For example, area, code, data, final, period, role, sex, score and site are contained in both the AWL and MWAL. These words are also listed in the basic 2,000 English words, announced by Taiwan’s Ministry of Education as the benchmark for junior high school English curriculum. That is, medical students excelling in the nationwide college entrance exams should have had a grip of English words at the junior high school level. The general nature of the MAWL, in contrast with its original purpose (namely, the establishment of a more focused, medicine-specific academic word list) may be ascribed to the lower selection criterion, namely compiling medical words outside the first 2,000 general English words rather than outside the top 3,000 most frequent words.
In view of notoriously difficult lexis in pharmacology, Fraser (2007) built a 601-word Pharmacology Word List (PWL). This single discipline-based word list, together with the most frequent 2,000 words (the GSL; West, 1953) and Coxhead’s (2000) AWL, provided only 88% coverage of a corpus of pharmacology research articles (RAs). By breaking down the divisions between general, academic and technical vocabulary, Fraser (2009) then created a Pharmacology Word list as the core vocabulary of pharmacology, containing 2,000 most frequently-used word families in the Pharmacology RA Corpus. Comparing with the 70.5 % lexical coverage of the Pharmacology RA Corpus by the GSL2000 and the AWL570 (totaling 2,570 word families), the 2,000-word PWL gives a far higher coverage of the Pharmacology RA Corpus (89.1%). Regardless of a 601-word PWL after the GSL and the AWL or a 2,000-word PWL that does not distinguish between general, academic and specialized vocabulary, what both mean is that learners will spend their limited time on words they only need to know.
4 Lexical coverage
Lexical coverage refers to ‘the percentage of running words in the text known by the reader’ (Nation, 2006, p. 61). The assumption made behind the lexical coverage is that there is a lexical knowledge threshold that marks the boundary between having and not having sufficient vocabulary knowledge for adequate reading comprehension. A lexical threshold is contingent upon the lexical coverage predetermined, because certain coverage points may symbolize a probability of gaining a certain level of comprehension. Essentially, more lexical coverage is better than less coverage, although 100% lexical coverage does not ensure 100% comprehension (Hu & Nation, 2000; Schmitt, Jiang & Grabe, 2011). Past studies have differed in the amount of lexical coverage that is needed for adequate comprehension to occur. Two putative coverage percentages in the literature have been suggested regarding the vocabulary threshold for successful reading comprehension: 95% for minimally acceptable comprehension (Laufer, 1989; Laufer & Ravenhorst-Kalovski, 2010) and 98% for optimal comprehension (Nation, 2001, 2006). These two coverage points signal the possible lower and upper boundaries associated with a lexical threshold and then determine the vocabulary size necessary to understand the text. The present study chose 98% as the aim since 95% seemed a relatively low threshold; therefore a more rigorous standard, 98% lexical coverage, was used as an index for calculation.
As opposed to the distinction among technical, sub-technical and non-technical vocabulary in the literature, the present study adopted a different division: that is, (1) the most frequent 3,000 word families within the grip of EFL medical students, (2) sub-technical and lay-technical vocabulary beyond the top 3,000 words, and (3) fully-technical terms (i.e. medical terminology of Greek/Latin sources). The most frequent 3,000 words may cover some words with technical meanings, which are likely to be known by the layperson or will be unproblematic to learners at this level. As mentioned above, Medical Terminology is a required course for medical undergraduates in Taiwan, since medical terms are inevitable in the medical field. This research therefore narrowed the focus on the second category. Deducting the lexical coverage of medical terminology and abbreviations and EFL students’ range of vocabulary, the study aimed to construct a high-frequency, wide-dispersion medical word list (dubbed as the MWL) with its cumulative lexical coverage contributing to reaching 98%. This restricted lexical repertoire, targeting medical learners’ specific need for the next set of vocabulary to learn, may contain sub-technical vocabulary as well as lay-technical vocabulary common to all medical disciplines. It is hoped that the vocabulary included in the MWL would be useful to medical undergraduates of different subspecialties.
This study sought to answer the following two questions:
What percentage of the running words in the Medical Textbook Corpus does the Medical Word List (MWL) need to cover in order to help to reach 98% lexical coverage?
What high-frequency words beyond the top 3,000 most frequent words make up a MWL?
III Research method
1 The corpus
Different from Wang, Liang and Ge’s (2008) MAWL derived from medical research articles for academic purposes at the advanced level, the present research targeted medical textbooks, because medical textbooks are first and foremost indispensable learning materials for medical undergraduates. The present corpus encompassed what medical teachers and students actually use in class.
The researcher compiled a corpus covering 155 medical textbooks across 31 medical subject areas from online resources, totaling approximately 15 million tokens (hereafter referred to as the Medical Textbook Corpus). All the sampled medical textbooks (five textbooks being selected for each sub-discipline) were downloaded full-text from the e-book databases such as My iLibrary ebook, Netlibrary, Medical Online, McGraw-Hill ebook, OVID LWW Collection, Oxford Scholarship ebook Online, CRCnetBase and ABC-CLIO ebook, which were purchased by the I-Shou University and can be used freely for research purposes (see Appendix 1). The subject areas included in the Medical Textbook Corpus based on the discipline of Medicine and Dentistry of ScienceDirect Online involved: (1) anesthesiology, (2) allergology/immunology, (3) alternative/complementary medicine, (4) cardiology, (5) dermatology, (6) dentistry, (7) endocrinology/metabolism, (8) emergency medicine, (9) forensic medicine, (10) gastroenterology, (11) hematology, (12) hepatology, (13) health informatics, (14) urology, (15) infectious diseases, (16) intensive care medicine (17) neurology, (18) nephrology, (19) obstetrics/gynecology, (20) oncology, (21) ophthalmology, (22) orthopedics/rehabilitation, (23) otorhinolaryngology, (24) perinatology/pediatrics, (25) psychiatry, (26) pathology, (27) pulmonary/respiratory medicine, (28) public health, (29) radiology, (30) surgery and (31) transplantation. Consequently, there were 31 sub-corpora/files. Excluding tables, notes and references, every sub-corpus consisted of an approximately equal number of running words, including almost 500,000.
2 The instrument
The present corpus was run on the RANGE program (Nation & Heatley, 2005) to calculate lexical coverage and frequencies. (The RANGE program and the word lists are free to download, available at http://www.victoria.ac.nz/lals/about/staff/paul-nation.)
RANGE is installed with fourteen 1,000 word lists made from the British National Corpus (BNC) plus proper nouns and marginal words. This series of ranked word lists involve word families. The criteria used in the RANGE program to make word families are based on Bauer and Nation (1993) six-level word building processes, which include inflections and over 80 derivational affixes. Word families are regarded as an important counting unit in terms of the learning load (Nagy, Anderson, Schommer, Scott, & Stallman, 1989). The concept of a word family is used to represent a group of words whose meanings can be inferred when the meaning of the base form in the group is known to a learner. For instance, the headword accommodate is grouped with its members accommodates, accommodated, accommodating, accommodation and accommodations to form a word family. Thus, the five family members are counted as the same word accommodate. With this computing principle, the RANGE program would read all inflections or derivatives of a word as its base form and count the range and frequency of them as one word family.
Moreover, the fourteen BNC 1,000 word list for use with RANGE provide an estimate of the vocabulary level of a text, because fourteen 1,000 word lists are ranked in terms of how frequently/commonly they occur. In the BNC scale, the 3rd 1,000 words are less frequent than the 2nd 1,000 words and more frequent than the 4th 1,000 words. By analogy, the 3,000-word level refers to the vocabulary of a text reaching a level that embraces the 1st, 2nd and 3rd 1,000 most frequent words of English.
Words that are not found in the most frequent 14,000 word families are classified as proper nouns (in the Basewrd List 15 in the RANGE program), marginal words such as oh, uh, mmm and ah (in the Basewrd List 16), and Not in the lists.
3 Data processing
The data collected from online resources at the library were first changed to conform to the spelling used in the British National Corpus word lists, with the aid of the spelling check function in Word. If the word forms were not changed into British spelling, they would have been classified by RANGE as words that are Not in the lists, namely beyond the 14,000- word level. After removing prefaces, illustrations, tables, notes and references, each Word file was then transformed to a text file. As for 155 medical textbook files, the Corpus Builder2 program (available at http://www.lextutor.ca/tools/corpus_builder2) can assemble text files (~.txt) as a single combined file for analysis on the RANGE program.
The proper nouns listed in the 15th BNC word list should be included in the cumulated lexical coverage. Nation (2006) took proper nouns into account in his analysis of novels, newspapers, graded readers and movies, because they were considered as having a minimal learning burden. Students can recognize the name of a person or a place from its spelling and transliteration without much effort. In this research, the coverage figures of proper nouns were added to the number of the ranked 1,000 high-frequency word lists needed until an accumulation of lexical coverage approached 98%. It is important to note that the 15th BNC word list in the RANGE program has over 13,000 entries of proper nouns, they are still not enough to account for all of the proper nouns in a corpus and hence some of the proper nouns in the present corpora may be classified by RANGE as words Not in the lists (words less frequent than the most frequent 14,000 word families). Proper nouns which appeared in the category Not in the lists were added to the 15th word list in RANGE.
As to spoken interjections and exclamations (so-called marginal words) listed in the 16th BNC word list, they were not factored in for this research. Since the present Medical Textbook Corpus was written texts, the very rare occurrence of marginal words (e.g. huh, erm, ooh and whew) with informal and spoken nature in the academic genre could be ignored.
In the process of data treatment, it was also noticed that there were many hyphenated compound words in the corpora, which are not listed in any of the BNC word lists. They were, for instance, acute-care, acid-base, adjunct-effect, aids-associated and clot-dissolving. It was not difficult to infer the meaning of these compounds from their individual components, which have already been included in the BNC fourteen 1,000 word-family lists. The hyphens were therefore replaced by spaces.
In addition to the 1st–14th 1,000 base word lists as well as the 15th word list of proper nouns and the 16th word list with spoken interjections and exclamations installed in the RANGE program, the medical terminology list was compiled and became the 17th word list. The medical terminology list was built mainly based on the book An introduction to medical terminology for health care by Hutton (2006), which contains a list of word components in forming medical terms. The term components involve 223 prefixes, 296 suffixes and 861 word roots, totaling 1,380 entries in the list. It is worth mentioning here that 1,380 items are still not enough to account for all the medical terms, as a medical fully-technical term is made of at least one word root and one or more affixes, and therefore 223 prefixes, 296 suffixes as well as 861 stems result in various combinations. Take the term component arthro meaning joint, as an example. A very large number of words can be generated by combining arthro as a stem, zero to several prefixes and suffixes as well as other word roots, for example, arthrocentesis (arthro + centesis, meaning ‘surgical puncture to remove fluid from a joint’), arthrodesis (arthro + desis, meaning ‘fixation of a joint by surgery’) and olecranarthropathy (olecran/o~elbow + arthro~joint + pathy~disease, meaning ‘disease of the elbow joint’). Due to diverse combining forms, some of the medical terminology in the present corpus may be classified by RANGE as words Not in the lists. Medical terms which appeared in the category Not in the lists were added to the 17th word list.
Last but not the least, medical abbreviations and acronyms appear in the medical register so frequently that they cannot be ignored (e.g. CT, MRI and RNA). The 18th word list was hence created. Hutton (2006) listed 1,143 commonly-used medical abbreviations and acronyms in his book and they were entered in the 18th word list. Likewise, some of the medical abbreviations and acronyms in the present corpus may be classified by RANGE as words Not in the lists. Medical abbreviations and acronyms which appeared in the category Not in the lists were added to the 18th word list and it was dubbed as the medical abbreviations list.
Finally, the modified corpus and base word lists were run on the RANGE program.
4 Selection criteria
Since this study set out from the minimal vocabulary size of EFL medical undergraduates (as mentioned earlier, a conservative estimate of 3,000 word families) and aimed to develop a medical word list with high frequency, three selection principles were as follows: specialized occurrence, range and frequency of a word family. If words occurred with very high frequency but appeared in only one or two medical subject areas, they would not be included in the MWL. This is due to the fact that a word count mainly based on the frequency may have been biased by topic-related words. Appearing across at least a half of the sub-disciplines as the range criterion for inclusion followed Coxhead (2000) in developing the AWL. The following number denotes the selection priority in sequence.
Specialized occurrence: The word families included had to be outside the top BNC 3,000 most frequently-occurring words of English.
Range: Members of a word family had to occur at least in more than a half of the 31 medical subject areas.
Frequency: Members of a word family, taken together, had to occur at least 863 times in the Medical Textbook Corpus (explanation as follows).
The cutting point of occurring at least 863 times was set after repeated trials until the lexical coverage of MWL contributed to reaching 98% (that is, lexical coverage of BNC 3,000 + proper nouns + medical terminology & abbreviations + MWL = 98%). If we adopted the frequency of occurrences less than 863 times as the selection criterion, the words included in the MWL would make up a cumbersome word list that contains more vocabulary than students urgently need.
IV Results
Research Question 1: What percentage of the running words in the Medical Textbook Corpus does the Medical Word List (MWL) need to cover in order to help to reach 98% lexical coverage?
The Medical Textbook Corpus contained 15,016,553 running words (tokens) and involved 10,668 word families (2,983 + 7,685) listed in the BNC 14,000 word list plus 5,952 proper nouns, 3,474 technical medical terms and 1,427 abbreviations & acronyms as well as 44 words Not in the lists. The first 3,000 word families from the BNC accounted for 70.68% of the running words, and this shows the relatively difficult nature of medical texts to read. In other words, if a student has a vocabulary of the most frequent 3,000 word families, a lack of familiarity with 29.32% of the running words in a text (equivalent to one unknown word in every 3.41 running words) may make reading an uneasy task. Even though one has already had a good command of medical terminology & abbreviations plus proper nouns, the cumulative coverage figure 87.28% (70.68% provided by the top 3,000 words + 1.44% by proper nouns + 14.39% by medical terminology + 0.77% by medical abbreviations) is still shy of 98%. EFL students may consider a vocabulary load of 12.72 unknown words per 100 words (87.28% known) more difficult reading, compared with 2 words unknown per 100 words (i.e. 98% known, the assumed threshold for adequate comprehension).
In reply to Research Question 1, the goal for the establishment of MWL was 10.72% lexical coverage, based on what was left after the BNC 3,000 and medical terminology & abbreviations plus proper nouns were counted (10.72% = 98% − 70.68% − 1.44% − 14.39% − 0.77%).
Research Question 2: What high-frequency words beyond the top 3,000 most frequent words make up a MWL?
Following Research Question 1, what remained after removing the items, BNC 1st–3rd 1,000, medical terms & abbreviations and proper nouns, was BNC 4th–14th 1,000 and Not in the lists. The next step was to select the most frequent words from these two items to be included in the MWL until their cumulative coverage approached 10.72% (to help to reach 98% altogether), as per the answer to Research Question 1. Table 1 shows that the words spreading among the BNC 4th–14th 1,000 word lists covered 12.42% in tokens and involved 7,685 word families, and the words beyond BNC 14,000 accounted for 0.3% of the running words in the present corpus. It should be noted here that before the data processing (see Section III.3), there were many frequent words occurring beyond the BNC 14,000. After the 17th and 18th word lists for fully-technical medical terms and abbreviations were compiled and added, the number of the words outside the BNC 14,000 was greatly reduced. They were classified by RANGE as medical terminology and abbreviations in the 17th and 18th word lists.
Lexical coverage of the BNC word lists and the medical terminology & abbreviations in the Medical Textbook Corpus.
With the aid of EXCEL, which can sort word frequencies in a descending order and can do arithmetic operations, 595 words were ultimately chosen and formed the MWL, whose cumulative coverage arrived at 10.72% and contributed to reaching 98% with the rest of the words lists.
All of the MWL word families appeared across more than a half of the 31 medical subject areas (i.e. dispersion in at least 16 medical sub-disciplines). When the cumulative coverage reached 10.72%, the last word for inclusion in the MWL (Item 595 in the EXCEL worksheet, ranked in an order of decreasing frequency) was vesicle, of which the total occurrences together with its family members were 863 times in the corpus containing 15 million running words and ranked in the BNC 12th 1,000-word level. Therefore, tracing back to the selection criteria for occurring frequency, the dividing point for 98% lexical coverage was that members of a word family altogether had to occur at least 863 times in the Medical Textbook Corpus.
Appendix 2 displays a full list of high-frequency medical interdisciplinary words after the BNC 3,000. Apart from the frequency of occurrences and range (occurrences across medical sub-disciplines), it also displays the word levels along the BNC scale. There were 122 words belonging to the BNC 4th 1,000-word level, 90 and 47 words from the BNC 5th and 6th 1,000 in turn (for the rest, see Table 2).
The number of the MWL in each BNC 1,000-word band and the cumulative coverage in the Medical Textbook Corpus.
By breaking down the divisions along the BNC 4th–14th 1,000 word-frequency levels, the 595-word MWL provided 10.72% coverage of medical texts, which was much better than Nation’s (2006) investigation into the coverage of the BNC ranked 1,000 word lists across several types of discourse (3% coverage of texts for the 4th and 5th 1,000 together; 2% collectively for the 6th–9th 1,000; less than 1% for the 10th–14th 1,000).
Moreover, referring back to Table 1, which demonstrates that 7,685 word families dispersing among the BNC 4th–14th 1,000 word lists covered 12.42% and contributed to 100% lexical coverage together with the other word lists, the 595-word MWL (contributing to 98% lexical coverage) was apparently a more reduced, more restricted word list. In terms of the return on learning efforts, it seems to be more worthwhile to focus on the 595-word MWL than additional 7,090 words (= 7,685 − 595). The former with only 595 words offered 10.72% lexical coverage while the latter with 7,090 more words gave merely 2% extra. A bold guess is that many of the 7,090 words may be of little immediate use to EFL medical students in terms of their very low percentages to the total. Consequently, 595 words appear to be a modest, plausible lexical target for learning.
It should be heeded here that there were 30 low-frequency words outside the BNC 14,000, occurring frequently in medical texts. They were, for instance, metastatic/metastases occurring 2,767 and 2,075 times respectively and both appearing in 28 medical subjects. Other examples include edema (in 30 medical subjects; 3,056 times), resection (28; 2,828), adrenal (27; 2,657), androgen (20; 2,649) and allograft (20; 1,905). They were the remaining words (thrown into the category Not in the lists by the RANGE software) neither in the BNC 14,000 word lists nor in the medical terminology list. Nevertheless, they should not be ignored as they frequented medical texts.
As is known, the division between fully-technical and lay-technical vocabulary is not always distinct. In some cases, an arbitrary decision was made to distinguish fully-technical vocabulary from lay-technical vocabulary. In this research, those with Latin/Greek origins were treated as fully-technical words and were assembled in the 17th word list: the medical terminology list. The words in this list were not considered for the inclusion in the MWL, since they would be taught in the course Medical Terminology. Medical abbreviations and acronyms were also excluded from consideration because students need to learn full forms first, before they are familiar with their short forms. As a result, the ranked BNC 4th–14th 1,000 word lists and Not in the lists were presumed as where lay-technical and sub-technical words may occur.
To sum up, totally 595 word families were found to have occurred beyond the BNC 3,000 (having fulfilled the Selection Criterion 1: Specialized occurrence) in more than 16 medical subject areas (Criterion 2: Range), and at least 863 times (Criterion 3: Frequency). These 595 words, directly from the users’ target texts, are the commonly-used words traversing the subfields of the medical domain. Regardless of whether they major in neurology, dermatology, otorhinolaryngology or other medicine-related subjects, medical students may encounter these words very often since they are shared by different medical specialist groups.
V Discussion
After 595 frequent medical interdisciplinary words beyond the BNC 3,000 were identified, a need for a further examination of these words in context followed, since it may not be easy to make out a word in isolation.
During the word screening, it was found that out of the 595-word MWL, there was only an overlap of 76 words with Coxhead’s (2000) 570 interdisciplinary academic words (AWL), scattering among the BNC 4th–10th 1,000 word lists, as opposed to Wang, Liang and Ge’s (2008) 623-word MAWL with 342 words from the AWL. They were, for example, enhance, formula, manipulate and sustain. A closer look at these words shows that a sizeable proportion of them have a hidden medical sense in addition to their commonly known meanings. For example, enhance in the medical register mostly means making greater treatment in effectiveness while it often refers to providing with advanced, improved or sophisticated features in computer software. The word formula has several meanings: (1) a symbolic representation of the composition of a compound, structure or molecule in chemistry, (2) an equation in mathematics, (3) a recipe or prescription for the preparation of a medicine, (4) an established model or approach to do something, (5) an utterance of conventional notions or beliefs for use in an occasion or a ceremony. The words with polysemous features may present learners with different degrees of difficulty.
In view of the hidden technicality a generic academic word may contain, it appears that there is a continuum of specialty for sub-technical and lay-technical words, ranging from the vocabulary of which the technical sense is frequently an extension of the general meaning, to the vocabulary of which the technical sense is primarily used. Consequently, referring to Baker’s (1988) and Fraser’s (2003, 2009) categorization of sub-technical and lay-technical vocabulary, the characteristics of the present MWL may be classified as follows:
Words, themselves and/or their family members, express some academic notions, approaches or procedures, and can be found across a wide range of disciplines (e.g. component, device, differentiate, mechanism, entity, undergo and parameter). These words can be found in the academic domains of social science, humanities and science & technology as well. The 76 AWL words in the MWL are the sub-technical vocabulary of such type.
Words of general use whose technical meaning may be hidden and only emerge from the context (e.g. ‘susceptible to colds’) or from the words that collocate with them (e.g. ‘tear duct’, ‘skin graft’, ‘anesthesia induction agents’, ‘dilate pupils’, ‘adrenergic receptor’, ‘respiratory tract’ and ‘cerebrospinal fluid’). Words in this category are usually not difficult for medical students to guess their technical meaning, as the hidden technical sense is closely related to their core meaning and can be viewed as a derivative of their general meaning. For example, the words acute in ‘acute pain’, circulation in ‘blood circulation’, conduction in ‘nerve conduction’, primary in ‘primary headache’ (versus ‘secondary headache’), placebo in ‘placebo effect’, seizure in ‘heart seizure’ and disorder in ‘functional disorder’, which have acquired a medical connotation, are the lexical items borrowed from non-technical/general language.
Words are used equally with general and specialized meanings (e.g. vessel, fracture, vein, serum, plasma, transplant, infiltrate, protein, glucose, posterior, anterior and lateral). These are words that may invisibly slip out of the medical field and into other specialized fields or everyday conversation. A ‘blood vessel’ versus a ‘shipping vessel’ or a ‘vessel holding liquids’; a ‘bone fracture’ (meaning ‘a break/crack’), a ‘fracture in the water pipe’ (‘a rupture’) and ‘reputation was fractured’ (denoting ‘destroy’); ‘blood plasma’ (in ‘blood transfusion’) and ‘plasma TV’ (in modern appliances); ‘gold vein’ (in geology), ‘vein’ as ‘a blood vessel’ (in medicine), ‘vein in a leaf’ (in botany), ‘vein in the wing of an insect’ (in zoology) as well as ‘comedy vein’ and ‘in the vein’ for jokes (in everyday conversation) are all such examples.
Words with a medical dress may undergo a semantic transfer when used in general language (e.g. trauma, nausea, immune, malign, morbid, cataract and chronic). The meaning may even become different from its original denotation after a shift from the medical field to another field. Such a feature can be shown in these instances like ‘removing a cataract’ versus ‘cataracts of rain’ and ‘morbid anatomy’ (pathological) versus ‘a morbid interest in horrors’ (unwholesome for general meaning).
Words reveal a technical sense in connection with anatomical, biochemical, demographic, epidemiological, semiological and topographical medicine, mainly used in the medical register (e.g. carcinoma, catheter, distal, pulmonary, renal, prostate, membrane, secretion, artery, sinus, steroid, mucus, enzyme, excrete, abscess, retina, pituitary). This category of words may cover lay-technical vocabulary (Fraser, 2003, 2009), namely, those medical words are easily understood by the layperson, e.g. gene, hormone, implant, relapse, suture, vascular and metabolism.
Words are used almost exclusively in the medical contexts (e.g. tomography, hyperplasia, biopsy, cyst, cirrhosis, necrosis, prognosis, dysplasia, macrophage, pathogenesis, thrombosis, sepsis, lesion and infarct).
The above six categories of MWL signify various degrees of technicality in an increasing order. It was found that the words in Category 6 with fully technical nature and nearly exclusively used in the medical domain were mostly spread in the latter bands of the BNC word-frequency scale (e.g. tomography, dysplasia and hyperplasia in the BNC 14th 1,000 word list; necrosis and thrombosis in the BNC 13th 1,000 word list). However, words in Category 5 or below (from lay-technical to general sub-technical vocabulary) scattered sporadically along the BNC scale. This again reflects that the learning sequence after the most frequent 3,000 words should be the MWL, which provides a direct access to the most frequently-used medical vocabulary for EFL medical novices, rather than continuing to progressively learn the BNC 4th–14th 1,000 word lists in turn.
The words from the 1st to the 4th categories tend to carry more distinct meanings than the words in the 5th and 6th categories. There seems to be a negative association between polysemy and technicality. In other words, the highly technical words are less likely to be highly polysemous, and vice versa. Accordingly, the words in the 5th and 6th categories have a specific, stable meaning and use and will not pose a major problem for medical students once the teacher has explained them. As for words in the 1st–4th categories, students should be made aware of the fact that some lexical items have different meanings and are used in a somewhat restricted way in the discipline concerned. Therefore, although the present MWL provides a window to the medical specialized content area, focusing on the frequent lexical items shared across various medical sub-disciplines is still not enough for EFL medical students like ours. It is suggested that data-driven corpus-based teaching materials should be provided so that students themselves can make direct discoveries about how technical and sub-technical words are used in the specialist context. Those lexical items that are common to both general and specialist register may have particular uses that will be revealed in concordances. Classroom exercises using concordance lines may thus be undertaken as follows.
For instance, aspiration is a family member of one of the words in the MWL and involves several meanings. It may refer to (1) a strong desire for high achievement/ an ambition (used in general language and its verb being aspire), (2) a speech sound produced with an aspirate (in phonetics), (3) the act of inhaling/drawing in (in particular referring to foreign materials such as vomitus and mucus into the trachea, lungs or the respiratory tract, and its verb being aspirate), (4) the removal of fluids, gases, or solids from a cavity with a suction device (used exclusively in the medical field). From the following concordance sample derived from the present Medical Textbook Corpus, students may be required to study the meanings of the word aspiration in the context. Activities include matching each of the four meanings above with the correct concordance examples below and indicating its original verb (see Answer Key One below). A series of more difficult exercises concerning the identification of the words that collocate with aspiration may pursue (see Answer Key Two).
Hiatal hernia with esophageal reflux symptoms, which increases the risk of pulmonary
In the 1970s, it was recognized that early suctioning by the obstetrician or pediatrician decreased the incidence of meconium
For women who choose a surgical procedure, there are two options for first-trimester abortion: manual vacuum
It may be omitted in cases where the patient has a depressed level of consciousness, 5-mL syringe for bone marrow
He has no
She has
In the event of a high spinal block that compromises ventilation or airway control, cricoid pressure should be applied and endotracheal intubation performed to prevent
If an ovarian cyst is drained and subsequently reaccumulates, two options are available-surgical excision and transvaginal needle
Answer key One:
Meanings for aspiration:
A strong desire for high achievement/ an ambition: Examples 5 & 6 (verb: aspire).
The act of inhaling/drawing in: Examples 1, 2 & 7 (verb: aspirate)
The removal of fluids or gases from a cavity with a suction device: Examples 3, 4, & 8 (verb: aspirate)
Answer Key Two:
Words that collocate with aspiration: pulmonary aspiration (adjective + noun); aspiration pneumonia; bone marrow aspiration; meconium aspiration; needle aspiration (noun + noun)
By using corpora, students can gain direct access to abundant examples of authentic language samples, resulting in a better understanding of the use and patterns of certain linguistic features. Therefore, corpus-based teaching can help train autonomous students who can take charge of their own learning processes. It may be helpful to provide students with a glossary of the MWL which explains their more usual meanings in the medical context and their most common collocates if possible.
VI Conclusions
Drawing upon the notion of lexical coverage, the 595-word MWL is a frequently-occurring medical word list across medical sub-disciplines, which gave a 10.72% coverage of a variety of English-medium medical textbooks outside of the BNC most frequent 3,000-word families, and distinct from fully-technical medical terms with Greek/Latin word components.
Under the constraint of scarce course time, the MWL appears to display a much more manageable size and still be of potentially great benefit to learners. Thanks to the advancement of computer technology, the MWL can be used with concordancing tools to help students raise their awareness of the usage of these words as well as the frequent sequence of words that accompany them. With more exposure to medical texts in the years that follow, EFL medical students will consolidate the vocabulary knowledge acquired from the MWL.
Last but not least, it must be acknowledged that although the proposed 595-word MWL was grounded in careful screening and can be useful, yet it is still not a panacea so as to ease the burden on much larger medical terminology, which involves various combinations of prefixes, suffixes and word roots of Greek/Latin origins. Rather, it was established to cover the gap between non-technical and fully-technical vocabulary. With high lexical coverage and hence great learning return, the MWL deserves as equal attention as medical terminology from Medical English teachers. They can complement each other and can be harnessed adaptably. The aim of this research has been to generate that awareness.
Footnotes
Appendix 1.
Appendix 2.
Acknowledgements
The researcher wishes to express her profound gratitude to Professor Paul Nation for the free RANGE program, which was indispensable to this study.
Funding
This research was supported by a monographic research grant from the National Science Council (NSC) in Taiwan.
