Abstract
Abstract
This essay studies Vietnamese demotic Nôm characters used to transcribe Sinitic loanwords in the text of Quốc Âm Thi Tập [Poetry Collection in the National Language] by Nguyễn Trãi (1380–1442) using an interdisciplinary approach that combines graphology, historical phonetics, and etymology. The text under study (with 11,067 unique instances of Nôm characters) has 1434 Nôm characters used to transcribe Sinitic loanwords, with 8040 instances of recurrence. These Nôm characters are divided into 10 categories, along two principal groups: (A) Nôm characters that are borrowed from Sinitic; and (B) Nôm characters that are self-generated. Statistics shows that the major trend of Nôm characters used to transcribed Sinitic loanwords is from borrowing (93.3%), while the minor trend is from self-generation (6.7%). Group (A) of Nôm characters mainly uses the method of graphemic borrowing from Sinitic, while group (B) uses two principal methods of graphemic creation: phono-semantic compound and phono-equivalent compound. The results show that basically the system of Nôm characters used to transcribe Sinitic loanwords inherits the graphological tradition of Sinitic as expressed on the following levels: (a) graphemic components from radicals and existing Sinitic characters; (b) methods of graphemic creation including phonetic loan and phono-semantic compound. However, the methods of graphemic creation and the evolution of the graphic form of Nôm characters all follow the common trends of the histories of Nôm characters and the Vietnamese language.
Introduction
This essay studies the principles of Vietnamese demotic Nôm characters (喃字) used to transcribe Sinitic loanwords in the Nôm work Quốc Âm Thi Tập [Poetry Collection in the National Language] by Nguyễn Trãi 阮廌 (1380–1442). It uses an interdisciplinary approach and the fruits of three fields of study: Nôm graphology, historical phonetics, and etymology. Theories of Sinitic and Nôm structural graphology (Nguyễn and Xtankevich, 1976; Gu, 2008; Nguyễn, 2008; Qiu, 2014; Trần, 2012a, 2012b) are used to analyze the functional structure of the system of Nôm characters used to transcribe Sinitic loanwords. Theories of historical phonetics (Maspéro, 1912; Wang, 1948; Gaston, 1967; Nguyễn, 2006; Trần, 2014) are used to determine the phonetic relationships between Sino-Vietnamese (SV) and Non-Sino-Vietnamese (NSV) and Nôm sounds, between the spoken sounds of Nôm characters and their Sinitic etymons and between the spoken sounds of Nôm characters and their phonetic notations (PNs). Finally, etymology (Schuessler, 2007; Trần, 2014, 2016) is used to determine which graphemes are Sinitic loanwords (with their original graphic forms, Sinitic etymograph).
In Vietnamese language texts written in the Nôm script from the 12th to the first half of the 20th century there have been different ways to deal with the phenomenon of Sinitic loanwords transcribed by Nôm characters (Nguyễn, 1990, 2006, 2014). Sinitic characters used to transcribe Sinitic loanwords, in theory, have been a tight and complete structure on all three levels of form–sound–meaning (Nguyễn, 2006: 83–84). Like Sinitic loanwords in Japanese, the Nôm characters used to transcribe these Sinitic loanwords do not need to be added to any other graphemic element. However, while the Sinitic characters used to transcribed SV words (pronounced in SV sounds, and recognized as loanwords by local native speakers) are almost all transcribed exactly by their corresponding Sinitic characters, the Nôm characters used to transcribe Sinitic loanwords read in NSV pronunciations show some changes in the structure and methods of graphemic formation. Based on statistics, this essay will present some observations on the major and minor trends of Nôm graphological structure.
Methodology and subject of study
The full text under study is the work Quốc Âm Thi Tập <国音诗集> [Poetry Collection in the National Language] (in the collection Ức Trai Di Tập <抑斋遗集,卷之七> [Inherited Collection of Ức Trai, Book 7]), with 11,067 unique instances of Nôm characters. The text is preserved in the Institute for Sino-Nôm Studies, call number VNv.143.
The transcription and study of the etymology of Sinitic loanwords in this essay are based on the book Nguyễn Trãi Quốc Âm Từ Điển [Dictionary of Nguyễn Trãi’s Poetry in the National Language] with 2435 entries (Trần, 2014). The Sinitic etymons are presented in the book <同源字典> [Dictionary of {Sinitic} Etymons] (Wang, 1982; Schuessler, 2007; Gu, 2008).
This essay does not include proper nouns for personal names, place names or dynasty names such as 姑射, 汉, 楚, 秦, 颜子, 颜渊. Neither does it include the titles of Confucian and Buddhist texts and classics such as 周易, 诸子, 谷风, 羲易, 羲经.
For etymologies with two interchangeable character forms, this essay will use a single form. For example, the morpheme “cầm” (to hold by hand) has two graphic forms, 扲 and 擒, which are considered equivalent. They are recognized Sinitic graphic forms, and do not belong to Nôm studies.
The morphemes written in popular forms (俗字), or character variants (异体字), are within the scope of Sinitic graphology (Qiu, 2014: 198–200), for these characters are only variants in graphic forms, and do not have value in the structures and methods of Nôm character formation, even if sometimes there are popular forms of characters that only appear in Vietnam (Trần et al., 2016: 81–82). For example,
is the popular form of the characters 夈 < 斎 < 斋, similarly for 迡 < 迟 or
< 飞. However, popular forms used as a component element, whether phonetic or semantic, in self-generated Nôm characters are under consideration here:
{悲+
}~
{拜+
}~
{拜+
} (Trần, 2014: 19).
A single graphic form with two different meanings, and read in two different pronunciations, is still considered two separate characters for study, for example 折 with the pronunciation “chiết” and the root meaning “to break,” and 折 with the pronunciation “chết” and the meaning “to die.” Another example is 主 with the pronunciation “chủ” and the meaning “host,” and 主 with the pronunciation “chúa” and the meaning “king/ruler/lord” (Trần, 2014).
A morpheme written as two different Nôm characters, but by the same method of formation, is still considered a single character. For example, the morpheme “chờ” (from the etymology 侍) is written as two Nôm characters,
and
.
In contrast, a Sinitic loanword morpheme written as two different characters in both form and structure will be considered as two different statistical entries, for example the morpheme “chợ” (from the etymology 市) but written in two Nôm character forms,
and
, and the morpheme “chữ” (from the etymology 字) written as two Nôm characters,
and
.
A morpheme with the same graphic form but unrelated meanings will be divided into two independent entries. For example, 底 (đáy) with the meaning “bottom” is different from 底 (để) with the meaning “to let, abandon, leave alone.”
Graphological differences due to the naming taboo practice will not be considered here, for example tông 宗~
(Ngô, 1997).
Differences in structural form due to the phenomenon of flexible positioning will be treated as uniform, for the two character forms still share component elements and the same functional structural principle. This is a difference between structures of Sinitic and Nôm characters (Nguyễn, 2012: 184, 191). For example,
~
do not differ as Nôm characters, while 吟 and 含 are two different Sinitic characters.
Based on such criteria and methodology, this essay has enumerated 1434 entries with 8040 instances of recurrence. These entries are classified in the following section and analyzed in the third section.
Classification model of Nôm characters used to transcribe Sinitic loanwords
Classification criteria and model
According to phonological criteria, I carry out a binary division of Nôm characters used to transcribe Sinitic loanwords into two main groups.
Sinitic loanwords read in Sino-Vietnamese pronunciations (abbreviated as SV); these words are also called SV words.
Sinitic loanwords read in Non-Sino-Vietnamese pronunciations (abbreviated as NSV); these words are called here NSV words.
SV pronunciation is a widespread pronunciation for Sinitic loanwords in the Vietnamese language; these pronunciations are concurrently used by Vietnamese people to read Sinitic characters in literary Sinitic texts (Nguyễn, 2002). NSV pronunciations are pronunciations only used to read Sinitic loanwords in the Vietnamese language that have been adopted from before or after the Tang period, with two groups that are Pre-SV and Post-SV pronunciations (Trần, 2016: 150). Pre-SV pronunciations are the pronunciations of some Sinitic loan morphemes that preserve the ancient Sinitic pronunciations before the formation of SV pronunciations (which are hardly ever used to read Sinitic characters in literary Sinitic). Post-SV pronunciations are pronunciations formed on the basis of SV pronunciations influenced by phonological evolution in the Vietnamese language in the period after the existence of SV pronunciations from the Tang and Song dynasties (Nguyễn, 2008: 196). However, phonological criteria are under the purview of linguistics, and therefore hidden under graphological classification. Nôm characters with pronunciations based on SV pronunciations are denoted by the symbol SV, and those with pronunciations based on NSV pronunciations are denoted by the symbol NSV in the classification table below (see Figure 2).

Relation between Nôm graphic forms and Sinitic Graphic forms.

Classification model for Nôm characters used to transcribe Sinitic loanwords: SV: denote Sino-Vietnamese pronunciation; NSV: Non-Sino-Vietnamese pronunciation.
According to the methods of character formation, Nôm characters used to transcribe Sinitic loanwords in the text can be divided into two main groups:
(A) Nôm characters that borrow Sinitic graphic forms (or without internal structure); and
(B) Nôm characters that are self-generated (with internal structure).
In relation to the treasury of Sinitic graphic forms (denoted as C), Nôm characters in group (A) will also appear in both Nôm and literary Sinitic texts, and those in group (B) will only appear in Nôm texts, as in Figure 1.
According to the above criteria, I have come up with the following classification model for Nôm characters used to transcribe Sinitic loanwords (Figure 2).
Therefore, the classification model has 10 types of Nôm characters in total: group (A) has four types and group (B) has six types. In the following I will go into detailed analysis of each type of Nôm character. Nôm characters in group (A) are greater in number, so I will only give five examples for each type. Although fewer in number, Nôm characters in group (B) show changes in graphological structures and are presented in greater detail.
Nôm characters that are Sinitic borrowings (A)
Group (A) is divided into two subgroups, (A+) and (A–). Subgroup (A+) includes Nôm characters that borrow Sinitic graphic forms exactly; subgroup (A–) includes Nôm characters that borrow Sinitic graphic forms phonetically. Subgroup (A+) includes two types of Nôm characters, (A1) and (A2); subgroup (A–) also includes two types of Nôm characters, (A3) and (A4). The form–sound–meaning relationship of these subgroups can be described as follows.
Subgroup (A1) includes Nôm characters used to transcribe SV characters that are fully borrowed in all three aspects of form–sound (SV pronunciation)–meaning (Trần, 2012a: 66, 2012b: 87). For example, the character 安 has the SV pronunciation “an,” meaning “peace.” See Table 1 for example.
Table of subgroup (A1).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
Subgroup (A2) includes Nôm characters used to transcribe Sinitic loanwords read in NSV pronunciation (Nguyễn, 2006: 85; Trần, 2012a: 66 (see Table 2). They are also borrowed in all three aspects of form–sound–meaning, but read in NSV pronunciations, in contrast with subgroup (A1) read in SV pronunciations. For example, the character 忧 has the SV pronunciation “ưu,” meaning “sad,” but is read in NSV pronunciation as “âu” in a Nôm text (Trần, 2014: 8).
Table of subgroup (A2).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
Subgroup (A3) includes Nôm characters used to transcribe SV characters by means of a homophonous or near-homophonous Sinitic character. For example, the character 奠 has the SV pronunciation “điện,” meaning “to arrange,” but is written in a Nôm text by the graphic form 殿, using a Sinitic graphic form that is homophonous but semantically different to transcribe the sound (see Table 3).
Table of subgroup (A3).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
Subgroup (A4) includes Nôm characters used to transcribe Sinitic loanwords read in NSV pronunciation. It is similar to subgroup (A2) with NSV pronunciation and semantic borrowing, but its graphic forms use a different Sinitic loanword that is homophonous or near-homophonous. For example, the character 飞 has the SV pronunciation “phi,” meaning “to fly,” which is read in a Nôm text by the NSV pronunciation “bay,” and for its Nôm character uses 拜, a Sinitic character that is homophonous to the NSV pronunciation to transcribe the sound (see Table 4).
Table of subgroup (A4).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
As mentioned above, the commonality of group (A) is that the Nôm characters all borrow the whole graphic form of a Sinitic character to transcribe Sinitic loanwords; therefore, group (A) is called the group of “Sinitic Borrowing” of graphic forms. The characteristic of this group is no change in the structure of those Sinitic characters, that is, no increase or decrease in character strokes, and no addition or reduction of components. Therefore, even if that Sinitic character has its own internal structure (whether phonetic compound or logical combination) on the level of Nôm graphological structure, they still are characters whose graphic forms are borrowed in toto from Sinitic.
Nôm characters that are self-generated (B)
Nôm graphologists used to consider this group as “self-generated Nôm characters” (Nguyễn and Xtankevich, 1976; Nguyễn, 2008; Trần, 2012a, 2012b). The characteristics of Nôm characters in group (B) is that they use different Sinitic elements (radicals, phonetic notations, Sinitic words) to generate new characters with new graphic forms, not coinciding graphically with any Sinitic word. 1 The greatest characteristic of this group of self-generated Nôm characters is that they (a) all use a phonetic method, including phono-semantic compound and pure phonetic; and (b) are all read in NSV pronunciation (and none in SV pronunciation).
Group (B) of self-generated Nôm characters is divided into two subgroups according to the criterion of “the extent of preservation and reinforcement of graphic form of the Sinitic etymon.”
Group (B+) comprises Nôm characters that keep the same graphic form of the Sinitic etymon (abbreviated as NSV) and with reinforced components (including radical and phonetic notation components). Group (B+) includes three subcategories, (B1), (B2), and (B3).
Group (B–) comprises Nôm characters that are created from new components, without preserving the graphic form of the Sinitic etymon. Group (B–) includes three subcategories, (B4), (B5), and (B6).
The common characteristic of characters in group (B) is that they are all self-generated Nôm characters, with a phonetic element. According to principles of character formation, group (B) includes the subcategories (B1), (B2), (B3), (B4), (B5) of phono-semantic characters, and only (B6) for purely phonetic characters.
(B1) = {B + radical 部首}.
2
For example, the morpheme “bầu,” meaning “gourd,” with its etymon 瓢 (SV: biều) is transcribed by the Nôm character
with the structure {瓢 + 艹}. However, the character 瓢 has a dual function as phonetic notation (PN) as well as semantic notation (SN) (denoting both sound and meaning, 兼声兼义).
3
The radical 艹 is a supplementary element. The functional structure of the characters in subgroup (B1) is {SN + PN/SN}. Table 5 is a statistical table of Nôm characters in subgroup (B1).
Table of subgroup (B1).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
(B2) = {B + a Sinitic word as PN}. For example, the morpheme “buồng,” meaning “room,” with its etymon 房 (SV: phòng) is transcribed by the Nôm character
{房 + 蓬}, in which 房 has the dual function as PN as well as SN (denoting both sound and meaning), whereas the Sinitic 蓬 is a supplementary PN. The functional structure of the characters in subgroup (B2) is {SN + PN/SN}, different from those in subgroup (B1) only in that the PN is a Sinitic word. It seems that in the mind of the character creator, the semantic function of 房 is given more attention, for the NSV pronunciation has become too distant from the SV pronunciation, so native speakers would consider “buồng” as a word that does not have a Sinitic origin (Table 6).
Table of subgroup (B2).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
(B3) = {B + a Sinitic synonym or near-synonym as semantic element}. For example, the morpheme “nghèo,” meaning “dangerous,” with its etymon 尧 (SV: nghiêu) is transcribed by the Nôm character
with the structure {尧 + 危}, in which 尧 has the dual function as PN as well as semantic element (denoting both sound and meaning), whereas the Sinitic word 危 (SV: nguy) is a supplementary SN (Table 7).
Table of subgroup (B3).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
(B4) = {radical + PN} (Table 8). For example, the morpheme “bão,” meaning “storm,” with its etymon 暴 (SV: bạo) as in 暴风 (SV: bạo phong)
4
is transcribed by the self-generated Nôm character
with the structure {
+ 包}, in which the radical
serves as SN and 包 serves as PN (Trần, 2014: 17).
Table of subgroup (B4).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
(B5) = {a Sinitic word as SN + a Sinitic word as PN} (Table 9). For example, the morpheme “năm,” meaning “year,” with its etymon 稔 (SV: nẫm, meaning “year”) is transcribed by the Nôm character
with the structure {南 + 年}. In this structure the Sinitic word 南 (SV: nam, meaning “south”) serves as PN, and the Sinitic word 年 (SV: niên, meaning “year”) serves as SN to specify the particular meaning, which is to use the synonym 年 to substitute for 稔 (Trần, 2014: 238).
Table of subgroup (B5).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
(B6) = {a Sinitic graph to transcribe the front syllable element of the initial consonant cluster + a Sinitic character as PN} (Table 10). For example, the morpheme “sống” with its etymon 生 (SV: sinh/sanh, meaning “to live”) is transcribed by the Nôm character
with the structure {古 + 弄}, in which the Sinitic word 古 (SV: cổ) is used to transcribe the front syllabic element / k- / of the initial consonant cluster, and the Sinitic word 弄 (SV: lộng, meaning “to jest”) is used to transcribe the fluid sound / l- / and the rhyme /-oŋ/. The reconstructed sound for this morpheme in the 15th-century Vietnamese language is /*kloŋ5/ (Trần, 2014: 211, 298, 313) (see Table 19).
Table of Subgroup (B6).
SE: denote Sinitic Etymon; SV: Sino-Vietnamese; SN: semantic notation; SubG: subgroup; WF: word frequency.
Analysis of Nôm characters used to transcribe Sinitic loanwords
Comparing Nôm characters in groups (A) and (B)
The amount of Sinitic loanword vocabulary takes up 1434 entries over 2335 morphemes in Quốc Âm Thi Tập, or 61.41%. Nôm characters used to transcribe Sinitic loanwords appear in 8040 instances or 72.64% of the text (11,067 instances of Nôm characters) (see Table 11). The number of Nôm characters in groups (B) and (A2), (A3) and (A4) account for 77 entries recurring in 505 instances. The number of Nôm characters in subgroup (A1) is 1357, recurring in 7535 instances. The above statistics can be presented in detail in Table 12 and Figure 3.
Statistical table of groups (A) and (B).
Note: The above statistics show the relationship between groups (A) and (B).
Ratio table of groups (A) and (B).
Note: Table 12 can be seen in the ratio chart between groups (A) and (B) (Figure 3).

Ratio chart between groups (A) and (B).
This chart shows that graphemic borrowing is the most prominent trend of Nôm characters used to transcribe Sinitic loanwords in the text of Quốc Âm Thi Tập. Graphemic self-generation is the minor trend. This is understandable because graphemic borrowing ensures the systemic stability of the written script. As mentioned above, in theory, each Nôm character used to transcribe a Sinitic loanword is in itself a complete closed structure in terms of form–sound–meaning, so it is not necessary to create too many new graphic forms to cause a “rupture” in the script and language.
Comparing subgroups (A+) and (A–)
Here I compare the two subgroups (A+) and (A–). Subgroup (A+) includes (A1) and (A2). The total entries in subgroup (A+) are 1242 {= (A1) + (A2) = 1146 +96}. The frequency of appearance in subgroup (A+) is 6631 {= (A1) + (A2) = 6154 + 477}. Subgroup (A–) includes (A3) and (A4). The total entries in subgroup (A–) are 116 {= (A3) + (A4) = 5 + 111}. The frequency of appearance in subgroup (A–) is 940 {= (A3) + (A4) = 37 + 867} (see Table 13 and Figure 4).
Ratio table of groups (A+) and (A–).

Ratio chart between groups (A+) and (A–).
The above statistics show that Nôm characters that borrow exactly form–sound–meaning in toto is the main trend, while Nôm characters that are phonetic loans by means of a Sinitic homophone or near-homophone constitute only a minor trend.
Comparing subgroups (A1), (A2) and (A3), (A4)
In theory, as introduced by the heading, all Nôm characters that are Sinitic loanwords are complete in toto in terms of form–sound–meaning, so their graphemic forms do not need to change. These Sinitic loanwords (especially SV characters) must be written exactly as in a Sinitic text. However, just as Nguyễn (2001: 205) has pointed out, it is not always the case in reality. The phenomenon of using a Sinitic homophone or near-homophone to transcribe a Sinitic loanword is a peripheral trend, but it needs to be studied. As we have seen, subgroup (A3) exhibits a way to transcribe differently and distortedly characters read in SV pronunciations as opposed to subgroup (A1), and subgroup (A4) exhibits a way to transcribe differently characters in subgroup (A2). This can be demonstrated in the form–sound–meaning comparison table between the (A) subgroups (see Table 14).
Comparison table between subgroups (A1) and (A2), (A3), (A4).
Note: The (+) sign shows the equivalence between a Nôm character and its Sinitic etymograph. The (–) sign shows the non-equivalence.
SV: Sino-Vietnamese; NSV: Non-Sino-Vietnamese.
The practice of using a Sinitic homophone or near-homophone to transcribe a SV character appears not only in Quốc Âm Thi Tập, but also in many other Nôm texts. As observed by Nguyễn (2001: 205–206), there are 180 cases in the Nôm dictionary by Vũ Văn Kính, 13 cases in Nhị Độ Mai Diễn Ca [Verse-tale about the Twice-Bloomed Plum Blossom] and eight cases in Quan Âm Chú Giải Tân Truyện [New Annotated Story of Bodhisattva Avalokitesvara], at a ratio of under 2%. He explains that this is a flexible trend in the writing of Nôm characters, which still occurs in texts by highly educated Confucian scholars. However, in my opinion, such flexible writing manner is a disadvantageous trend for both script and language. What is the cause of this phenomenon? It is due to the etymological knowledge of native speakers. It is possible that entries in subgroup (A3) are not considered by native speakers as SV characters, but as basic characters in Vietnamese vocabulary. The deep and long entrance of monosyllabic characters into the local language has made native speakers consider them as autochthonous morphemes and forget completely their roots. The morpheme “điện” (to arrange, with the etymon 奠) is transcribed by the character 殿 (điện: palace) because 奠 is a Sinitic character with a rare meaning. Therefore, the word-creator has used a common Sinitic character to transcribe the sound. Among four cases (Table 3) there are three entries with high frequencies of appearance (cả, cái, cha). The delicate point about these Nôm characters is that their Sinitic etymographs have two SV readings. The character 嘏 is a basic vocabulary in Confucian classics, with the common SV pronunciation “hỗ,” meaning “blessing,” 5 but a different pronunciation “cả,” meaning “big,” is recorded in the Shuowen and Kangxi Dictionaries. 6 There are cases of the characters 㸙 and 介 respectively with SV pronunciation “già” and “giới,” and secondary SV pronunciations “cha” 7 and “cái.” 8 The character 介 is read in the pronunciation “cái,” meaning “big,” an antiquarian reading in the Book of Poetry. While the character 㸙 is glossed as a Sinitic character to transcribe the Wu dialect, the character 嘏 is also recorded as a Sinitic dialect.
Therefore, I can now explain the phenomenon of the emergence of characters in subgroup (A3) as follows: (a) a few Sinitic loanwords have two readings; (b) but the second reading (which can arise from the classics or a Sinitic dialect) has not been recognized by native speakers as a loanword; (c) these morphemes have entered Vietnamese vocabulary (through both oral and textual transmissions) and become a basic vocabulary, to the point that native speakers consider them as autochthonous vocabulary and no longer exotic words; (d) on some level, they have forgotten or never known the linguistic and scriptural roots of these morphemes; (e) therefore, they have used Sinitic homophones or near-homophones to transcribe these morphemes.
However, it is worth repeating here that the phenomenon of forgetting etymons and etymographs is only a peripheral trend. This fact will be more clearly demonstrated in the relationship between subgroups (A2) and (A4) and especially group (B).
Characters in subgroups (A2) and (A4) are all Nôm characters used to transcribe Sinitic loanwords read in NSV pronunciations. Statistics show that characters in subgroup (A4) appear more frequently than those in (A2) (see Figure 5).

Ratio chart between subgroups (A1) and (A2), (A3), (A4).
As we know, Sinitic morphemes read in NSV pronunciation are a heavily Vietnamized vocabulary group in both pronunciation and meaning. This phenomenon of Vietnamization has taken place over a fairly long period during the process of SV linguistic contact over 2000 years. These sporadic morphemes have entered the Vietnamese vocabulary fairly profoundly and solidly. Native speakers have considered some of this vocabulary as “pure Vietnamese” morphemes (Trần, 2016: 139), so they have used Sinitic homophones to transcribe them. They still consider some other morphemes as Sinitic loanwords, and have used their exact Sinitic etymographs in Nôm texts. The situation that subgroup (A4) is far more numerous than (A2) shows that Sinitic loanwords read in NSV pronunciation by means of Nôm phonetic loan reflect a much stronger trend than SV characters. This fact can be demonstrated in Table 15:
Statistical table between subgroups (A2) and (A4).
Therefore, with the pronunciation criteria (SV and NSV), I tentatively establish the boundary between the subgroups “complete borrowing” (A+) and “phonetic loan” (A–) and “self-generation” (B). Undoubtedly, the SV characters in subgroup (A1) have brought graphological stability to the system of Nôm characters in Quốc Âm Thi Tập in particular, as well as in other extant Nôm texts. The trends of “phonetic loan” (A–) and “self-generation” (B) are governed by subjective reasons of the graphological subject (whether conscious or unconscious about etymographs). This fact will be discussed in greater detail in the following section.
Discussion of Nôm characters in group (B)
As in “Ratio Table of Groups (A) and (B)” (Table 12), we have seen that there are only 76 self-generated Nôm characters used to transcribe Sinitic loanwords in 649 instances. This is a minor trend compared to Nôm characters that are Sinitic borrowings. However, this trend is dealt with in many different methods of character formation. In theory, Nôm characters in group (B) are considered by the cultural subject as “pure Vietnamese” words, and are therefore dealt with by several different methods, chief among which is the reinforcement of “foreign components,” such as SNs and phonetic notations. On one level, the classification model for Nôm characters of Sinitic origins is a “mini-system” within the diachronic classification table presented by Nguyễn and Xtankevich (1976). Therefore, the types of Nôm characters under investigation are all governed by rules of character formation in Nôm (Nguyễn, 2006: 86).
Nôm characters in group (B) within the Nôm classification model in general. Looking at the Nôm classification model (Figure 2), I find many similarities with the Nôm structural model proposed by Nguyễn and Xtankevich (1976), and the classification model of Nôm characters that borrow NSV pronunciation investigated by Nguyễn (2006) as follows (Table 16).
Classification model for Nôm characters by Nguyễn.
SV: Sino-Vietnamese; NSV: Non-Sino-Vietnamese.
Nguyễn’s general model to express the major trends in the structure of Nôm characters cannot reflect all the subcategories that rarely appear, especially Nôm characters to transcribe Sinitic loanwords. Only recently, the diachronic classification model of Trần (2012b: 84–87) with 32 subcategories has included all the phenomena of the structure of Nôm characters over eight centuries.
The paper by Nguyễn and Nguyễn (1986: 19–23) is the first to investigate the transcription of SV characters in Nôm texts. The authors divide them into three categories: (a) Nôm characters that use a Sinitic homophone to transcribe an SV word; (b) Nôm characters that use a Sinitic near-homophone to transcribe an SV word; (c) Nôm characters that use the exact etymograph but reinforce it with a new element (a radical). Groups (a) and (b) correspond to subgroup (A3) in this paper. Group (c) corresponds to subgroup (B1).
Nguyễn’s paper (2006: 85) is the first one to study the transcription of Nôm characters with respect to Sinitic loanwords read in NSV pronunciation. However, the author mainly focuses on studying the following: (1) “The impact of foreign components to structures of self-generation Nôm characters;” (2) Meanwhile, such Nôm characters still preserve the graphic form of the Sinitic etymons (denoted as NSV) with five principal subcategories: (a) NSV for Nôm characters that are Sinitic borrowings; (b) NSVk (combining NSV with supplemental sign k to adjust the sound); (c) NSVa (combining NSV with the Sinitic phonetic component (a); (d) NSVb (combining NSV with the indirect Sinitic semantic component b, in which b represents long-standing meaning, to express general meaning, namely a Sinitic radical); (e) NSVc (combining NSV with the direct semantic component c, in which c directly represents the concept, to express the exact meaning, which is a Sinitic word). Compared with the classification model (Figure 2) in this paper, subgroup NSV corresponds to subcategory (A2), subgroup NSVa corresponds to the two subcategories (B2), subgroup NSVb corresponds to subcategory (B1), and subgroup NSVc corresponds to subcategory (B3).
Comparing the two classification schemes, we see that my model does not have subgroup NSVk, for the text of Quốc Âm Thi Tập does not show Nôm characters in this subgroup, whereas Nguyễn’s model does not have the subgroups (A1), (A3), (A4), and (B4), (B5), (B6). This is because his paper investigates neither SV characters nor Nôm characters that no longer preserve their Sinitic etymographs. As for subcategory (B6), it does not appear in Nguyễn’s classification table because this is the type of Nôm character to transcribe the particular pronunciations of the Vietnamese language during the XV–XVII centuries, whereas the author’s investigative samples are from Lục Vân Tiên Truyện 蓼云仙传 [The Story of Lục Vân Tiên], a Nôm text from the XIX century. Subcategory (B6) is also group (Đ) in the classification model of Nguyễn. Here, I want to study the methods of character formation in group (B) in relation to the existence or disappearance of NSV.
Analysis of Nôm characters in subgroup (B+). Structurally, Nôm characters in subgroup (B+) are all characters formed by the method of phono-semantic compound, a characteristic method of Sinitic characters. The common characteristic of subgroup (B+) is the graphological structure {PN + SN}; the etymograph NSV participates like a bifunctional component with “both sound and meaning.” There are three issues that need to be discussed:
the functional role of NSV;
the functional role of supplementary components;
the structure of Nôm characters in subcategories (B1), (B2), (B3).
The structural role of component NSV as mentioned above is an element with “both sound and meaning.” This phenomenon has been proven by Sinitic philologists in that a Sinitic component can carry the function of representing both sound and meaning, as in the characters bị 备, châu 洲, chiếu 照, chúc 蠾, di 怡, đạo 导, etc., which are called “phonetic notation cum semantic representation” (Cao and Su, 1999: 18, 99, 624, 680, 700, 704). Based on this theory, Nguyễn (2006: 88) opines:
In the capacity of a derivative script on a Sinitic material basis, Nôm characters have had to consult in no small measure methods of creation from that originary script, it is thus understandable that in Nôm characters there arises forms of script creation similar to Sinitic. Therefore, to answer the question above, I’m inclined toward the second possibility, the NSV component carrying both functions concurrently, to represent both sound and meaning, that is a characteristic I tentatively term NSV duality (Vietnamese: tính song quan; Chinese: 双关性).
In the cases of subcategories (B1), (B2), (B3) in Quốc Âm Thi Tập, they all exhibit this NSV characteristic, as analyzed in Table 17.
Comparison table between subgroups (B1), (B2), and (B3).
NSV: Non-Sino-Vietnamese; PN: phonetic notation; SN: semantic notation.
In terms of character formation methods, the Nôm characters in subcategories (B1), (B2), and (B3) are all phono-semantic compounds (形声字/谐声字). However, the functional structures of each component in character formation are different, particularly the following.
(B1) = {义符/声符 + 部首}. Subgroup with two semantic components, one indicating particular meaning and one indicating long-standing meaning.
(B2) = {义符/声符+ Chinese character as 声符}. Subgroup with two phonetic components.
(B3) = {义符/声符+ Chinese character as 义符}. Subgroup with two equivalent semantic components.
Regarding the function of foreign elements, we see that even though NSV is already a complete structure in all aspects of form–sound–meaning, the supplementary components (+) are still added for three types of functions “to adjust long-standing meaning” (with a radical, as in B1), “to adjust pronunciation” (with a Sinitic PN, as in B2) and “to specify particular meaning” (with a Sinitic synonym, as in B3). However, the problem is not so simple when we compare them to their etymographs and pronunciations. We all know that there exist concurrently SV pronunciations for the NSV etymons, that is, a Sinitic loanword can be read by at least a phonological bimodality, including SV and NSV pronunciations. For example, 瓢 biều~bầu,
phi~bay and 代 đại~đời. This phonological bimodality exists permanently and parallelly in reading and writing life, especially in the practice of explicative reading of Nôm texts. However, native speakers cannot always determine when to read in SV or NSV pronunciation. Therefore, the foreign elements here are a formal sign to contrast with the etymograph and its SV pronunciation, signaling to the reader that the graphic form at hand can be read by an NSV pronunciation. The foreign elements thus hold three concurrent functions:
to differentiate the graphic form of a Nôm character with its Sinitic etymograph;
that graphic form helps the reader eliminate the SV pronunciation to guide toward a NSV pronunciation;
the NSV pronunciation of the NSV character is its “meaning,” as recognized by native speakers.
These three functions of a foreign component coexist to different degrees in the three subgroups (B1), (B2), and (B3). However, the foreign component in subgroup (B3) is a “Sinitic semantic notation,” whose main function is to specify the particular meaning. Therefore, the character 世 holds all at once the three functions: show form–adjust sound–determine meaning. This can be demonstrated in Table 18.
Comparison table between subgroups (B1) and (B2), (B3).
SVN: Sino-Vietnamese Pronunciation; SE Sinitic etymograph.
Therefore, in subgroup (B+) we see the reinforcement by foreign elements has direct influence on the graphic form and structure of the NSV etymon to generate a new Nôm character. This new graphic form has created contrast in form between Nôm characters and their etymographs, creating contrast between NSV and SV pronunciations, thereby guiding the reader toward a pronunciation according to meaning.
Analysis of Nôm characters in subgroup (B–). Subgroup (B–) is a group of self-generated Nôm characters. The graphic form of Nôm characters does not show the linkage to the graphic form of a Sinitic etymograph. Perhaps in the view of the character creator, these Vietnamese morphemes are not considered Sinitic loanwords.
Subgroup (B–) can be divided into two binary categories according to methods of character formation. The first category includes subgroups (B4) and (B5), which are characters generated by the phono-semantic compound principle {PN + SN}. The second category only includes subgroup (B6), which are characters generated by the phonetic combination method {front-consonant element + PN}. As a special subgroup, (B6) will be discussed separately in a later section.
As for the subgroups of phono-semantic compounds (B4) and (B5), their functional structure is {radical + PN}. In terms of form, (B4) is similar to (B1) in nature, and (B5) similar to (B2). However, the contrast in these pairs of subgroups is the presence or absence of NSV. Because NSV appears regularly in all Nôm characters, this group has already been called (B+). The subgroups (B4), (B5), and (B6) all lack NSV and are called (B–). This contrast causes (B4) and (B5) to be fundamentally different from (B1) and (B2). In group (B+) the preservation of NSV in the Nôm characters shows that the character creating subject may have clear recognition of Sinitic etymons and etymographs, whereas in group (B–), on the contrary, it shows that the character creator no longer knew about Sinitic etymographs.
The similarities between the two subgroups can be demonstrated in the following formulae:
(B4) = {radical SN + Sinitic PN} = (B1)
(B5) = {Sinitic SN + Sinitic PN} = (B2)
The differences between the two subgroups can be demonstrated in the following formulae:
(B4) = {radical SN + Sinitic PN} ≠ {radical PN + NSV} = (B1)
(B5) = {Sinitic SN + Sinitic PN} ≠ {Sinitic PN + NSV} = (B2)
Or a more abbreviated formula for all the subgroups:
(B4) ~ (B5) = {SN + PN} ≠ (B1) ~ (B2) = {SN + PN/SN}
This result shows that once the characters in subgroups (B4) and (B5) have been considered pure Vietnamese morphemes, their graphological structure has been processed according to Nôm principles of character formation. This is the phenomenon that Nguyễn (2008) calls “Nôm-ization.”
In any case, with statistics from the five subgroups (B1), (B2), (B3), (B4), and (B5) we see the following two trends of “Nôm-ization” that have occurred in the structure of Nôm characters used to transcribe Sinitic loanwords.
Trend 1: the phenomenon of increasing semantic representation in self-generated Nôm characters (such as the addition of radical SN, and a Sinitic SN to specify the particular meaning).
Trend 2: the phenomenon of increasing phonetic representation in self-generated Nôm characters.
Both of these trends exhibit the combination of a larger trend in the evolution of Nôm characters: the increase of phono-semantic compound Nôm characters, and the increase of self-generated Nôm characters.
Analytical table of subgroups (B6).
Analysis of Nôm characters in subgroup (B6). Under the criterion of the presence or absence of NSV, Nôm characters in subcategory (B6) are placed into group (B–), which represents a fairly unique type of self-generated Nôm character. Their structure is also created from two components of Sinitic borrowing, but both of these components are phonetic in nature. Precisely for that reason, this type of Nôm character is also called a Nôm character by “phonetic combination.” The functional structure of Nôm characters in subgroup (B6) can be demonstrated as follows:
(B6) = {secondary phonetic element + primary phonetic element}
A Nôm character in subgroup (B6) uses two Sinitic words combined in a square shape to transcribe a Vietnamese morpheme with the syllabic structure CCVC. Nguyễn (2008: 200–201) calls this type of character a “primo-secondary phonetic combination.” 9 This type of Nôm character is created from two Sinitic words in which each word represents a part of the pronunciation of the Nôm character with an initial consonant cluster. The first graph is only used to “cut out” the front consonant to transcribe the front syllabic element of the initial consonant cluster. The second graph is used to transcribe both the rear component of the initial consonant cluster and its rhyme. Table 19 shows a more detailed analysis.
According to Nguyễn’s assessment (2008: 248–249), in creating the method of Nôm character formation by a “primo-secondary phonetic combination” structure, Vietnamese people did not learn anything from existing methods of character formation in Sinitic. This is really an important creation in traditional Vietnamese philology, allowing Nôm characters to suit Vietnamese phonological structures at the time. It is possible to arrive at the hypothesis that the type of Nôm character by “primo-secondary phonetic combination” was inspired by the fanqie (反切) method in traditional Chinese phonology. The use of the fanqie method in the formation of Nôm characters can be seen as an interesting phenomenon of “Nôm-ization” that needs further study.
Returning to the issue of NSV sound, we see that the cognitive blurriness of native speakers with NSV pronunciation has generated a system (B) of Nôm characters. We see that group (B+) is derived from subgroup (A2) by the reinforcement of a PN or SN to NSV, and group (B–) is derived from subgroup (A4) by using elements with different graphic forms from NSV to transcribe the sound. Once again, we see here that the Nôm character is a phono-semantic script but that the phonetic elements hold a major position. As observed earlier, Nôm characters with PN account for 99.54%, while those with SN 52.44% (Trần, 2012a: 111). 10
Conclusion
From the above analysis, this essay offers the following preliminary observations.
The system of Nôm characters used to transcribe Sinitic loanwords basically inherits the Sinitic graphological tradition. The inheritance is expressed on the following levels: (a) components using radicals and existing Sinitic words; (b) methods of character formation including phonetic loan and phono-semantic compound.
The phenomena of Nôm characters that are Sinitic borrowings in group (A) hold a major position, in which Nôm characters in subgroup (A1) are the most numerous. This fact shows the systematic and parallel nature of the borrowing of vocabulary (in the Vietnamese language) and script (Nôm characters). It also helps us arrive at the observation that while SV words have taken an important part in the history of formation and development of the Vietnamese language, Nôm characters used to transcribe SV words have also played a critical role in the history of formation and development of the Nôm script.
The overwhelming ratio of Nôm characters in subgroups (A1) and (A2) also shows that the major trend in Nôm characters is the complete borrowing of form–sound–meaning. This trend caused stability in the Nôm script.
The low ratio of Nôm characters in subgroups (A3) and (A4) is a minor trend, but the appearance of these phonetic loans show the fairly clear nature of phonetic transcription of the Nôm script. The character creator sometimes might forget the etymon and could use Sinitic homophones or near-homophones to transcribe a certain morpheme.
The phenomenon of “forgetting the etymon” is more clearly expressed in subgroups (B) and (A4), when the character creating subject considers the morphemes that need transcription as “pure Vietnamese” words (unrelated to Sinitic). Precisely from there, Nôm characters used to transcribe Sinitic loanwords have exhibited aspects of graphemic “Nôm-ization”/transformation through three principal methods: (a) phonetic loan with a Sinitic homophone or near-homophone (as in subgroups (A3) and (A4)); (b) phono-semantic compound (subgroups (B1), (B2), (B3), (B4), and (B5)); and (c) “primo-secondary phonetic combination” (subgroup (B6)). Among those three methods, phono-semantic compound is the dominant one. True to an earlier observation by Nguyễn (2006), Nôm characters used to transcribe Sinitic loanwords form a “subset” of the whole system of the Vietnamese Nôm script. Therefore, all methods of character formation and issues of the evolution in graphic form of the Nôm script have moved along the common trends of the history of the Nôm script and the history of the Vietnamese language (trans. Nguyễn Quốc Vinh, Harvard University).
Footnotes
Declaration of conflicting interests
The author declares that there is no conflict of interest.
The phenomenon of graphical coincidence is possible in reality, but it is beyond the scope of this essay.
The structures of the subgroups (B1), (B2), and (B3) are represented here by the formula {B +}, even though on the graph the NSV (i.e., B) element can stand in front of or behind the foreign element. Because NSV is the issue at hand, as introduced above, the exchange in position of elements in the structure of a Nôm character does not create a new word, but only the different ways of writing the same word, whereas the structure of subgroups (B4), and (B5) is written by the formula {radical + PN}.
Or 形声兼会意字 (Gu, 2008).
Bão (storm) is an abbreviated form for the word bạo phong 暴风 whose Vietnamized pronunciation is bão bùng [Trần, 2014: 17]. 风 *pjuwng (Baxter, 1992: 185).
嘏:【广韵】福也。【诗·大雅】纯嘏尔常矣。【笺】子福曰嘏。【礼·礼运】祝嘏莫敢易其常古,是谓大假。(Zhang, 1958: 205).
嘏:【唐韵】古雅切【集韵】【韵会】【正韵】举下切,音贾。【尔雅·释诂】嘏,大也。【疏】《方言》云:凡物壮大谓之嘏。又【说文】大远也。(Zhang, 1958: 205).
㸙:【广韵】正奢切【集韵】之奢切,音遮。【玉 篇】父也。【广韵】吴人呼父。(Zhang, 1958: 690).
介:又大也。【诗·小雅】神之听之,介尔景福。(Zhang, 1958: 91).
This is group (Đ) in Nguyễn and Xtankevich’s model (1976), and subgroup E2 in Trần’s diachronic classification table of Nôm characters (2012b: 85–86).
Nguyễn (2012: 191), based on a different survey sample, also gives a similar measure: Nôm characters with phonetic notation account for 99%, while those account for SN 51–56%.

bầu
bay
đời