Abstract
This study investigates the contextual disparities in the meanings of the Korean Hanja word “행복 [hɛŋbok]” and the Chinese Hanzi word “幸福 [ɕiŋ51 fu35],” both conceptualized similarly to “happiness” in English. Utilizing computational linguistics methodologies, we analyzed 192,249 instances of “행복” from the Modu Corpus and 324,585 instances of “幸福” from the CCL Corpus, employing collocation, bi-gram, and tri-gram techniques.
Our analysis reveals that these words embody distinct cultural, social, and political dimensions in their respective contexts. Korean usage of “행복” shows a strong association with community well-being, social infrastructure, and educational development. Frequently co-occurring terms include “가족” (family), “교육” (education), and “사회” (society), reflecting deep-rooted Confucian values. Conversely, Chinese usage of “幸福” emphasizes national prosperity, collective harmony, and individual life quality. Common collocates include “国家” (country), “人民” (people), and “生活” (life), influenced by traditional philosophies and modern political ideologies.
The research highlights significant differences in how happiness is conceptualized in relation to government policies. Korean contexts focus on specific quality-of-life improvements, evident in phrases like “행복 주택” (happiness housing) and “행복 도시” (happiness city). Chinese contexts stress the link between national development and individual well-being, as seen in expressions like “人民幸福” (people’s happiness) and “幸福生活” (happy life). These findings underscore the complex interplay between personal fulfillment and social responsibility in East Asian cultures.
Our study demonstrates the effectiveness of computational linguistics in analyzing complex cultural concepts, while also revealing limitations in capturing individual subjective experiences. It contributes to contrastive lexicography between Korean and Chinese, offering insights for cross-cultural communication and policymaking in social well-being initiatives.
Keywords
Introduction
The concept of happiness, while universal in its appeal, finds diverse expressions across different cultures and languages. This study aims to examine the contextual meaning differences between the Korean Hanja word “행복[hɛŋbok]” and the Chinese Hanzi word “幸福[ɕiŋ⁵¹ fu³⁵]” using computational linguistics tools within a quantitative linguistics framework. While these words are often considered homonymous synonyms across the two languages, their usage and connotations may differ significantly, reflecting distinct cultural and linguistic contexts. 1
The relationship between these two words is complex and historically nuanced. The Chinese Hanzi word “幸福” first appeared in the 新唐書 (New Book of Tang), 2 published between 1044 and 1060. In contrast, the Korean Hanja word “행복” seems to have first appeared in 日東記遊 (Travelogue to the East), 3 published in 1877. Interestingly, a recent scholarly discussion posits that the modern meaning of the Chinese Hanzi word “幸福” did not originate in Chinese but was transmitted from Japanese in the 20th century. 4
Our research seeks to address two primary questions: To what extent do the contextual uses of “행복” in Korean and “幸福” in Chinese differ? How do these differences reflect the social, cultural, and linguistic characteristics of each language community?
The importance of this research lies in its potential to enhance our understanding of cross-cultural semantics and contribute to the fields of comparative linguistic and cultural studies. While previous studies have examined the etymology and dictionary definitions of these terms, little research has been conducted on their actual usage in contemporary contexts using large-scale corpus analysis.
Our study employs a corpus-based approach, utilizing extensive datasets from both Korean and Chinese sources. Specifically, we analyze 192,249 instances of “행복” from the Modu Corpus and 324,585 instances of “幸福” from the CCL Corpus. These datasets comprise a wide range of text types, primarily focusing on newspaper articles, ensuring a comprehensive representation of contemporary language use. The Modu Corpus, established by the National Institute of Korean Language, encompasses approximately 100 million sentences from 3,536,491 articles produced between 2009 and 2018. The CCL Corpus, provided by the Center for Chinese Linguistics at Peking University, offers a similarly extensive collection of modern Chinese texts. 5
Methodologically, we will employ Python-based computational linguistics tools, focusing on the following: Collocation analysis: to identify words frequently occurring with “행복” and “幸福.” N-gram analysis: to examine the immediate linguistic context of these terms.
These quantitative methods will allow us to move beyond intuitive perceptions and provide empirical evidence for the contextual differences in the usage of “행복” and “幸福.”
6
The significance of this research extends beyond academic linguistics. By elucidating the nuanced differences in how happiness is conceptualized and expressed in Korean and Chinese, this study can inform cross-cultural communication strategies, translation practices, and even policymaking in areas related to social well-being.
The structure of this paper is as follows. The next section details the data collection and processing methods. This is followed by the results of our analyses and their interpretations. The final section summarizes our findings, discusses their implications, and suggests directions for future research.
Through this study, we aim to contribute to a more nuanced understanding of cultural semantics and demonstrate the value of computational methods in cross-linguistic research.
Data overview and processing methodology
Data collection and preprocessing
Korean corpus
Data source: The Modu Corpus, established by the National Institute of Korean Language, was utilized for collecting Korean materials. This extensive newspaper corpus encompasses approximately 100 million sentences from 3,536,491 articles produced between 2009 and 2018. 7
Preprocessing steps: Extraction: 192,249 sentences containing “행복” were extracted using Python regular expressions. Part-of-speech tagging: The Bareun morphological analyzer was employed for its efficiency and accuracy in Korean language processing.
8
Stopword removal: Grammatical elements such as particles, endings, and certain nouns were removed to focus on semantically meaningful parts.
Chinese corpus
Data source: The CCL Corpus, provided by the Center for Chinese Linguistics at Peking University, was used for Chinese data collection. A total of 324,585 sentences containing “幸福” were extracted. 9
Preprocessing steps: Web scraping: LISTLY was used to collect a larger dataset beyond the initial 20,000 downloadable instances. Text segmentation: The Jieba morphological analyzer was employed for accurate Chinese text segmentation. Stopword removal: A three-step approach was implemented to remove special characters, unnecessary spaces, and common stopwords.
Table 1 summarizes the data collection and preprocessing of the Korean and Chinese corpora.
Summary of data collection and preprocessing.
Data analysis methodology
Our study employed two primary analytical techniques to explore the contextual usage and semantic nuances of “행복” in Korean and “幸福” in Chinese. These methods were chosen for their ability to provide both broad and detailed insights into word usage patterns and semantic relationships.
Word co-occurrence collocation
Word co-occurrence collocation analysis examines the tendency of words to appear together within a specified context. This method allows us to identify words that frequently co-occur with “행복” and “幸福,” providing insights into the semantic fields associated with these concepts in their respective languages.
Process: Data processing: The corpus data was iteratively read, with each sentence treated as a separate unit of analysis. This step ensures that we capture the full context of each instance of “행복” and “幸福.” Word extraction and frequency aggregation: Python’s Counter class from the collections module was utilized for efficient frequency counting. Words co-occurring with “행복” and “幸福” were extracted and their frequencies tallied. This step allows us to identify the most common words associated with our target terms. Data sorting and visualization: The aggregated word frequency data was sorted in descending order of frequency. Results were stored in a pandas.DataFrame for improved readability and ease of analysis. This organization of data facilitates the identification of significant patterns and trends.
N-gram analysis
N-gram analysis examines sequences of n contiguous words in a text. In our study, we focused on two types of n-grams: Bi-grams (2-word sequences): Analyzed to capture immediate word pairings with “행복” and “幸福,” revealing common phrases and close associations. Tri-grams (3-word sequences): Examined to understand broader contextual usage patterns and more complex phrasal structures involving “행복” and “幸福.” Stopword and POS filtering (for Korean data): - Stopwords were removed and part-of-speech (POS) filtering was applied to focus on content words. - This step helps to reduce noise in the data and focus on meaningful word combinations. N-gram extraction: - Bi-grams and tri-grams containing the target words (“행복” or “幸福”) were extracted from the text. - We extracted sequences where the target word appeared in any position. - This allows us to examine the immediate and broader context in which these words appear. Frequency-based analysis: - The frequency of each bi-gram and tri-gram was calculated. - This identified the most common 2-word and 3-word combinations involving our target terms. Documentation and visualization of results: - Results were documented in tables and visualized to facilitate interpretation. - This presentation of data allows for easy identification of patterns and trends in the usage of “행복” and “幸福.”
Process:
By analyzing these two levels of n-grams, we can provide a comprehensive view of how “행복” and “幸福” are used in context, from immediate word associations to broader phrasal patterns.
These methodologies provide a comprehensive understanding of the contextual usage patterns of “행복” and “幸福” in their respective languages, forming a crucial foundation for our comparative analysis. The word co-occurrence analysis offers a broad view of the semantic fields associated with these concepts, while the n-gram analysis provides more specific insights into common phrases and expressions.
By combining these methods, we can capture both the broader semantic associations and the more specific phrasal patterns associated with “행복” and “幸福.” This approach allows us to uncover subtle nuances in how these concepts are used and understood in Korean and Chinese contexts, enabling a richer comparative analysis.
Data analysis and discussion
Data analysis
Our analysis begins with a visual representation of the most frequent co-occurring words with “행복” in Korean and “幸福” in Chinese, as shown in Figure 1.

Word clouds of co-occurring words with “행복” in Korean and “幸福” in Chinese.
These word clouds provide an intuitive grasp of the frequency of co-occurring words but may overlook nuanced details. Therefore, we conducted more in-depth analyses using word co-occurrence collocation, bi-gram, and tri-gram techniques as described above.
Word co-occurrence collocation analysis
Utilizing the word co-occurrence collocation method outlined above, we analyzed the frequency of words appearing in proximity to our target terms. This analysis allows us to identify the semantic fields associated with “행복” and “幸福” in their respective languages.
Korean “행복” analysis
Table 2 reveals that in Korean, “행복” is closely associated with various aspects of life, including social relationships, personal emotions, economic factors, health, and education. The high frequency of education-related words (36,301 occurrences) suggests a strong cultural emphasis on education as a path to happiness in Korean society. This phenomenon, often referred to as “education fever” (교육열), reflects the deep-rooted Confucian values that prioritize learning and self-improvement as means to achieve personal and societal well-being.
Top co-occurring words with “행복” in Korean.
Chinese “幸福” analysis
In Chinese, “幸福” shows strong associations with national and societal aspects, as well as individual life experiences (see Table 3). The high frequency of words related to the nation and society (26,686 occurrences) indicates that happiness in the Chinese context is often viewed through a collective lens. This perspective aligns with traditional Confucian and Taoist philosophies, which emphasize harmony between the individual and society, and has been reinforced by modern Chinese political ideology that stresses collective welfare.
Top co-occurring words with “幸福” in Chinese.
Bi-gram analysis
Following the N-gram analysis method described above, we examined bi-grams (2-word sequences) to capture immediate word pairings with “행복” and “幸福.” This analysis reveals common phrases and close associations in both languages.
Korean “행복” bi-grams
The bi-gram analysis for Korean reveals a strong emphasis on social structures and collective happiness (see Table 4). The high frequency of “행복 주택” (happiness housing) and “행복 도시” (happiness city) suggests that urban planning and housing policies are seen as crucial factors in promoting happiness in Korean society. These concepts are part of broader government initiatives to address social welfare issues, particularly focusing on housing affordability for young adults and newlyweds.
High-frequency bi-grams with “행복” in Korean.
Chinese “幸福” bi-grams
The Chinese bi-gram analysis highlights a strong connection between happiness and national or societal well-being (see Table 5). The high frequency of “人民 幸福” (people’s happiness) and “国家 幸福” (country's happiness) indicates that happiness in the Chinese context is often framed in terms of collective national prosperity. This aligns with the political concept of “以人民为中心” (people-centered approach), which has been a key principle in Chinese governance, especially in recent decades.
High-frequency bi-grams with “幸福” in Chinese.
Tri-gram analysis
Extending our N-gram analysis to tri-grams (3-word sequences) as outlined above, we examined broader contextual usage patterns and more complex phrasal structures involving “행복” and “幸福.”
Korean “행복” tri-grams
The Korean tri-gram analysis further emphasizes the role of political entities and social infrastructure in the discourse of happiness (see Table 6). The appearance of political figures and parties in these tri-grams suggests a strong link between political leadership and the pursuit of collective happiness in Korean society. This reflects South Korea's vibrant democratic culture, where politicians often campaign on promises of improving citizens’ well-being and happiness.
High-frequency tri-grams with “행복” in Korean.
Chinese “幸福” tri-grams
The Chinese tri-gram analysis reinforces the interconnectedness of personal happiness with national prosperity and social harmony (see Table 7). The repetition of “幸福” in the first tri-gram (“幸福 享 幸福”) suggests a cyclical or self-reinforcing nature of happiness in Chinese thought, reminiscent of philosophical concepts where happiness is seen as both a means and an end.
High-frequency tri-grams with “幸福” in Chinese.
Discussion
Our analysis of the contextual usage of “행복” in Korean and “幸福” in Chinese reveals significant insights into how happiness is conceptualized and expressed in these two cultures. This discussion aims to interpret our findings within broader cultural, social, and linguistic contexts, highlighting the implications of our research.
Cultural nuances and societal values
The distinct patterns in the usage of “행복” and “幸福” reflect deep-rooted cultural values and societal norms in Korean and Chinese societies. In Korean contexts, the strong association of “행복” with community well-being and social structures aligns with the collectivist orientation of Korean culture, which emphasizes interpersonal harmony and social cohesion. 10 The frequent co-occurrence of “행복” with terms related to education and personal growth (e.g., “교육,” “성장”) reflects the high value placed on self-improvement and academic achievement in Korean society, a phenomenon often referred to as “education fever.” 11
In contrast, the Chinese usage of “幸福” shows a greater emphasis on national prosperity and collective harmony, reflecting the influence of both traditional Confucian values and contemporary political ideologies. The strong association between “幸福” and terms like “国家” (country) and “人民” (people) suggests a more macro-level conceptualization of happiness, where individual well-being is closely tied to national development and social stability. This aligns with the concept of “xiaokang society” (小康社会), a political ideal that emphasizes moderate prosperity for all. 12
These findings highlight how seemingly equivalent terms for “happiness” can embody distinct cultural connotations, shaped by historical, philosophical, and social factors unique to each society.
Role of government and policies
Our analysis reveals significant differences in how happiness is framed in relation to government and policies in Korean and Chinese contexts. In Korean data, the frequent occurrence of terms like “행복 주택” (happy housing) and “행복 도시” (happy city) suggests a more direct and specific approach to promoting happiness through targeted policy initiatives. This reflects the Korean government's active role in addressing social welfare issues, particularly in areas such as housing and urban development. 13
In Chinese data, the strong association between “幸福” and broader concepts of national development (e.g., “发展,” development; “经济,” economy) indicates a more holistic approach to happiness, where the state's role is seen as creating overall conditions conducive to well-being. This aligns with the Chinese government's emphasis on “people-centered development” (以人民为中心的发展), where national progress is viewed as a key pathway to individual happiness.
These contrasting approaches reflect not only different governance styles but also varying conceptualizations of the relationship between the state and individual well-being in these two societies.
Individual vs collective pursuits of happiness
While both Korean and Chinese cultures are often characterized as collectivist, our analysis reveals nuanced differences in how they balance individual and collective aspects of happiness. In Korean contexts, the emphasis on personal growth and education alongside community well-being suggests a dual focus on individual achievement and social harmony. This duality reflects the complex interplay in Korean society between traditional collectivist values and more recent individualistic influences. 14
Chinese usage of “幸福,” with its strong links to national prosperity and social stability, appears to place a greater emphasis on collective well-being as a precondition for individual happiness. This perspective aligns with traditional Chinese philosophical concepts like “datong” (大同, great unity), where individual fulfillment is seen as inseparable from social harmony.
These findings challenge simplistic categorizations of cultures as purely individualistic or collectivistic, highlighting the need for more nuanced understanding of how different societies conceptualize the relationship between individual and collective happiness.
Implications for cross-cultural understanding and communication
The disparities in how “행복” and “幸福” are contextualized have significant implications for cross-cultural understanding and communication. In international dialogues on well-being and quality of life, awareness of these nuances can help prevent misunderstandings and foster more effective communication. For instance, policy discussions on promoting happiness may need to account for different emphases on individual versus societal measures in Korean and Chinese contexts.
Moreover, these findings have practical implications for fields such as international marketing, diplomacy, and global health initiatives. Strategies that resonate with the concept of happiness in one culture may not be equally effective in another, necessitating culturally tailored approaches.
Methodological reflections and future directions
Our computational approach to analyzing large-scale linguistic data has proven effective in uncovering subtle patterns in the usage of happiness-related terms. However, we acknowledge that this method has limitations. While it provides insights into general trends, it may not capture the full complexity of individual experiences or the historical evolution of these concepts.
Future research could benefit from complementing this quantitative approach with qualitative methods, such as in-depth interviews or historical text analysis. Additionally, extending this comparative framework to other East Asian languages or to different time periods could provide a more comprehensive understanding of how concepts of happiness are shaped by cultural and historical factors in the region.
In conclusion, our analysis of “행복” and “幸福” reveals that these terms, while often treated as direct translations, embody distinct cultural, social, and political dimensions. These findings not only contribute to the field of comparative linguistics but also offer valuable insights for cross-cultural understanding, policymaking, and international cooperation in matters related to social well-being and happiness.
Conclusion
This study conducted an in-depth analysis of the cultural and linguistic contextual differences between the Korean word “행복” and the Chinese word “幸福” using collocation, bi-gram, and tri-gram analysis techniques. Through this investigation, we have elucidated that the concept of happiness extends beyond personal emotions, intricately connecting with social, economic, and political dimensions.
The key findings are as follows: While the Korean concept of “행복” focuses on community-centric approaches and social infrastructure improvement, the Chinese “幸福” is more closely associated with national prosperity and social harmony (see “Word co-occurrence collocation” and “Bi-gram analysis” above). In Korea, education and personal growth are strongly linked to happiness, whereas in China, economic development and political stability are emphasized as primary elements of happiness (see “Cultural nuances and societal values” above). Although government policies play a crucial role in promoting happiness in both countries, Korean policies tend to emphasize specific improvements in quality of life, while Chinese policies stress the correlation between national development and individual happiness (see “Role of government and policies” above).
These findings contribute significantly to our understanding of how happiness is conceptualized and expressed in Korean and Chinese societies. The nuanced differences revealed by our analysis highlight the importance of considering cultural context in discussions of well-being and happiness.
The methodological approach employed in this study, combining computational linguistics techniques with cultural analysis, proved effective in quantitatively analyzing complex concepts. However, it also demonstrated limitations in fully capturing individual subjective experiences of happiness. Future research should consider integrating qualitative research methods, such as interviews or focus groups, with this quantitative analysis to address these limitations and provide a more comprehensive understanding of happiness in these cultures.
The insights gained from this study have important implications for various fields: Cross-cultural communication: Understanding the nuanced differences in happiness concepts can enhance communication and mutual understanding between Korean and Chinese cultures. Policymaking: Policymakers in both countries can benefit from these insights when designing and implementing well-being initiatives, ensuring they align with cultural perceptions of happiness. Linguistic studies: This research provides a model for comparative lexical analysis that can be applied to other culturally significant concepts across different languages. Examining how these concepts of happiness have evolved over time, potentially through analysis of historical texts. Investigating how different demographic groups within each society conceptualize happiness. Extending the analysis to other East Asian cultures to build a more comprehensive regional understanding of happiness concepts.
Future research could expand on this work by:
In conclusion, this study provides an important foundation for developing contrastive lexical resources for Korean and Chinese, enhancing linguistic understanding between the two languages. Moreover, by illuminating the cultural specificities of happiness concepts, this research contributes to more nuanced cross-cultural dialogue and cooperation, particularly in areas related to social policy and well-being initiatives.
Footnotes
Acknowledgments
The development of this paper involved extensive programming efforts and the meticulous translation of the original manuscript from Korean into English. In these endeavors, ChatGPT 4.0 and Claude 3.5 Sonnet proved to be invaluable companions, offering essential support that significantly contributed to the completion of this work. I am profoundly grateful for the assistance provided by ChatGPT 4.0 and Claude 3.5 Sonnet, which were instrumental in bringing this research to fruition.
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Ministry of Education of the Republic of Korea and the National Research Foundation of Korea (grant number NRF-2018S1A6A3A02043693).
