Abstract
Animation has a rich history of over a century, and scholarly interest in the critique of animated films has evolved alongside its development. With the advent of the Internet and advancements in computer technology, traditional methods of animated movie appreciation are increasingly challenged by new complexities. This paper addresses these challenges by presenting an advanced intelligent system for evaluating animated films, utilizing cutting-edge deep learning models. Our approach begins with a pre-trained deep learning model to classify image categories and extract scene characteristics, while a primary color extraction algorithm captures key color features. Building on this, we propose an innovative deep convolutional neural network (CNN) augmented with fully connected layers to integrate multimodal features, enhancing the accuracy of film evaluations. Furthermore, we introduce a novel deep learning framework designed to detect pixel regions in images that evoke emotional responses. By employing visualization techniques within the convolutional network, we identify emotionally significant regions and enhance these features during initial training, improving the system’s emotional sensitivity. Experimental results demonstrate that our system not only improves the accuracy and efficiency of animated film evaluations but also surpasses existing methods, offering a more nuanced, emotion-aware approach to animation critique. This work advances the field by providing a robust, automated tool for evaluating animated films, offering deeper insights into their emotional and visual components, and contributing to the ongoing evolution of AI-driven film appreciation systems.
Keywords
Introduction
Animated film criticism, as the name suggests, refers to film criticism written by film scholars in strict accordance with film criticism methods and with a solid theoretical foundation. At the beginning of its existence, film criticism relied on the exploitation of the artistry of film to save its low social status, so a group of enlightened artists took on the role of film critics. Therefore, film criticism in the 1920s and 1930s already possessed seriousness and rigor. 1 In the 1930s and 1950s, the flourishing of film theories made film criticism no longer live under the hedge of other art forms but have its own theoretical roots.
All along, the understanding of animation film appreciation has mostly focused on its excavation of film art attributes, but few people have included it in the film industry system. Especially film criticism, which in fact participates in the whole process of film from production to sales to promotion, has become the key to constructing a good object-based relationship between film and audience. In the production machine, film criticism influences production decisions, enabling creators to better understand audience desires and seek to satisfy them; in the consumption machine, the “cinematic atmosphere” created by film criticism is an intrinsic motivation for audiences to enter theaters and an important guide for choosing consumption targets; and in the promotion machine, film criticism in the promotion machine, movie reviews are responsible for cultivating audiences. 2 The reason why film can become an important part of lifestyle is that it has transcended the meaning of ordinary consumer goods and has become a form of culture to be chewed and savored.
Some studies categorize the appreciation of animated films into three distinct types based on content: critical reviews, which aim to guide audiences in choosing which films to watch; analytical reviews, which focus on examining the specific features and characteristics of the films; and interpretive reviews, which seek to validate particular critical theories and methodologies through the analysis of films.
3
An example of this appreciation process, specifically focused on action scenes in animated films, is illustrated in Figure 1. The process of appreciation of animated movies based on action scenes.
However, film criticism, in general, remains an underexplored area of study. While there is a wealth of research dedicated to analyzing film texts and the creative process behind film production, the field of film criticism itself often lacks the same level of attention. In contrast, various forms of film criticism have flourished across different media platforms, with passionate discussions and debates about current films and prominent filmmakers taking center stage.
The biggest difference between online movie reviews and popular movie reviews lies in their two-way communication. Good movies have cultivated a large number of high-quality movie fans, who use the Internet to share and interact with their movie-viewing experience, which is also the original source of online movie reviews. They have their own judgment and interpretation after watching a movie, and they also have the desire to share their feelings with their relatives and friends. The existence of social networks and movie forums provides a great convenience for such sharing. Many online movie reviews are not limited in length, opinion or style. 4 However, the reason why online reviews are included in the film review system is that they break the monopoly of traditional media and achieve a high degree of democratization in film criticism.
At present, the distribution of animated films is a regular business of the distribution department, and its decision makers are objectively in a special position as gatekeepers of the market, and this threshold of access also plays a role of pre-evaluation, because the selection criteria of the distribution company directly affect whether the film can be seen by the audience. The media’s public evaluation of animation appreciation is often notable for its utilitarian goal of guiding the message to the audience and accepting the praise and criticism made by the journalists. 5 Most of the media evaluations appear in the form of news reports and newsletters, while the entertainment journalists use the style of interleaved narrative and discussion, which is more arbitrary and with distinct personal opinions, often presenting the tendency to praise or kill a certain film.
In addition, the box office of animated films is also an important aspect of film appreciation. Box office is not enough to make a fair and ultimate evaluation of a film’s value, but it can at least verify whether a film is popular with the audience from the market perspective. 6 Box office revenue and attendance can be regarded as the audience’s selective evaluation of a film through consumption, and the weekly box office rankings published by the national cinema computerized ticketing terminal system can provide immediate feedback to the film’s investors, producers and distributors. Audience appeal is a hard indicator, but good appeal is flexible. Different audiences often have different esthetic interests due to different national cultures and regional backgrounds, as well as age, gender, and education. 7 It can also cause very different reactions depending on whether the release schedule is timely and opportune, and there are many inevitable or accidental factors that restrict audiences from entering the cinema.
Based on the above background and combined with deep learning model, this paper proposes an intelligent animation movie appreciation system. Firstly, we propose a method to extract and fuse multimodal features by using deep convolutional neural network superimposed with fully connected layers. Then, for the emotions embedded in animated movies, a deep learning framework that can automatically discover the pixel regions that induce visual emotion expressions in images is proposed. The framework uses an enhancement algorithm to strengthen the features in salience regions, which makes the convolutional neural network more sensitive to the features that induce visual emotion excitation in the learning process, so as to improve the accuracy of the neural network model. Finally, the experimental validation proves the effectiveness and superiority of the proposed method for animated movie appreciation.
The purpose of this research is to explore the role of film criticism in the appreciation of animated movies, focusing on its influence on audience behavior and the film industry. It aims to integrate deep learning models to improve the appreciation process by extracting and fusing multimodal features from animated films. The study seeks to develop a deep learning framework that identifies emotion-inducing visual features in animation, enhancing the model’s accuracy in movie appreciation. Finally, the research validates the proposed method’s effectiveness in advancing the automatic analysis and understanding of animated movie content.
Related works
The development status of animation film appreciation system
As an important part of quality education, art inculcation, especially the inculcation of animation film appreciation is particularly important for contemporary people. In the process of learning film appreciation, teachers and students can appreciate the infinite charm of animated films and improve their esthetic ability, which is very helpful for students’ physical and mental health development and knowledge experience enhancement. 8 For example, the access to animation film knowledge is limited to receiving lectures from teachers in class, and the animation films that can be appreciated and analyzed are limited to the materials provided by teachers.
Outside of animation film classes, there are almost no activities to discuss and appreciate animation films. Even students of professional animation film schools and art colleges can only download animation films on the Internet for appreciation and interpretation, while many niche animation films and animation films made by students themselves are even more difficult to be appreciated and analyzed on a large scale. As an important aspect of quality education, animation film appreciation is being included in the curriculum of humanities education by many universities. 9 As a long-established culture, animation film is as vast as a vast ocean and can profoundly influence the quality of people, and animation film education has a pivotal position in quality education.
Animation film appreciation can not only improve people’s abstract thinking ability, but also improve people’s image thinking ability. However, how to make students deeply feel the beauty of animation films and improve their ability to appreciate animation films in a short period of time has been a problem in animation film education. Therefore, the importance of collecting and organizing the materials for animation film appreciation has been highlighted. The development of a platform that allows students to appreciate and communicate with each other is one of the most effective ways to collect and organize animation film appreciation materials.
It is widely loved by students for its more convenient and richer content, and also provides richer resources and more colorful forms of communication for teachers and students, effectively promoting quality education: (1) Information technology plays an easy and convenient role in helping students to memorize thematic animated movies. (2) Information technology plays a better role in animated movie education and the popularization of animated movie knowledge. 10 (3) Information technology enables students to better understand the connotation of the work. (4) Information technology provides a platform for students’ animation movie creation.
The process of appreciating the animation movie is the process of emotional experience, which is not only the process of experiencing the emotional connotation of the animation movie but also the process of intermingling and resonating between the appreciator’s own feelings and the feelings expressed in the animation movie, in this process, combined with the melody, harmony, rhythm, speed, strength, timbre, style, genre, and other animation movie and the text outside the animation movie. 11 In this process, combined with the understanding of melody, harmony, rhythm, speed, strength, timbre, genre, and other expressions of animated movies as well as textual factors outside of animated movies, we can correctly understand animated movies. Secondly, it closely combines with the life experience and emotional desire of the appreciators and brings the appreciators into the mood of the animation movie quickly, so that the appreciators can appreciate the animation movie closely with their own life experience and get more profound emotional experience.
Current status of multimodal feature extraction and application
In the field of multimodal learning, the key problems are the following two: (1) how to understand the semantics in different modal data, and (2) how to construct semantic correlations between different modalities. Specifically, in the image text retrieval task, for problem (1), the semantic understanding of image modality is mainly analyzed by computer vision, which completes the analysis process from image processing and pattern recognition to image understanding by identifying key points and objects in the image; the semantic understanding of text modality is usually parsed by natural language processing domain knowledge, often based on the statistical natural language composition laws for analysis and understanding. 12 The image text retrieval task, as a problem at the intersection of computer vision and natural language processing, therefore requires not only the fusion of knowledge from both domains but also the communication and interaction between the two modalities to establish cross-modal semantic alignment.
Feature fusion is the integration of features from different data sources into a single representation through optimal combination, and then further processing based on this fused representation according to the relevant task requirements. In the field of multimodal learning, feature fusion has been widely used for tasks such as visual question and answer, multimodal sentiment analysis, image description generation, and image text retrieval. 13 Some of these feature fusion methods perform integration between different modalities through simple linear fusion operations, such as stitching, weighted summation, and point multiplication. These fusion operations do not directly link the internal information of image and text features but automatically adapt the fused multimodal unified representations through a later designed network layer.
To address this problem, a series of related models are approached from different perspectives to learn the low-rank representation of the cross-modal tensor while preserving its ability to express cross-modal interactions as much as possible. The researchers propose multimodal compact bilinear pooling fusion, which takes into account the dimensional catastrophe problem described above, and the method uses a sketch computation algorithm to improve the execution efficiency by sacrificing some accuracy. Here sketch computation achieves dimensionality reduction of visual features and text features, and the model obtains fused features by convolution of visual sketches and text sketches. 14 The subsequent multimodal low-rank bilinear pooling fusion then performs matrix decomposition for the tensor product of image features and text features to obtain multiple low-rank two-dimensional features before multimodal feature combination. This method not only achieves dimensionality reduction but also promotes multimodal feature interaction during the combination process.
Multimodal decomposition bilinear pooling fusion is a further extension of the above model. When performing rank-k decomposition on the fused feature tensor product, this method can flexibly choose the rank value for decomposition and finally obtain the similarity by and pooling. Some researchers have proposed effective dimensionality reduction of fused features using decomposition in a visual quiz task, and the resulting core tensor after decomposition is the basis for interaction between vision and text. Based on this, some researchers also used the tensor product of decomposed cross-modal features in an image text retrieval task and introduced a reordering strategy to further improve the retrieval effect. 15 Recently, some researchers have proposed block decomposition pooling architectures, which aim to decompose diagonal matrix blocks so that the sparse interactions between two modalities, image and text, can be learned using the properties of diagonal matrices while reducing the computational complexity.
In addition, multimodal feature fusion can be applied not only between two modalities, but it can also be widely used in tasks involving multiple inputs or multiple modalities. Some researchers have performed selective fusion based on the nearest neighbors of query input modal features so that visual similarity and linguistic similarity can fit well with each other. 16 Based on this, some researchers use three parallel LSTMs to embed and learn the visual, speech and text modal information contained in short videos, and introduce a filtering fusion operation based on convolutional neural network when fusing the three modal features, that is, the result of the convolution of the three modal features is regarded as a combination of multiple filters, and then the semantic related information in the fused features is perceived by further filtering.
Algorithm design
Multimodal feature extraction for animated movies
This section focuses on the multimodal feature fusion based on CNN. Convolutional neural networks have made remarkable achievements in the field of image and video, and the multimodal depth fusion features proposed in this paper will be extracted by using convolutional design via network. Then the depth features obtained from the 1D convolutional neural network are fused with the category features of the image to become multimodal depth fused features.
The overall flow of the CNN-based multimodal feature fusion model is shown in Figure 2. The CNN is generally used for image and video related tasks, and the convolutional energy can effectively extract various features of the image. As you can see, the shallow convolutional layers can extract simple features such as edges and corners. At the deeper convolutional layers, the CNN is able to extract abstract features, which have been fused with features from different locations of the original image. Overall framework of CNN-based multimodal feature fusion model.
In this paper, three types of features will be extracted from images: primary color, image category, and image scene: (1) Since the pictures in the dataset are in RGB space, we first use the OpenCV library to convert the pictures to HSV. (2) Image category is an important property of images, and there are many open source advanced image classification depths, including the Inception series of deep learning models published by Google.
17
(3) This paper adopts Places2-365-CNN algorithm for picture scene feature extraction.
An example of the process of obtaining multimodal deep fusion feature vectors from a 1D convolutional neural network. Different colors in the input vector indicate data of different modes, and the input is mapped to the corresponding position after the convolutional layer, and the multimodal data starts to be fused in an orderly manner:
In this paper, we use the maximum pooling layer, which is used to reduce the dimensionality and map the most prominent features to a higher level. So we get a dimensionality reduction feature:
The loss function for training the convolutional neural network is summarized as follows:
In the optimization process, a penalty term is usually added to prevent the model from overfitting, and the MSE and L2 penalty terms are used in this paper. The final objective equation is obtained in the following form:
After the convolutional layer, we use a fully connected layer to integrate the global multimodal feature information. The output of the fully connected layer is the prevalence prediction value, and in this paper, we use the real label value to supervise the whole network so that the previous layer of the network output learns the multimodal deep fusion features.
Sentiment analysis model for animated movies
In the sentiment analysis of visuals, one problem that all scholars have not been able to avoid is the sentiment gap, which is defined as the lack of agreement between measurable signals, often called features, and the expected emotional state that the perceived signal brings to the user. In order to bridge the sentiment gap, some scholars, inspired by psychology and art, have devised different manual features aiming at finding the visual emotionally motivating features by empower computers to understand human perception of emotions in images. 18 However, the manual features extract the most intuitive primary features of images, which contain fewer semantic features and always have a one-sided understanding of the visual emotion expression, so the results in experiments have been unsatisfactory until the rise of deep learning broke this deadlock, and later, scholars began to use convolutional neural networks to automatically learn the deep features in images and use them in the emotion state analysis. 19
The general architecture of the network model used to analyze the emotion of animated movies is shown in Figure 3. Based on the initialized framework, the emotion classification of a given image can be summarized as follows: given a test image, we first compute the image saliency map by computing the image of that image. To increase the robustness and portability of the network, we use a gradient evocation class activation graph approach by passing back the gradient as the weight of the feature map and by superimposing it between the elements of different channels into the sentiment saliency map of that image. The gradient weights computed here are linked to the sentiment scores, which are also both important determinants of the sentiment salience regions, which are expressed to some extent as the emotional content that easily draws people’s attention. Then, for each layer of the convolutional features of the backbone network as well as the sentiment salience features, the CNN obtains a C-dimensional prediction and fuses it into the final prediction. Visual sentiment analysis model based on salient regions.
In the past research on image sentiment prediction, the more common CNNs are usually VGGNet and AlexNet, but in the traditional feed-forward model architecture, the state is propagated from layer to layer, each layer reads the state of its previous layer, changes the state and keeps the needed information to the next layer. In order to get a better training effect, we use DenseNet, where the input of each layer of DenseNet comes from the output of all the previous layers, which ensures the maximum information transmission between layers in the network and connects all the layers together to alleviate the gradient disappearance problem. Then the feature map of row l is:
After research, it is known that the emotion of data observed in a picture is often motivated by certain specific emotional regions in the picture, and in the research of neuroscientists and biologists has long proven that the reason why humans are able to recognize data of visual scenes quickly is the human attention mechanism, and the existence of the attention mechanism allows humans to acquire information of interest in a short time and skim the background and foreground information by themselves. The existence of the attention mechanism allows humans to acquire information of interest in a short period of time and to locate key information areas on their own, thus enhancing the understanding of visual scene information.
In other words, the main determinants of human visual emotion from images come from certain emotionally salient regions. The purpose of this paper is to propose a method to automatically discover emotionally salient regions of images and enhance the features of emotionally salient regions through feature fusion, so that the convolutional neural network can be more sensitive to features related to emotional categories during retraining and improve the performance of the model in visual emotion analysis. The performance of the model in image visual sentiment analysis is improved. For the input image, the features of the input image are augmented as follows:
The meaning of using the ReLU activation function is the part of the feature map that takes a positive value, only the region that is relevant to the category is focused on. Although the computation of Grad_CAM is different from that of CAM, it can be derived from each other, so both show the same results, but Grad_CAM does not need to use the global average pooling layer to replace the fully connected layer and retrain it. This makes the use of Grad_CAM less restrictive than CAM.
Experiments
Multimodal feature extraction experiments
This section describes the dataset utilized for the experiments. The dataset chosen for this study is SMHP, which originates from the media headline prediction task in the ACMMM2018 Grand Challenge. The data was collected from Flickr, a popular social photo-sharing platform, and consists of 305,613 tagged entries, each corresponding to an image or video uploaded by a user, along with its associated 15 attributes. Each data entry is tagged with a floating-point value greater than 1, with the highest tag in the dataset being 16.56. A higher tag value indicates greater popularity of the social media post. The dataset contains diverse modal data, including images, text, geographical information, and timestamps. In this study, we will extract features from these rich multimodal data using deep learning algorithms and combine them through feature fusion.
Comparison of evaluation results between models.
The comparison of experimental results presented in Table 1 demonstrates that the regression model utilizing multimodal deep fusion features outperforms all four traditional regression algorithms across all three evaluation metrics (MSE, MAE, and SPR). The multimodal features, though initially scattered and ambiguous due to coarse preprocessing, are significantly enhanced after being processed by the convolutional neural network. This transformation maps the original features into a more structured and prominent feature space, leading to improved model performance. Furthermore, the model’s ability to prevent overfitting, combined with an optimized learning algorithm, results in superior performance compared to the other regression models. Given the promising results, the multimodal deep fusion feature approach based on deep convolutional neural networks is recommended as the preferred model for regression tasks in this context.
Figure 4 shows the MAE loss during the training of the algorithm, which converges after 10,000 iterations. It can be seen that the performance of text modal features is significantly higher than that of the other four modal data. Among them, the geographic information modal data performs the worst. The visual modal features are followed by the statistical numerical features, and the worst performance is for the geographic information modal data. This shows that among the multimodal data, textual modal data still occupies the most influential factor in social media popularity. Comparison of MAE loss in training.
As the largest source of information available to humans from the outside world, vision influences people’s social interactions. Time is also an important factor influencing people’s online activities, as work and work life constrain people’s time online. Also, real life creates fluctuations in the time of online social interaction, making temporal modal characteristics a significant factor affecting the popularity of social media. Geographic information modal characteristics are the worst performers in social media popularity prediction, reflecting the global nature of the Internet.
Experiments on sentiment analysis of animated movies
In the sentiment analysis experiments, we evaluate the proposed multi-level sentiment salience enhancement network model using two publicly available datasets: the Emotion6 dataset and the Abstract dataset. The Emotion6 dataset contains 1,980 images sourced from Flickr, with 330 images labeled for each of six emotions: anger, disgust, fear, joy, sadness, and surprise. These images were classified by 230 participants in a comprehensive psychological study, which categorized the emotions into eight distinct groups, encompassing both negative emotions (e.g., anger, disgust, fear, and sadness) and positive emotions (e.g., amusement, awe, satisfaction, and excitement).
To strengthen the credibility of our findings, we conducted comparison experiments on both the Abstract and Emotion6 datasets, extracting a range of features for evaluation. These features were classified into two primary categories: manually extracted features and deep learning-based features. We employed four evaluation metrics in this study: sum of squared differences (SSD), Kullback-Leibler (KL) divergence, and the baroscopic coefficient (BC), to assess the performance of the proposed model.
Average performance of different models in the dataset.
In addition to the common observations mentioned above, there are some different results among the data sets. Features based on artistic principles and artistic elements perform even better than high-level features on the Abstract dataset. This may be due to the fact that the images in abstract paintings are abstract paintings without identifiable objects, and their emotions are mainly induced by art theory and aesthetics.
The model performance with different reinforcement weights is shown in Figure 5. Among them, the visual sentiment analysis model based on sentiment salience regions designed in this paper is to first obtain the feature pixel points related to sentiment classification by convolutional neural me visualization, and do feature strengthening of the pixels with the original image input information, and then retrain the model from the processed data. In this process, the selection of feature enhancement weights is a very critical issue. In this experiment, the enhancement weights are taken to be between 0 and 1, and it is finally found that the enhancement weight of 0.2 has the best integrated effect. Evaluation indexes with different strengthening weights.
Conclusion
Animated movies have a fascinating temperament, and the charm of movies is largely exuded in the appreciation. Film appreciation is the excavator and pioneer of the artistic value of movies, and with the development of the network era and computer technology, new requirements are put forward for the appreciation of animated movies. In order to build an animation movie appreciation system with better performance, combined with deep learning model, this paper proposes an intelligent animation movie appreciation system. Firstly, CNN is used to extract multimodal deep fusion features for text, time, geography and statistical numerical modal information in animated movies, and the CNN is trained with labels, and the multimodal deep fusion features are obtained by taking the first layer of CNN output. Then, for the emotions contained in animated movies, the proposed method imitates the physiological structural properties of human eyes, extracts the salient regions first, and then augments the features in the salient regions to retrain to obtain a deep model that is more sensitive to the relevance of emotion classification. Finally, it is experimentally verified that the method in this paper can be applied to an animated movie appreciation system with advantages in speed and performance. In this paper, three types of features are extracted for animated movies, and in order to make better use of visual information, more diverse visual feature extraction methods will be investigated to improve the model performance.
While this study uses CNN to extract multimodal deep fusion features, there is potential to integrate more sophisticated visual feature extraction techniques to better capture the nuances of animation. Exploring methods like Generative Adversarial Networks (GANs) or Transformer-based architectures for visual analysis could significantly improve the system’s ability to interpret complex visual patterns and emotional cues in animated movies. Additionally, incorporating spatio-temporal feature extraction could help capture motion dynamics and their emotional relevance.
Although multimodal fusion is applied in the current system, future research could experiment with more advanced fusion strategies, such as attention mechanisms, which allow for more dynamic and context-aware feature integration. This could enhance the system’s ability to prioritize certain modalities (e.g., visual, audio, and textual data) based on their emotional salience and context within a particular animation. Exploring hybrid models that combine CNN with Recurrent Neural Networks (RNNs) or Long Short-Term Memory (LSTM) networks might also improve the system’s temporal understanding of animation sequences.
The current method imitates the physiological structure of the human eye for emotion extraction, which can be further refined by incorporating more detailed models of human visual processing. Exploring neuroscience-inspired models, such as those based on visual attention mechanisms, could enhance the system’s ability to detect emotional cues in a more human-like manner. Additionally, understanding how different age groups or cultural backgrounds perceive emotional content in animated films could lead to more personalized systems that cater to diverse audience segments.
One area that holds great promise is the development of systems that adapt to individual user preferences over time. By leveraging user feedback, interaction data, and adaptive learning algorithms, the system could tailor its recommendations and emotional analysis based on individual viewing history and emotional reactions. Incorporating techniques from reinforcement learning or online learning could allow the model to continuously improve and adapt in real-time. While the study uses the Emotion6 dataset and the Abstract dataset, the system would benefit from a broader, more diverse set of annotated datasets. Animated movies from different cultures, genres, and production styles could provide a wider range of emotional contexts for the model to analyze. Additionally, more granular emotion labels and fine-grained annotation could help improve the specificity of emotion detection, allowing for a more accurate and nuanced emotional profile of animated films.
Future research could explore cross-domain analyses by integrating the appreciation system with other forms of media, such as video games or virtual reality (VR), where emotional dynamics play a crucial role. By adapting the system to work across various forms of digital entertainment, researchers could gain a deeper understanding of how animated film emotion is experienced in different contexts, ultimately leading to a more generalized and robust appreciation system. As the field of automated movie appreciation and sentiment analysis grows, it will be important to consider the ethical implications of such systems. Researchers should ensure that the emotional analysis provided by these models does not perpetuate harmful biases or stereotypes, especially in the context of animated movies, which often reflect cultural values and social norms. Exploring the ethical considerations of AI-driven emotional recognition systems and their impact on the creative industries will be critical in ensuring responsible development and application.
Statements and declarations
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial supports for the research, authorship, and/or publication of this article: This study has received support by Achievements of Culture Research of Special Project on Henan Program about Culture Revitalization (No. 2023XWH129), Humanities and Social Sciences Research Project of Henan Provincial Department of Education (No. 2024-ZZJH-340) <Research on the inheritance of shadow puppet modeling and animation application and dissemination in southern Henan>.
