Abstract
Affective computing (AC) has been regarded as a relevant approach to identifying online learners’ mental states and predicting their learning performance. Previous research mainly used one single-source data set, typically learners’ facial expression, to compute learners’ affection. However, a single facial expression may represent different affections in various head poses. This study proposed a dual-source data approach to solve the problem. Facial expression and head pose are two typical data sources that can be captured from online learning videos. The current study collected a dual-source data set of facial expressions and head poses from an online learning class in a middle school. A deep learning neural network using AlexNet with an attention mechanism was developed to verify the syncretic effect on affective computing of the proposed dual-source fusion strategy. The results show that the dual-source fusion approach significantly outperforms the single-source approach based on the AC recognition accuracy between the two approaches (dual-source approach using Attention-AlexNet model 80.96%; single-source approach, facial expression 76.65% and head pose 64.34%). This study contributes to the theoretical construction of the dual-source data fusion approach, and the empirical validation of the effect of the Attention-AlexNet neural network approach on affective computing in online learning contexts.
Introduction
The COVID-19 global pandemic has witnessed online learning becoming an indispensable supplementary form in education. However, the wholly online learning model has impaired peers’ communications and teacher-student interactions, which results in inefficient learning performance and many psychological issues (B. Liu et al., 2021; Wei & Chou, 2020). This dilemma may be because of the lack of timely and effective responses to students’ affection, which may detrimentally impact interactions and learning information processing in online learning environments (Han et al., 2021; Z. Liu et al., 2019). Therefore, optimizing the AC approaches in online learning settings is of great theoretical and practical significance.
The cognitive psychological status of learners’ affection was regarded as a reliable predictor of their academic behaviours (Xu et al., 2014). It was also considered a critical indicator of assessing formative learning outcomes (Faria et al., 2017; Zhang et al., 2020). With the increasing progress in the construction and accumulation of the emotion data set, more and more deep learning and neural network models have been developed for affective computing using emotional data, which in return have facilitated the optimization of algorithms to improve the accuracy of affection computing (Kumar, 2021; Fouladgar et al., 2021). However, computing online learners’ affections have many barriers. Firstly, online learning requires lightweight technologies to meet the massive implementation. Prior research mainly employed heavyweight physiological feedback technology, such as EEG, and MRI, to understand learners’ behaviours and cranial nerve mechanism and recognize their emotions (H. Liu et al., 2021, 2021; Shu et al., 2018). However, heavyweight physiological technologies are typically utilized in small-scale experiments, which are challenging to be employed in a massive online learning context. Meanwhile, collecting physiological data was commonly invasive, distracting learners from a learning status and impair their affective expressions. For example, EEG data collection requires external EEG equipment to be in contact with the human head. Facial expression identification is a typical method in AC with many challenges. Although facial expression identification is a relatively lightweight AI technique for persons’ affection recognition (Barrett et al., 2019), sole facial expression data cannot accurately analyse learners’ affections. Given that varied instructive designs used online would generate distinct interactive behaviors between online learners and tutors, requiring multi-source data such as head movements poses to analyse learners’ affections collaboratively. Online learners are typically in an unsupervised learning context, and their facial expressions are vulnerable to the influence of head poses (Wang et al., 2020). Introducing complementary data sources for a single facial expression would improve the accuracy of AC to realize cross-validation and mutual compensation of data information (Wu et al., 2019).
Although deep learning algorithms, such as Convolutional Neural Network (CNN), have achieved good results in AC (Hassan et al., 2019), developing an optimized AC algorithm in online learning settings has two challenges. Firstly, some typical neural network algorithms are deficient in weights allocation, leading to the impairment of enhancing key features and weaken useless features. Wu et al. (2021) reported that CNN could not obtain useful feature information due to its small number of convolution layers, which challenge the accurate extraction of facial expression features and further undermine the facial expression recognition rate. Additionally, traditional CNN has the disadvantage of translation invariance which lead to the concealing of key information from the data (Hou et al., 2021; Sabour et al., 2017). Therefore, the compelling features of the input should be given importance, and computing resources should be allocated to the features. The second challenge is a lack of an affective computing data set to identify online learners’ affections. Students’ affection and expressions may differ from other working situations, which warrants developing a learner affection data set suitable for online learning. Especially in the post-pandemic era, massive online learning has become a new normal learning mode, and online learners’ affections may be affected by pressures from various dynamics in the learning settings (Lockee, 2021).
Tackling the above challenges, this study collected a dual-source affection data set in online learning settings in a middle school, including students’ facial expression images and head pose images. The facial expression images include emotions such as happy, confused, calm and boredom. The head pose images include happy, confused, calm and boredom, which consists of feature points of head contour, eyes, eyebrows, mouth, and nose. Based on AlexNet (Krizhevsky et al., 2012), this study employed the attention mechanism to construct the Attention-AlexNet model that can determine important regions in facial expressions and head poses image data, and allocate more weights on them. By constructing dual-source affection data and AlexNet with attention mechanism, the current study compares the effects of different affection computing models, obtains an optimal model for assessing students’ learning status in an online learning context, and verifies the effectiveness of affective computing a proposed dual-source fusion strategy.
Related Work
Affective Computing in Online Learning Environments
The COVID-19 pandemic has been significantly reforming learning from physical environments to online settings, enabling the effect of affective computing on online learners. However, online learning, primarily conducted in an unsupervised and monotonous context, may increase the risk of perceived stress and anxiety, even leading to learning exhaustion and career burnout (Mheidly et al., 2020). Some researchers found that the positive effect of students’ affection on the performance in the test, students with positive affection showed better performance (Kaur et al., 2021). Another study has certified students’ learning outcomes were related to their affections by investigating the relations between students’ negative affections, acceptance of online education and learning outcomes during the COVID-19 crisis (Tzafilkou et al., 2021). Because the positive affectivity improves students learning motivation and then generate influence in learning performance (Jiménez et al., 2018), while some negative affection, such as boredom lead to insufficient engaged concentration which further impair learning efficiency (Pardos et al., 2013). The above shreds of evidence suggest that affection is a crucial component of meaningful learning. Meanwhile, it is an essential measurement of learning performance. However, many students still fail to regulate their affections in an unsupervised context especially in the remoting learning settings, which appeal to teachers to to obtain students’ learning affections and make appropriate interventions efficiently.
Learners’ affections can be measured by varied data sources, such as facial expression, pose, and physiological parameters. Facial expression is an essential predictor to evaluate an individual’s affective status. Based on facial identification, Z. Liu et al. (2017) constructed a human-robot interaction system, by which human affections could be recognized and artificial facial expression could be generated. Besides, some researchers have found that head pose are another essential factor driving reliably for learners’ affection. Stephens-Fripp et al. (2017) carried out a research to identify learners’ affections and obtain a reliable recognition rate based on body gait and pose. Furthermore, some researchers employed physiological parameters to recognize students’ learning affections. Based on electrocardiograph signals, Hsu et al. (2017) presented an automatic recognition algorithm for exploring the recognition of students’ current affection. The above studies conclude that most AC studies are confined to single-source data and rarely explore multi-source data fusion.
Multi-Source Data Fusion in Affective Computing
Multi-source data fusion is an emerging subject of further affective computing research but poses many challenges. On the one hand, the data relevant to affection was diverse, including EEG, ECG, facial expression, pose and behavior. The syncretic effect of multi-source data can outperform the limitations posed by processing a single form of data and offer more references for AC, which improves the accuracy and robustness of general computing (Cimtay et al., 2020). On the other hand, multi-source data fusion’s AC ability shows an obvious advantage compared with single-source data. Poria et al. (2015) proposed a multimodal information extraction agent, which used three-modal (text, audio, and video) features to enhance the multimodal information extraction process, inferred and aggregated the affective information and achieved an accuracy 87.95%. Salama et al. (2021) compared three approaches (EEG, face, fusion-based) in AC and concluded that the fusion-based affective computing approach had the best accuracy.
Fusion of facial expression and head pose is an effective method to improve affection recognition accuracy. Head pose affects the accuracy of facial expression recognition, considering the facial expression with varied head poses may present different affections (Valstar et al., 2017). For example, a smile face may denote excited affection or helpless affection, while head pose may help to be an interpretive variable to certify the feature of smile. Additionally, most facial expression recognition methods were conducted through the picture captured by a frontal or near frontal head posture, while its accuracy is usually significantly reduced when using a non-frontal posture for testing (Vieriu et al., 2015). Moreover, head poses carry affective information that is complementary rather than redundant to the affection content in facial expressions (Adams et al., 2015), which presents a contextual cue for the interpretation of the affection by facial expressions (Hess et al., 2007)”. B. Abisado et al. (2020) constructed a multi-modal affection model to identify examinees’ five academic affections based on facial expressions and head pose data, and the accuracy is significantly increased to 92.66%. The head pose and facial expressions was also fused in the context of online learning, and it was certified as a reliable approach for measuring students’ engagement and affective states (Yu et al., 2021).
Current research mainly focused on two fusion approaches of multi-source: feature-level fusion and decision-level fusion (Jiang et al., 2020). Feature-level fusion connects features extracted from each modal to a new feature vector. Firstly, feature-level fusion needs to extract features’ representation by each modal data to obtain the commonness of different data sources. It syncretizes multiple independent data sets into a single feature vector and inputs into the machine or deep learning classifier to make decisions. Decision-level fusion syncretises the different data training results, which use the AC models to train affection data and then syncretize the models’ output. Different from feature-level fusion, decision layer fusion emphasizes the distinctions among additional features and can sharpen the discriminative ability. In the current studies, multi-source data fusion has been carried out to conduct AC. Bandara et al. (2016) combined ECG and EDA data to implement affection classification utilizing the decision-level fusion approach. Using EEG and facial expression data, Huang et al. (2017) recognized four affection states (happiness, neutrality, sadness, fear) by decision-level fusion methods. Syncretizeing electroencephalography (EEG) data and facial expression data, D. Li et al. (2019) employed a decision-level fusion framework for detecting affections continuously. Lately, using a feature-level fusion technique, Bairavel and Krishnamurthy (2020) proposed an audio-video-textual-based multimodal affection analysis approach to reach a satisfactory identification rate. Furthermore, based on speech and image data, Y. Li et al. (2018) compared different data fusion methods in AC and found that the decision-level fusion approach is superior to the feature-level fusion approach.
Nevertheless, feature-level fusion has some deficiencies. Firstly, representing the time synchronization between multi-source data features likely results in redundancy between different data (Wu et al., 2006). Secondly, highly correlated feature-level data likely result in a strong correlation between errors in multi-source data. Thus, decision-level fusion may be a more suitable approach for affective computing based on facial expression and head pose data.
CNN and Attention Mechanism
Compared with traditional machine learning methods, deep learning has become the core driving force in the era of artificial intelligence because of its strong learning ability, satisfactory adaptability, higher efficiency, and accuracy. As a feedforward neural network with depth structure including convolution calculation, CNN is one of the representative deep learning algorithms. The basic architecture of the CNN consists of the input, convolution, pooling, full-connection and output layers. According to the structural classification of the convolution layer and pooling layer, the convolution neural network model is divided into network structures with different performances, such as AlexNet, VggNet, ResNet, and GoogleNet.
The existing research outcomes on attention mechanism have been adopted for computer vision research. Several types of attention mechanisms are commonly used in CNNs, such as the Squeeze-and-Excitation Network and the Convolutional Block Attention Module. In CNN, visual attention was typically classified into spatial and channel attention. Spatial attention functions as a transformer of extracting the key information from the spatial information and commonly ignores the information in the channel, which confines the spatial transformation method to the stage of feature extraction of the original picture. Furthermore, to focus on the image’s task-related area effectively, channel attention enables the neural network to determine automatically which channel is important or unimportant and then assign the appropriate weights (Guo et al., 2022). However, the information was typically pooled into the channel’s attention globally, ignoring the local information in each channel. To compensate for the disadvantage of only one kind of attention, Woo et al. (2018) designed a network structure with an attention mechanism in the two dimensions of feature channel and feature space, called the Convolutional Block Attention Modul. It is a lightweight and generic module that can focus properly on target objects, and accurate attention and noise reduction of irrelevant clutters.
CNN and Attention Mechanism for Affective Computing
The Convolutional Neural Network has shown outstanding performance in affective computing (Aydoğdu, 2021). Many CNN structures are applied to AC, such as AlexNet, ResNet, and VggNet (Gan, 2018). AlexNet has more layers and stronger learning capability among various CNNs, and has advantages in online learners' affective computing (J. Li et al., 2021). On the one hand, containing five convolution layers and three fully connected layers, AlexNet has a simple network structure avoiding redundant network parameters, which improve the convergence rate and generalization ability (Krizhevsky et al., 2012). Khare and Bajaj (2020) evaluated four CNNs in affection recognition, and AlexNet offered quicker training and testing. Based on flexible strucures, AlexNet is more convenient to upgrade the structure and obtain a better effect for different online learning image data sets. On the other hand, AlexNet has good effects on facial expression recognition. Giannopoulos et al. (2018) used AlexNet to judge seven different emotions, which were superior to the accuracy of GoogLeNet given its shallow architecture. Demircan and Örnek (2020) used AlexNet architecture for affective computing from speech data and achieved more discrimination. Kartali et al. (2018) compared three deep-learning approaches (fine-tuned AlexNet CNN, Affdex CNN and FER-CNN) based on convolutional neural networks to achieve the real-time affective computing of four basic affections (happiness, sadness, anger, and fear) from facial images. The results show that fine-tuned AlexNet CNN has better generalization power and performance in real-time applications. CNN has significant strengths in AC, but still has exist some limitations. It barely ignores irrelevant information and focuses on crucial details. Thus, paying more attention to data’s significant features has become a bottleneck problem in neural network algorithms applied in AC.
In recent years, a neural network with an attention mechanism has become an essential principle in deep learning research, which is used to distinguish the critical differences of different features (Lin et al., 2020). The attention mechanism is a combinatorial function to highlight the impact of a critical input on the output by calculating the probability distribution of attention. One of the effective approaches to CNN structure optimization is employing an attention mechanism. Attention mechanism, enlightened from the human vision, focuses on allocating input weight. It is a data processing method widely used in classification tasks, such as natural language processing, image recognition and speech recognition through an attention map, to improve the receptive field of the underlying features and highlight the more favourable features for classification (Wang et al., 2017). The previous study has proposed three ways to use attention in CNN to model a pair of sentences, including ABCNN-1, ABCNN-2, and ABCNN-3. The result shows that attention-based CNNs had a better accuracy and recall rate performance than CNNs without attention mechanisms. ABCNNs are beneficial in tackling the selection and textual entailment tasks, with competitive performance in paraphrasing identification. ABCNN-3 generally surpasses ABCNN-1 and ABCNN-2 because it can consider the attention of finer-grained granularity in each convolution-pooling block (Yin et al., 2016). Furthermore, introducing the attention mechanism into CNN can make the resultant neural network pay more attention to valuable features and generate more reasonable output (J. Li et al., 2020). For example, CNN with an attention mechanism was employed to identify the occlusion regions of the face and shift the attention from the occluded patches to other related but unobstructed ones. The approach focuses on the most discriminative un-occluded areas, significantly improving the recognition accuracy (Y. Li et al., 2018). Thus, CNN with an attention mechanism could contribute to better AC outcomes.
Methodology
This section includes three parts. The first part establishes the architecture of the dual-source affective computing model. The second part introduces the convolutional neural network approach with the attention mechanism. The third part introduces the decision-based fusion approach for affective computing.
The Architecture of Dual-Source Affective Computing Model
As shown in Figure 1, the architecture of the dual-source affective computing model contains three sequential stages: transformation, recognition, and fusion. Architecture of dual-source affective computing model.
The input data are facial expression image and head pose image. In the stage of transformation, the image data of facial expression and head pose was converted into matrix pixel data that could be figured out in the recognition stage. In the recognition stage, AlexNet with an attention mechanism was employed to process facial expression and head pose data. Firstly, the convolutional layers with attention layers were employed to extract features of the facial expressions data and head poses data. Then, the fully connected layers and SoftMax activation function were employed to generate the prediction results. In the fusion stage, the decision-level fusion approach was employed for syncretizing the data accessed from prediction results in the recognition stage. By giving different weights to the results of the recognition model, we obtained the syncretic calculating results of facial expression recognition and head pose recognition.
Convolutional Neural Network with Attention Mechanism
AlexNet
This study employed the AlexNet network as the primary neural network structure for learning AC, which has eight layers of weights, including five convolution layers and three full connections (Figure 2). On the one hand, AlexNet, as a deep convolution neural network model, is superior at extracting features and image classification and has been widely deployed in image recognition (Yuan & Zhang, 2016). On the other hand, the AlexNet has been proven as an efficient model in terms of training time, for example, employing “dropout” method that reduces overfitting in the fully-connected layers (Krizhevsky et al., 2012). Furthermore, AlexNet offers a reliable compromise between network simplicity and high performance and shows a good recognition effect on facial expression datasets that are Cohn-Kanade (CK+) dataset and FER dataset (Stolar et al., 2017; Shaees et al., 2020). Thus, it is an effective and stable way to identify students’ learning affections through AlexNet neural network. AlexNet neural network model.
Attention mechanism in convolutional neural networks
This study employs Convolutional Block Attention Module that applies attention to spatial and channel dimensions to maximize the key features and minimize the chaos of inefficient features, (Woo et al., 2018). The computation process of each attention (Figure 3) includes the following stages: 1. The feature map is compressed into two one-dimensional vectors according to its length and width input into the channel attention module. The weight is outputted through the channel attention module. 2. The attention weight is a dot product with the feature map, inputted through the spatial attention module to obtain the output weight and multiplied with the feature map. 3. The result of point multiplication is taken as the output of the attention module. Computation process of each attention map. AlexNet with attention an mechanism (Attention-AlexNet).

From the perspective of space and channel, this study developed an “Attention-AlexNet” neural network fusing the attention mechanism to the AlexNet convolutional neural network. The spatial attention mechanism enhances the contrast between salient and potentially irrelevant regions of affection image data. The channel-wise weight mechanism emphasizes the informative affection features while suppressing fewer valuable features (B. Li et al., 2021). In the Attention-AlexNet neural network, the attention module localizes within the third layer of the AlexNet network. The data processing of the Attention-AlexNet neural network is shown in Figure 4. Data processing of the Attention-AlexNet.
Firstly, the convolution layer of the Attention-AlexNet is responsible for the convolution operation to extract the feature map. Secondly, the pooling layer is accountable for reducing the size of the feature map, and the attention module is responsible for distributing the weight. Finally, the full connection layer is responsible for making decisions and outputting the prediction results. Procedure is as follows: 1. Adjust the size of the original facial expression or head pose image to (224, 224, 3). 2. Employ a convolution operation, the output feature layer size is 96, and the output size is (55, 55, 96); Employ a maximum pool operation, and the output size is (27, 27, 96). 3. Employ a convolution operation, the output feature layer size is 256, and the output size is (27, 27, 256); Employ an attention module distributing the weight; Employ a maximum pool operation, and the output size is (13, 13, 256). 4. Employ three convolution operations, the output feature layer size is 384, and the output size is (6, 6, 256); Employ a maximum pool operation, and the output size is (28, 28, 256) 5. Input full connection layer, the output size is (1, 1, 4096), two times in total. 6. Input the full connection layer, and the output size is (1, 1, 4). 7. Output predictions for each affection.
Decision-Level Fusion for Affective Computing
Decision-Level Fusion Method
Weighted rule fusion, sum and product rules and highest confidence fusion were principally used in the decision level to determine the output results (Gokberk & Akarun, 2006). Considering that distinct data sources will bring different learning emotion recognition results, the data with higher correlation will significantly impact the results. This study implemented the method of weight rule fusion, which is a fusion method that shows the importance of a part in the whole and distributed different scale coefficients to the dual-source data output results. The output results of the facial expression recognition and head pose recognition are in a probability distribution matrix, which can be represented as follows
The output of the recognition results is given different weights, represented as
Decision Level Fusion Process
The decision-level fusion approach syncretizes learners’ facial expression and head pose data (Figure 5). Firstly, the facial expression and head pose data are inputted into the corresponding affective computing model. Secondly, through model training, the recognition output results of each single-source data of facial expression and head pose are obtained. Finally, the decision-fusion approach syncretizes the facial expression and head pose, and affective computing models’ results are considered the final output result. Decision level fusion process of dual-source data.
Experiment
Dataset
Datasets play an essential role in the application of AC algorithms, but they may suffer from apparent biases caused by different cultures and collection conditions (Li & Deng, 2022). At present, affection data sets about middle school students for online learning are lacking. Thus, we constructed middle school students’ dual-source affection data sets to support deep learning model training and AC research.
Participants
36 middle school students, including 19 males and 17 females, were recruited in this experiment in Hangzhou, the East city of China The middle school is an affiliated school to the University where one of the authors of this study working for. One of the research members in this study was assigned to take off-campus internship there during the pandemic period. The facial expression and head pose data set were collected while students were participating in online learning. All students recruited were aware of the research aim and experimental procedure; all participants signed the consent confirm before participating in this experiment.
Data Acquisition and Screening
The physiological experiments are typically classified into invasive experiments and non-invasive experiments. To avoid learners’ distraction in online learning, web camera was used to collect facial expressions and head poses data as sources rather than invasive approach. Data acquisition shown in Figure 6. was carried out in the classrooms and meeting rooms of the middle school. The selected videos ranging from 6 to 10 minutes are popular science videos of different topics, with text introduction, explanation and some background music. The sample videos selected conform to two criteria to maximize learners’ attention during participating the online courses: the content is in the range of learners’ comprehensive ability, and the presentation is diversified and interesting, which help to inspire students' varied affections in the learning process. The data acquisition process is as follows: The staff guided the subjects into the experimental site. Participants filled in the basic information questionnaire and informed consent form. The participants conducted learning tasks online in a nature and quite learning environment, watched online videos, and filled in emotional questionnaires, and the research assistants recorded them. Assistants imported the recorded video files of the issues watching the teaching video into Elan software to assist the subjects in self-assessment of learning emotion. Assistants collected questionnaires and saved video data and students' self-evaluation data. Data collection in a case.
Screening Steps.
Affection Judgment
Specific affections can frequently be perceived from facial expressions and head poses. The face is an affect display system, while the pose shows the person’s adaptive efforts regarding affect (Ekman & Friesen, 1967). The affection judgment of facial expression in this study was based on Ekman’s FACS (Facial Action Coding System) describing the relationship between inner affection and facial expression in detail. The FACS can express the regularity of facial muscle movement when human beings have the same affection, and is not affected by gender, age, race, education level and other factors. Additionally, learners’ varied affections could be detected by different head posture characteristics. The emotional judgment of head posture features is also based on FACS, focusing on the orientation of the face and the position of the cheeks (Trad et al., 2012).
Data Annotation
Data annotation is a process of processing data with the help of marking tools. The procedure of the data annotation in this study is classified labelling. The corresponding labels are selected from the established labels for labelling. This study has four learning affection labels: happy, confused, calm, and boredom. Data annotation is based on head pose images and facial expression images captured from recorded videos.
The data annotation in this study includes two parts: subject annotation and experimenter annotation. Firstly, the participants marked andpassed Elan6.0 (Sloetjes & Seibert, 2016) to evaluate their expressions and keep the corresponding affection labels. The purpose of selecting subject labeling is to provide more direct and accurate affection judgment criteria for experimental labeling personnel. Secondly, three experimenters annotated the image. The data annotation rules are as follows: 1. The image was determined to be a valid data label and was marked if the labels determined by the three annotators were the same. 2. The image was determined to be the data label determined by two annotators and was marked if the labels determined by two annotators were the same and that determined by the other annotator was different. 3. The image was determined as an invalid label, if the three annotators’ definitions were different.
Dataset Division
The data volume of facial expression and head pose data sets include 3939 images. The two data sets contained training, verification and test sets. The training set accounts for 60% of the total data, the verification set accounts for 20%, and the test set accounts for 20%. The facial expression data set (local) and the head pose data set (local) are shown in Figure 7. Although the head pose data and facial expression data are in the same emotional categories, the differences between head pose data and facial expression data in the modeling process are reflected in the content and form of data presentation: Firstly, the features of the data are different. The facial data focused on the features of facial organs, while the head pose data focus on the learners’ head skeleton. Furthermore, single data source may not be a viable judgment of supporting certain affections. For example, if the head pose is askew when facial expression is smiling, it may present a confused affection rather than a happy affection. The facial expression data set (local) and the head pose data set (local).
Model Training
Number of Datasets.
Firstly, model training needs to import the training data set into the AlexNet and Attention-AlexNet models and compute the model’s optimal weight matrix and a bias value. Secondly, it achieves the accuracy of the two neural network models on the validation data set. Furthermore, the progress of model training is determined by the output value of the loss function. The more visible the loss function’s output value is, the closer the model is to the truth. The model training is completed when the output value of the loss functions is zero.
When completing each training iteration of the model, the model was tested by the verification set. The accuracy of the validation set is the proportion of the correct prediction results to the total number of validation sets. The accuracy of the verification set can be employed as the evaluation criterion of the model, which with the highest accuracy will be the best model. In the end, the test set data are imported into the Attention-AlexNet model, and the prediction accuracy of the four types of affections is obtained. The accuracy of the test set is the proportion of the correct prediction results against the total number of test sets.
Results and Discussion
Results of Different Learning Affective Computing Models by Single-Source Data Set
After model training, the loss functions of facial expression and head pose data in the AlexNet, and Attention-AlexNet models converge and have a stable value, as shown in Figure 8. The accuracy of these models on the training set is more than 95%. The training set prediction accuracy and the loss function’s convergence show that these models have good performance and optimal overall prediction effect. Convergence of training set loss function of AlexNet model and Attention-AlexNet model based on single-source data.
The recognition accuracy results of the AlexNet model and Attention-AlexNet model based on single-source data for different learning affections in the verification set are shown in Figure 9. The facial expression model based on Attention-AlexNet has the best recognition effect on happy (88.70%) and confused (76.77%). The facial expression model based on AlexNet has the best recognition effect on calm (81.26%) and boredom (80.79%). The recognition accuracy of AlexNet model and Attention-AlexNet model for different learning affections based on single-source data set.
The results show that affective computing is effective based on single-source data and deep neural networks. Firstly, affective computing based on deep learning can implement the automated extraction of emotional features, simplify the data processing steps, decrease the access threshold, and enhance the algorithm’s efficiency. Secondly, CNN in deep learning has advantages in image data extraction, prediction, and large-scale image data processing. The CNN model based on attention mechanism can effectively improve the accuracy of learning affection recognition and optimize the affective computing model. Lastly, learning affective computing can also be identified based on the single-source data of facial expression and head pose, and the effect of using facial expression data is the best.
Results of Different Learning Affective Computing Models by Dual-Source Data Sets
The Results of the Fusion Model that are Giving Different Weights.
From Table 3, it is found that when
The dual-source data fusion model showed a better learning AC effect than the single-source data model in this study. On the one hand, learning affective computing based on dual-source data is a practical recognition approach that can effectively increase the accuracy of online learning affective computing. On the other hand, with the increasing application scenarios of AC, the problems to be solved become more in-depth and detailed. The available open-source data sets supporting AC can no longer meet the needs of current education scenarios and normalized applications. Constructing multi-source data sets that satisfy specific educational approaches’ needs will become a meaningful way to solve the challenge of personalized AC.
Optimal Model and Test Set Results
The Accuracy of Different Models for Affection Recognition.
Figure 10 is the different affections’ confusion matrix, which shows the recognition results of learning affections of the optimal model with the test set as the data set. The abscissa of the matrix represents the predicted classification results, and the ordinate of the matrix represents the accurate classification results. The coordinates include happy, calm, boredom and confused. From the confusion matrix diagram’s colour depth and value, the model correctly identifies the number of calm affections up to 218. The number of correctly identifying happy affections is 99. The number of correctly identified boredom affections is 166. The number of correctly identified confused affections is 152. The confusion matrix of different learning affections of the optimal model.
The Recognition Accuracy of the Four Types Affection From the Optimal Model.
From the results of the optimal model, the recognition effect of happy and boredom learning affections was more satisfactory. The possible reason is that the happy expression features are more outstanding, representing curved eyebrows, raised cheeks and exposed teeth. The head pose features an elongated jaw and mouth, which are easy to recognize. Boredom is mostly eyelid closure, lip up or down, and facial features and mostly head tilt and offset, are easy to identify correctly. The recognition accuracy of calm and confusion is relatively low. The possible reason is that their facial expression and head pose features are not prominent.
The above results show that: the deep learning AC model of dual-source data fusion has higher accuracy than single-source data, and the effect of the dual-source data fusion model based on Attention-AlexNet is the best. The AlexNet model with an increased attention mechanism has a better recognition effect than the AlexNet model. The optimal AC model (Attention-AlexNet) has the most satisfactory recognition effect on happy, followed by boredom, calm and confusion affection. Based on these findings, educational researchers can optimize the neural network model from the perspective of the integration of diverse data and attention mechanisms to obtain better algorithm’s efficiency. Furthermore, the developed dual-source AC model may be conducive to providing affective care to students, bringing a decent flow state for them, and further decreasing the dropout rate of online learning (Semerci & Goularas, 2021).
Conclusion and Limitations
This study contributed to developing CNN with the dual-source data set and fusion model. Based on establishing a dual-source emotion data set in online learning settings, the current research constructed the neural network deep learning model combined with attention mechanism, which significantly increased the accuracy rate for online learners’ AC. However, some limitations still existed in this study. Firstly, the amount of data collected in this study is relatively small in deep learning research. In the future, we will expand the range of research objects, such as college students and learners in special education, by which we will establish a relatively holistic online learning affection data set. Secondly, the data fused in this study only includes facial expression and head pose; more source data generated from learners, for example, EEG, voice, and eye-tracking are expected to be captured. In addition, the limited number of sample size was a practical constraint due to COVID-19 pandemic, therefore the data set is from one school inone country which may lead to cultural bias of this research. However, we will expand our data set from different countries and regions to reduce cultural bias and further validate the applicability of this approach in the future.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by the National Natural Science Foundation of China (62177042; 6217021982); Humanity and Social Science Funding of China’ Minister of Education (20YJC880118); The 2021 key lab funding of Beihang University of China (VRLAB2021D10; The 2021 key lab funding of Anhui Jianzhu University of China (IBES2020KF02).
Ethical Approval
Before the study commenced, the authors had applied for ethics approval and get the approved. The consent forms were sent to all participants. The data collection approaches were approved by all participants; participants' facial images are processed with mosaic to avoid infringement of portrait rights.
