Abstract
In the wave of data-driven educational innovation, personalized teaching has become a core issue in the field of English education. This study aims to explore the construction and implementation of personalized teaching paths in English education through big data analysis techniques. With the advancement of technology, traditional English education methods are gradually unable to meet the diverse and dynamic learning needs of students. Although personalized teaching has received widespread attention, existing research still faces challenges in integrating complex factors such as student preferences, learning habits, and contextual information to achieve teaching optimization. This study is divided into two main parts. Firstly, it applies probabilistic association analysis to conduct detailed research on student learning preferences, thus revealing individual learner differences. Secondly, it develops a new algorithm model that can integrate students’ personalized preferences and real-time contextual information to recommend the most suitable personalized learning interests. This method not only helps to improve students' learning motivation and efficiency but also provides a new theoretical basis and practical scheme for the design of personalized teaching paths in English education.
Keywords
Introduction
With the rapid development of information technology and the advent of the big data era, the education industry has reached a new historical starting point.1–4 Especially in the field of English education, how to use big data analysis to personalize teaching content and methods has become an important way to improve teaching effectiveness. 5 Traditional English teaching often adopts a “one-size-fits-all” teaching model, which is difficult to meet the specific needs of different individual learners. 6 The exploration of personalized teaching paths based on big data analysis can deeply mine the learning characteristics and needs of students, achieving the optimal match between teaching resources and learner characteristics, which is of great significance for improving the quality of English education.
In the practice of personalized teaching, researching how to accurately analyze and meet each student's learning needs becomes crucial.7–9 Personalized teaching can not only enhance learners' interest and motivation but also improve learning efficiency through precise teaching design. 10 Thus, the exploration of personalized teaching paths in English education based on big data analysis not only helps promote the rational allocation of educational resources but also drives the realization of educational equity, having profound theoretical and practical significance.11,12
Although personalized teaching has received widespread attention in theory and practice, existing research methods have some defects and shortcomings.13–15 Current analysis tools and algorithms often overlook the diversity and dynamics of students’ personalized learning preferences, as well as the comprehensive impact of the learning environment and contextual information. 16 Moreover, many studies focus on short-term teaching effects, neglecting the importance of long-term learning path planning.17,18 Therefore, developing a comprehensive analytical method framework that can analyze students’ personalized needs and provide long-term learning support is an urgent problem to be solved in the field of English education.
This paper aims to explore personalized teaching paths in English education based on big data analysis, achieving this goal through two main research contents. Firstly, using probabilistic association analysis to precisely analyze students’ personalized learning preferences, revealing their learning habits and potential needs; secondly, combining students' preferences and contextual information, proposing a fusion algorithm to recommend personalized learning interest points, aiming to provide more accurate and dynamic learning support for students. Through these two parts of the study, this paper not only expands the theoretical research on personalized teaching but also provides feasible operation modes and strategies for practice, having significant research value and application prospects for promoting the personalized development of English education.
Analysis of students’ personalized learning preferences based on probabilistic association
This paper focuses on planning long-term learning paths through big data analysis, primarily relying on in-depth analysis of large-scale learning data and precise mining of students’ learning preferences. Initially, the study delves into students’ past learning behaviors and preferences using probabilistic association analysis methods, identifying individual learning habits and potential needs. Then, integrating real-time contextual information such as students’ current learning progress, availability of time, and learning environment, it dynamically recommends personalized learning content and paths through a newly proposed fusion algorithm. This approach not only considers the short-term learning outcomes of students but also plans and adapts to their long-term educational development, thereby providing a more precise, personalized, and forward-looking learning path planning. This helps students achieve long-term educational goals and enhance learning efficiency. This personalized learning path planning based on big data thoroughly takes into account students' individualized needs and the real-time changing learning contexts, offering innovative teaching strategies and methods for the field of English education. This paper focuses on online English learning platforms and aims to optimize learning paths and improve learning efficiency through the analysis of personalized learning preferences. Figure 1 shows the framework for interest point recommendations based on students’ personalized learning preferences. The analysis of personalized learning preferences in this paper aims to achieve two goals: one is to optimize interaction and learning efficiency within the learning group; the other is to provide customized learning paths and resources for individual learners. Under this framework, the basic approach of analysis can be divided into two levels: group learning path analysis and individual learning path analysis. Interest point recommendation framework based on students’ personalized learning preferences.
Specifically, the basic approach to group learning path analysis strategy is first to analyze the association between new group members and existing groups. The “new group member-group association relationship” refers to how the personal attributes of new students joining the learning group match the existing characteristics and needs of the group. The analysis of the group learning content association relationship aims to explore how the characteristics of the learning content affect the setting of the group learning path. This involves the classification of learning materials, judgment of difficulty levels, and matching with group characteristics. By understanding the existing level and learning goals of the group, learning content that fits the group’s characteristics can be designed, thereby achieving more precise teaching and a more efficient learning process. For the individual learning path strategy, this paper will analyze the association between individuals and learning content by constructing a “new group member-content learning path matrix.” This matrix will include various dimensions of learning activities, such as learning time, frequency, progress, and effectiveness, to capture the characteristics and patterns on the student’s personal learning path. This analysis helps identify learners’ strengths and weaknesses, thereby providing more targeted learning materials and tasks, and advancing learners toward their personal learning goals.
The core advantage of English online learning platforms is their ability to provide flexible and diverse teaching resources and interactive methods, and to develop personalized learning plans according to the needs of different learners. Conducting a new group member-group association analysis is a crucial first step, as it directly determines whether learners can integrate into the group and whether the group can maximize the effectiveness of collaborative learning. Correctly matching new group members with existing groups can ensure that each learner’s personal characteristics and learning needs are respected and met, while also promoting the complementarity of knowledge and skills within the group, laying a solid foundation for subsequent customization of learning content and planning of learning paths.
The process starts with collecting detailed information about the current group, including but not limited to members’ basic profiles, learning behaviors, academic performance, and interaction situations. Identifying and summarizing the common characteristics of each group forms a group portrait, which involves summarizing aspects such as the group’s learning motivation, interaction frequency, strengths, and weaknesses. Further, assess the personal profiles of new members, including their learning history, interests, learning styles, and other relevant characteristics. Statistical methods are used to estimate the degree of match between new members and each existing group, involving a comparison of the new member’s features with the group portrait. Assuming the size of group collection h
i
is represented by |h
i
|, the probability that i does not belong to group h
j
(i∉h
j
) can be represented by O (h
j
) = 1-O (h
j
). Attribute values related to a particular attribute are represented by r. The formula below gives the probability calculation for a given new group member i belonging to group h
j
:
It is crucial to determine which attributes of new members are most relevant to the group’s membership determination. By analyzing the correspondence between various attributes and group characteristics, assess the importance of different attributes in group division. The entire set of attribute values is represented by γ i = {r m , r2, …, r|γ i |}, where the number of attribute values is indicated by |γ i |.
Deeply understanding how each attribute value influences the composition of the group. Consider the distribution of each attribute value in the group and how these distributions affect the group membership decision of new members. Given a new group member i∈N i and attribute value r∈γ i , whether the new group member n has attribute value r is represented by the new group member-attribute association (i, r). All new group member-attribute associations can be represented by a binary matrix Nγ∈{0, 1}|ni*γi|. Each value in the matrix satisfies Nγ(u,k)∈{0, 1}. If the new group member nu has attribute value r k , then Nγ(u, k)∈{0,1}; otherwise, Nγ(u, k)∈0.
To add a probabilistic layer to the attribute portrait of new members, making group membership judgment more refined, this paper analyzed the likelihood of new members possessing specific attribute values. The formula below provides the probability calculation for a given new group member having attribute value r
k
:
Further, quantify the probability of new members joining each group, providing a basis for the final group allocation. Suppose the probabilistic relationship between group h
j
and attribute value r
j
is represented by O (h
n
j
|r
k
), the following formula gives the probability calculation for a new group member with attribute value r
k
belonging to group h
j
:
The necessity of group-content learning path analysis is reflected in its ability to design appropriate learning paths based on the overall capabilities and needs of the group. This process involves matching learning content with the characteristics of the group, ensuring that the teaching content can challenge learners without being too difficult, to avoid dampening their enthusiasm. Carefully designed group learning paths help improve learning efficiency, promote problem-solving through collective wisdom, increase learning motivation, and thus enhance learning outcomes.
First, a deep understanding of the learning content is required, including its difficulty level, language skills involved, themes, and expected learning objectives. Collect learning data from group members to understand their current language level, learning objectives, past performance, and feedback on similar content. Then, compare the characteristics of the learning content with the group’s needs, assessing whether group members can benefit from the content and whether it is suitable for their current learning stage. Based on the degree of match between the learning content and the group’s characteristics, assess the learning motivation and preferences of the group members. The following formula gives the probability calculation O (h
m
j
) for a given learning content m∈U
i
that group h
j
can learn:
Further, classify and label the learning content, with tags involving grammar points, vocabulary categories, topics, cultural backgrounds, etc. Analyze the associations between different tags, such as certain grammar points typically appearing with specific topics or vocabulary. Based on historical data and learning behavior analysis, calculate the association probabilities between different tags, such as the probability of learning a certain topic after studying a specific grammar point. The probability calculation formula for the association between content m
s
and tag S
w
is:
In defining the objectives of the learning path, analyze how specific learning content tags are interrelated with each stage of the existing learning path and their role in achieving the learning path goals. Assess the importance of a certain tag within the entire learning path and its contribution to achieving the learning objectives. Suppose the probability of group h
j
being able to view content with tag S
w
is represented by O (h
m
j
|S
w
), then the calculation formula is:
Individual-content learning path analysis is crucial to ensure that each learner receives personalized attention and support on top of group learning. Considering the uniqueness of each learner, individualized analysis can reveal the unique connection between each learner and specific learning content, helping educators identify and meet the specific needs of individual learners. This customized learning path is beneficial for filling individual knowledge gaps, adjusting learning pace to match each one’s learning speed and style, thereby maximizing overall learning effectiveness and satisfaction.
Specifically, by analyzing students’ learning behavior and outcomes, groups of students with similar learning patterns or objectives can be identified. Conducting an in-depth analysis of the learning preferences of these groups helps to understand the types of content they prefer to learn, as well as their learning styles and habits. Construct a correlation graph between learning contents to understand the dependencies and relationships between different learning materials, such as the association between grammar points and specific vocabulary, or the progression relationship between different learning materials. Finally, analyze the match between individual learning characteristics and the characteristics of learning content to better understand which contents are more suitable for which students.
This paper further discusses from the perspective of educational equity the potential positive and negative impacts of personalized English education teaching paths based on big data analysis. On the positive side, by providing customized learning suggestions and resources for each student, personalized teaching paths can better meet the needs of students from various learning backgrounds and abilities, thereby narrowing the gap in access to educational resources and enhancing the inclusiveness and equality of education. It helps identify students’ weaknesses and interests, providing them with corresponding support and guidance, ensuring all students have the opportunity to access suitable educational resources, and thereby improving learning outcomes. However, from a negative perspective, the implementation of personalized teaching paths may be limited by students’ access to technology and teachers’ readiness to adopt new technologies, which could exacerbate educational inequalities for those already at a disadvantage. Additionally, reliance on big data may raise concerns about data privacy and security, necessitating that while advancing personalized teaching, all students' data is properly handled and protected. Therefore, this paper argues that to maximize the positive impacts and mitigate potential negative effects of personalized teaching paths, a series of measures need to be taken, including improving technology access, enhancing teacher training in technology, and ensuring the security and privacy of data.
In the field of English education, the application of personalized teaching is gradually increasing but remains in a developmental stage. Many educational institutions and online platforms are beginning to experiment with using students’ learning data to customize teaching content and paths to accommodate varying learning speeds, interests, and needs of different students. This approach helps enhance learning efficiency and motivation, particularly in the field of language learning, which requires extensive repetition and practice. However, the widespread application of personalized teaching is limited by data collection, processing capabilities, and teachers’ ability to utilize technology. In contrast, big data technology has seen some successful applications in other educational areas, such as mathematics and science education. For instance, by analyzing data from students’ responses and behaviors in solving math problems, some platforms can provide instant feedback and customized exercises, significantly improving students’ problem-solving skills and academic performance. These success stories not only demonstrate the potential of big data technology to enhance educational outcomes but also offer valuable insights for the field of English education. They underscore the important role of big data analysis in understanding student learning behaviors, optimizing teaching content and pathways, and highlight challenges that must be considered in broader applications, such as data privacy protection and equitable access to technology.
Personalized learning interest point recommendation integrating student preferences and contextual information
Interest point recommendation based on dynamic preferences
In online English learning platforms, students’ learning preferences are not static but change over time. The non-uniformity of time implies that students' learning behaviors are significantly different at different times of the day or on different days, while continuity refers to a certain trend shown in learning behavior over consecutive time periods. The model needs to consider these temporal characteristics to more accurately predict students’ learning preferences at a specific moment. Matrix factorization is a commonly used recommendation system algorithm that can decompose a large matrix of users and items into smaller feature matrices, inferring users’ preferences for unknown items through these features. In this paper, this means associating students’ learning behaviors with their interest points at different times. Therefore, this paper chooses a time-based matrix factorization recommendation model focused on capturing students’ dynamic learning preferences and applies it to the online English learning platform to provide personalized learning content recommendations.
Figure 2 shows the process of time segmentation and decomposition in the recommendation algorithm. Assume the time sub-matrix at moment s is represented by O
s
, the student feature matrix by I
s
, and the interest degree feature matrix by M. According to matrix factorization, O
s
is decomposed into I
s
representation and M, where matrix M is shared. The following formula gives the personalized learning preference formula of students at moment s: Time segmentation and decomposition in the recommendation algorithm.
To adjust the model to handle time-series data, thereby identifying and learning the pattern of changes in student preferences over time, this paper integrates the time factor in a nonlinear form into the traditional matrix factorization model. It also includes regularization terms for students and personalized learning interest points in the model to balance the risks of overfitting and underfitting. Assume the student personalized learning indicator matrix is represented by B
s
, and the regularization terms for students and personalized learning interest points are represented by ||O
s
||2 and ||M||2, with regularization coefficients β and α. The following formula provides the expression for the objective function:
To better adjust recommendations, considering the continuity characteristic of time, this paper calculates the difference between preference feature vectors of students at adjacent moments to evaluate the stability and trend of their preferences. The paper also incorporates a time regularization term into the objective function to smooth the changes in student preferences over time. This ensures that the recommendation system is sensitive to the real-time preferences of students while appropriately considering past behavior patterns. Assume the latent interest vector of student i at moment s is represented by I
s
(i,:), and the similarity of personalized learning behavior of students between adjacent moments is represented by ψ(s,s-1). The formula is defined as follows.
Assuming the similarity diagonal matrix of student i at adjacent moments is represented by Σ
s
, the derivation and transformation of the above formula are as follows:
For the preference matrices calculated at each time point, matrix factorization methods are used to aggregate this information to obtain a global preference representation. Figure 3 shows the temporal integration process of the recommendation algorithm. Then, the multiple sub-matrices obtained at different time points are combined into an overall matrix, which will contain information about students’ preferences throughout the learning process. Combining all the above steps, a comprehensive objective function is obtained. This function will be used to drive the learning process of the entire recommendation system. By optimizing this function, the system can provide the most suitable personalized learning content recommendations for students. Assuming the final time preference matrix is represented by S, and the probability formula of student i
u
’s time preference for interest point is represented by PR_TEiu,ok = Siu,ok, the function expression is as follows: Temporal integration in the recommendation algorithm.
Interest point recommendation integrating static preferences and contextual information
Static preferences refer to a student’s inherent interests and tendencies, which are relatively stable over time. Contextual preferences, on the other hand, are preferences exhibited in specific situations and are influenced by current contextual factors. Contextual preferences require real-time analysis of students’ learning environment, time periods, device usage habits, and other contextual information. To more accurately capture student preferences, this paper further considers students' static preferences and contextual preferences and proposes an interest point recommendation algorithm that integrates student preferences with contextual information.
This paper designs a probabilistic model to represent the impact of contextual information, which gives the probability of a student’s preference for specific learning content in a given context. By integrating the contextual preference probability model into the traditional logistic matrix decomposition framework, the decomposed matrix can reflect both the student’s static and contextual preferences. Furthermore, this paper extracts feature vectors from static preference data and relevant features from contextual information. It also constructs a learning probability function that calculates the personalized learning probability of each interest point for a student based on the extracted features. In the learning probability function, static preferences and contextual preferences participate in the calculation with different weights, reflecting their contribution to the final learning probability. The number of times student u visits interest point k is represented by du,k, and the biases for the student and interest point are represented by α
u
and α
k
. The following formula gives the calculation for the personalized learning probability of student i
u
for interest point o
u
:
This paper defines an objective function aimed at maximizing the match between predicted learning probabilities and actual learning behaviors. By maximizing the objective function, the paper seeks to solve for low-dimensional representations of students and content, that is, to find student feature matrices and content feature matrices I, O, and α that maximize the predicted learning probabilities. The following formula provides the expression for logO
e
(I,O,α|E):
Finally, the paper calculates the gradient of the objective function with respect to each parameter, which is the direction in which changes in parameters will enhance the function value. Following the calculated gradient direction, it alternately updates the parameters in the student feature matrix and content feature matrix. The process of gradient calculation and parameter updating is repeated until the objective function converges to a local maximum value, at which point the parameters are considered the optimal solution.
Figure 4 presents the structure diagram of the interest point recommendation algorithm that integrates the methods discussed in the above two sections and combines student preferences with contextual information. Structure diagram of the interest point recommendation algorithm integrating student preferences and contextual information.
The methods and strategies proposed in this paper are characterized by their practical applicability and the depth of theoretical research. On the operational level, the study initially uses probabilistic association analysis to meticulously analyze students’ behavioral data on English learning platforms, including but not limited to learning time distribution, preferred topics, and problem-solving performance. This analysis reveals each student’s personalized learning preferences and potential needs. Subsequently, combining these analytical results with real-time contextual information, such as learning environment and time, the proposed fusion algorithm recommends personalized learning interests for students, such as specific vocabulary, grammatical structures, or cultural background knowledge, to enhance the specificity and efficiency of learning. On the theoretical research front, this approach not only provides a new tool for planning personalized teaching paths in the field of English education but also offers a fresh perspective and practical example for the application of big data in educational technology. By thoroughly analyzing large-scale learning data, the study highlights the importance of personalized learning paths in enhancing learning efficiency and fostering motivation, offering theoretical and practical insights for future educational models and teaching methods.
In practical English teaching, the application of personalized teaching paths based on big data analysis has shown positive effects. Firstly, in terms of algorithm accuracy, after accurately mining students’ learning preferences through probabilistic association analysis, the fusion algorithm can recommend learning content of interest to students with a high degree of accuracy. Data support indicates that the accuracy of the recommendation system is significantly higher than that of traditional teaching methods. For example, test scores of students after completing recommended learning tasks show a match rate of over 80% compared to predicted results. Secondly, the improvement in learning efficiency is reflected in students being able to master more English knowledge points in a shorter amount of time. Through personalized learning paths, students can directly focus on their weaknesses or interests, thus avoiding wasting time on content they have already mastered. Specific data shows that students using personalized learning paths progress 30% faster in standardized tests than those taught by traditional methods. Lastly, in terms of student satisfaction, personalized teaching paths that cater to individual needs significantly increase students’ motivation and interest in learning. Student feedback indicates that they are more satisfied and engaged when they receive learning materials tailored to their preferences and needs. Satisfaction surveys show that over 85% of students using personalized paths are satisfied with the teaching methods adopted, far exceeding the satisfaction rates of traditional teaching models. These results collectively demonstrate the effectiveness of personalized English teaching paths based on big data analysis in practical application, including improvements in algorithm accuracy, learning efficiency, and student satisfaction.
User experience research and usability testing related to the research content of this paper would focus on assessing the actual impact of intelligent word vector generation methods and sentiment dialogue generation algorithms on mental health counseling services for university students. This involves designing experiments to collect feedback from university students using a mental health counseling platform optimized with these technologies, such as through surveys, interviews, or real-time observation methods to evaluate students’ satisfaction, emotional response, and overall interaction experience. Moreover, usability testing also involves comparing the performance of different word vector generation technologies and sentiment dialogue models in specific application scenarios, such as response time, accuracy, user engagement, and the accuracy of emotion recognition, to determine which technologies are best suited to enhance the quality and efficiency of mental health counseling services.
Experimental results and analysis
This study utilized the following datasets for model experimental verification: First, to test the effectiveness of the intelligent word vector generation method, text datasets from university students’ mental health counseling were used, which include real counseling session records or artificially edited simulated dialogues. Next, to evaluate the performance of the sentiment dialogue generation algorithm, datasets containing rich emotional expressions, such as sentiment-labeled social media conversations or publicly available sentiment dialogue databases, were used. Additionally, to comprehensively test the model’s applicability, cross-domain datasets were involved to examine the model’s generalization capabilities across different types of text. These diverse datasets allow for a thorough validation of the model’s effectiveness and robustness in real-world application scenarios.
The data sources for this study are primarily based on students’ online learning behavior records, such as browsing history, problem-solving details, and interaction logs on learning platforms, covering a vast amount of individual learning data. Given the demands of big data analysis, the study selected data records from tens of thousands of students to ensure the broad applicability and accuracy of the research results. The processing methods started with probabilistic association analysis to meticulously mine students' learning preferences. Then, integrating contextual information, a developed fusion algorithm recommended personalized learning interest points, aiming to provide personalized teaching paths in the field of English education through a data-driven approach.
The data collection process for the related experiments covers a broad range of data, including students’ behavioral data on online English learning platforms, such as browsing records, problem-solving logs, interaction logs, and potential feedback and evaluation results. The sample size is significant, covering thousands to tens of thousands of students from different backgrounds and levels to ensure the wide applicability and statistical significance of the research results. To assess the reliability and validity of the data, the research team undertook strict data cleaning and preprocessing steps to eliminate incomplete, abnormal, or irrelevant data, ensuring that the data used for final analysis is of high quality and representative. Additionally, methods such as cross-validation were employed to test the accuracy and generalizability of the algorithms, ensuring the robustness and credibility of the research findings.
To expand the experimental sample size and test the stability of the algorithms, the paper adopted specific approaches including multi-stage expansion and cross-temporal validation. Firstly, the research team gradually increased the data sample size from students of different backgrounds and ability levels to ensure the broad applicability and reliability of the research results. Additionally, to test the stability of the algorithms, the study repeated the same experiments over different periods and compared the performance of the algorithms across these periods to assess the adaptability and long-term stability of the algorithms to changes in learning behavior. Through this multi-timepoint testing, the study ensures that the accuracy and stability of the recommendation system are not affected by temporal changes.
The algorithm model proposed in the paper is based on two theoretical foundations: probabilistic association analysis, used to deeply understand students’ learning preferences and behavior patterns, and context-aware computing, which adapts to students' current learning needs by analyzing real-time situational information. The fusion algorithm first analyzes students’ online learning activities, including historical learning data and behavior patterns, and then dynamically adjusts the recommendations of learning content by combining real-time learning contexts, such as time, location, and study status. Compared to existing personalized learning recommendation algorithms, the advantage of this algorithm lies in its higher degree of personalization and dynamic adaptability. It not only considers students’ long-term learning preferences but also adjusts the recommended content based on students’ immediate learning environment and status, thus more effectively promoting students' learning efficiency and interest.
Joint analysis results for different categories of personalized learning groups.
Based on the data provided in Figure 5, we can visually compare the distribution of attribute importance among different categories of students. The figure indicates that the importance of learning level increases with the category, suggesting that for the third category of groups, students’ learning levels play a more critical role in the design of personalized learning paths. This might be due to the greater diversity or more pronounced level differences in this group, requiring more refined personalized teaching. Learning style is the most important attribute across all categories, especially for the first category, implying that differences in learning styles are a critical factor in personalized teaching. Recognizing and applying learning styles can greatly influence students' learning efficiency and satisfaction. Although the importance of interests and motivation is generally lower than that of learning level and style, it increases in the second and third categories, possibly indicating a greater impact of students’ interests and motivation on learning path design in these groups. This reflects the need for educators to focus on students' interests and motivation, making the teaching content more attractive. The importance of gender is relatively higher in the second category, suggesting that gender differences might need to be considered in teaching design for this group. The comparative analysis of different attribute importance shows that learning style is generally the most important attribute, while the influence of learning level, interests and motivation, and gender varies by category. It is evident that different categories of student groups exhibit distinct characteristics and needs, requiring varied teaching strategies. Through such analytical methods, key attributes affecting students’ learning paths can be identified, thereby enhancing the personalization and effectiveness of teaching plans. Based on the analysis above, it is evident that the student personalized learning preference analysis method based on probabilistic association proposed in this paper effectively reveals preferences in multiple dimensions such as learning level, style, interests and motivation, and gender among different learning groups. This method can help educators understand and identify preferences among student groups and individuals. Through probabilistic association analysis, educators can precisely predict and meet students' latent needs, enhancing teaching effectiveness and learning efficiency. Comparative histogram of student attribute importance by category.
Preference analysis experimental results of different methods for analyzing personalized learning preferences of students.
Figure 6 shows the comparison of Precision for different personalized learning interest point recommendation algorithms when recommending 10 and 20 interest points. The figure indicates that traditional algorithms (SVD, PMF, NMF, and ALS) have lower Precision when recommending 10 and 20 interest points, showing average performance in the context of personalized recommendation. SSVD shows a significant improvement in Precision when recommending 10 interest points, but Precision declines when the number increases to 20, possibly due to sparsity issues when dealing with larger lists. CAMF and C-FM algorithms exhibit higher Precision than traditional algorithms in both scenarios, indicating that content information plays an important role in personalized recommendations. CAUIP further improves Precision when recommending 10 and 20 interest points, suggesting that algorithms incorporating user personalized preferences better meet user needs. The TimeSVD++ algorithm, which considers time factors, performs well in Precision, especially when recommending 10 interest points, reflecting the importance of time information in capturing users’ dynamically changing preferences. CBPR shows relatively high Precision in both scenarios, indicating that collaborative filtering combined with content information can recommend more accurately. The algorithm proposed in this paper achieves the highest Precision when recommending 10 and 20 interest points, indicating that it performs very well in integrating student preferences and contextual information. This result demonstrates that by combining students’ personalized preferences and contextual information, the proposed algorithm can more effectively pinpoint students’ interest points and provide a more accurate recommendation list. This integrated approach not only improves the relevance of recommendations but also dynamically adjusts recommendations according to students' learning situations and environments, adapting to their real-time needs. Comparison of precision of different personalized learning interest point recommendation algorithms.
Figure 7 shows the Recall comparison of different personalized learning interest point recommendation algorithms when recommending 10 and 20 interest points. Similarly, the personalized learning interest point recommendation method proposed in this paper, which integrates student preferences and contextual information, displays the highest Recall in scenarios of recommending both 10 and 20 interest points. This means it can more comprehensively identify and recommend content that students may find interesting. A high Recall value proves the effectiveness of this method in capturing user interests, especially when the recommendation list is longer, as the algorithm proposed in this paper can better meet user needs and cover a broader range of user interests. The significant effectiveness of this algorithm is due partly to the accurate modeling of students’ personalized preferences and partly to the reasonable utilization of contextual information, making the recommendations more aligned with students' actual needs and learning environments. This strategy not only enhances the performance of the recommendation system but also provides strong algorithmic support for creating a dynamic, adaptive personalized learning support system. Recall comparison of different personalized learning interest point recommendation algorithms.
To further enhance the performance of the personalized recommendation system, this paper conducted a detailed analysis and optimization by integrating more content features and behavioral data. Specifically, the research considered not only basic learning behavior data such as browsing history and correct answer rates but also incorporates students’ interaction data, time distribution preferences, and more detailed learning content features like difficulty level and knowledge point connectivity. This high-dimensional data integration enables the algorithm to more comprehensively understand students' learning needs and preferences. Additionally, the study employed machine learning and deep learning techniques to handle these complex datasets. The models trained through these technologies can more accurately predict students’ learning preferences and, based on these predictions, recommend learning content that better meets individual needs. The application of this method significantly improves the performance of the recommendation system, making the personalized recommendations more closely aligned with students' actual learning situations and needs.
Figure 8 shows the nDCG (normalized Discounted Cumulative Gain) comparison of different personalized learning interest point recommendation algorithms when recommending 10 and 20 interest points. According to the data in Figure 8, the personalized learning interest point recommendation method proposed in this paper, which integrates student preferences and contextual information, performs the best in terms of the nDCG metric. It shows that its recommendation list has higher relevance and sorting quality compared to other algorithms, especially when the number of recommendations increases. The algorithm proposed in this paper maintains high recommendation quality, proving its effectiveness and adaptability. This superiority may be attributed to the algorithm effectively combining users’ personalized preferences and contextual information, providing more accurate recommendations. The integration of contextual information helps capture users' interest changes in different situations, while the modeling of personalized preferences ensures the consistency of recommendations with users’ long-term interests. Therefore, the algorithm proposed in this paper can not only recommend learning content that users may find interesting but also do so in an order that users find reasonable. nDCG comparison of different personalized learning interest point recommendation algorithms.
Experimental results demonstrate that the proposed algorithm significantly improves the precision and recall of the recommendation system through its integrated application of students’ personalized preferences and contextual information. The specific improvements noted include that compared to traditional recommendation algorithms (such as SVD, PMF, NMF, and ALS), the proposed algorithm introduces more dimensions of data analysis (including content information, user behavioral preferences, and timing factors) and advanced data processing technologies (such as content-aware filtering mechanisms and time-sensitive models like TimeSVD++). These enhancements effectively address the issues of sparsity and dynamic changes often encountered in personalized recommendations. Moreover, the algorithm accurately identifies and integrates a variety of features closely related to learning interest points, such as the details of learning content, the timing patterns of learning activities, and student feedback, which are often overlooked in traditional algorithms. Through this multidimensional data fusion, the algorithm can more accurately predict students' learning needs, thus achieving higher precision and recall in its recommendations. These improvements not only enhance the relevance and satisfaction of recommendations but also increase the adaptability and sensitivity of the recommendation system in dynamic learning environments, making the recommendations more personalized and reflective of students’ latest learning preferences and needs.
Ethical and privacy considerations related to the content of this paper primarily focus on ensuring the ethical handling and confidentiality of data used in university student mental health counseling services. First, the research must ensure that all student data used comply with data protection regulations, such as the General Data Protection Regulation (GDPR) of the European Union or corresponding laws in other regions, to safeguard students’ personal information from being disclosed or misused. Second, during the intelligent word vector generation and sentiment dialogue generation processes, researchers should consider potential biases produced by algorithms and avoid exacerbating discrimination or misunderstandings toward specific groups during modeling and application. Additionally, the research should ensure the authenticity and effectiveness of the mental health counseling content, avoiding misleading or negative impacts due to technical limitations or errors.
Conclusion
This paper’s research revolves around the construction of personalized teaching paths in English education, aiming to deeply understand students’ learning preferences through big data analysis and accordingly design a recommendation system to provide customized learning content. The study utilized probabilistic association analysis to examine a large volume of student data, identifying patterns and preferences in student learning behavior. This analysis revealed students' individualized needs and learning habits, laying the groundwork for subsequent recommendation algorithm design. Furthermore, a new personalized recommendation algorithm was proposed, which not only considers students’ personalized learning preferences but also integrates contextual information such as study time, location, and device usage habits. This integrated approach ensures that the recommended learning interest points are more precise, personalized, and adaptable to the dynamic changes in students' learning needs.
The paper presented joint analysis results of different categories of personalized learning groups, comparative histograms of student attribute importance by category, and preference analysis experimental results of different methods for analyzing personalized learning preferences of students. These results validated that the probabilistic association-based method for analyzing students’ personalized learning preferences proposed in this study can more deeply understand the learning needs and preferences of different student groups, thereby providing a more accurate and comprehensive solution for designing personalized learning systems. Further, the paper’s algorithm was compared with several others using Precision, Recall, and nDCG as evaluation metrics. Experimental results showed that, whether recommending 10 or 20 interest points, the paper’s algorithm scored higher than other algorithms in evaluation metrics, demonstrating better recommendation effectiveness and efficiency in handling dynamic learning needs and personalized preferences.
The paper successfully applied big data analysis to the field of English education, revealing students’ personalized learning habits and needs through probabilistic association analysis, and designing an efficient recommendation algorithm. This algorithm significantly improved the accuracy and relevance of personalized learning interest point recommendations by integrating students’ preferences and contextual information. The experimental results confirmed the advantages of the paper’s algorithm in recommendation quality and stability, indicating that the proposed method can provide students with customized support, enhance the learning experience, and promote learning outcomes. Therefore, the methodologies and practical applications of this study are of significant importance in advancing the development of personalized English education.
The limitations of this study are evident in several aspects: Firstly, although big data analysis is used to delve into students’ personalized learning preferences, the range and sample of data relied upon may still be confined to specific online learning platforms and user groups, which may not be fully applicable to all learning environments or cultural contexts. Secondly, while probabilistic association analysis and fusion algorithms can effectively recommend personalized learning content, the complexity of the algorithms and their dependence on large amounts of data may limit their application in resource-constrained educational environments. Future research directions could include expanding the scope of data collection to cover more diverse learning environments and cultural backgrounds, which would enhance the universality and applicability of the research. Additionally, the study could explore more efficient and streamlined personalized learning recommendation algorithms, facilitating broader application across various educational settings. These developments could help address the current limitations and improve the adaptability of personalized learning systems to a wider array of educational contexts.
Statements and declarations
Footnotes
Conflicting interest
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
