Abstract
Artificial Intelligence (AI) is becoming increasingly indispensable across diverse domains as technology rapidly advances. As traditional energy sources dwindle, there's a noticeable pivot towards renewable energy sources (RES). However, to effectively meet energy demands, integrating these RES into smart grids to bolster efficiency is imperative. Despite the transition, ongoing technical challenges persist, specifically in accurately predicting and optimizing smart grid parameters. To tackle these hurdles and enhance smart grid efficiency, various AI techniques are being harnessed. This study leverages real-time energy generation data (MWh) from solar and wind plants over a year, dependent on parameters such as POA and wind speed, respectively. Prediction outcomes are derived using three machine learning (ML) models (XGBoost, CatBoost, and LightGBM) and three deep learning (DL) models (LSTM, BiLSTM, and GRU). From these individual models, two hybrid ML and DL models are developed, yielding promising results. Subsequently, these outcomes are further refined through a parallel fusion approach (PFA), resulting in heightened accuracy and reliability. The implementation of this technique notably reduces error rates to 15.05% for hybrid ML, 19.18% for hybrid DL, and 8.1432% for PFA. This methodology holds substantial potential for future research endeavors, supplementing existing AI models for enhanced efficiency.
Introduction
Recent advancements, including the application of blockchain in distributed renewable energy systems to combat cyber threats, multilayer cyberattack detection using machine learning in IoBC-based networks, and AI-based parameter prediction for solar power systems, have significantly enhanced the security and efficiency of smart grids, thereby supporting the transition to a sustainable energy future.1–3 Many other significant energy challenges necessitate exploring innovative solutions to ensure a sustainable and reliable energy supply. 4 Also, fossil fuel reserves diminish, and the urgency of transitioning to cleaner, more environmentally friendly energy sources becomes increasingly apparent, with environmental pollution from fossil fuels further reinforcing the need for green energy alternatives. 5 In this context, smart grids have emerged as a pivotal solution, offering advanced technologies to optimize energy distribution and consumption, particularly through solar, wind, and thermal energy. Installing these grids, especially for solar and wind energy plants, is relatively straightforward compared to conventional grids, making them accessible and efficient options for harnessing renewable energy. 6
In renewable energy resources, smart grids have emerged as a pivotal solution to address these energy challenges, offering advanced technologies to optimize energy distribution and consumption. 7 Key contributors to this transition are solar, wind, and thermal energy, which are essential components of smart grids. 8 Installing these grids, particularly for solar and wind energy plants, is relatively straightforward compared to conventional grids, making them accessible and efficient options for harnessing renewable energy. 9 Integrating intermittent wind and solar power into the electricity grid presents a significant challenge. 4 Therefore, these grids must demonstrate heightened precision, necessitating the development of advanced methodologies to ensure the accurate forecasting of power and other interrelated parameters within these grid systems. Presently, artificial intelligence assumes a pivotal role in facilitating this predictive process, with the implementation of refined machine learning and deep learning models at the forefront of these advancements. 10 The research in 11 proposes data-driven models, including ANN, to enhance predictability and quantify uncertainties in renewable energy generation. These models, which can accurately predict wind and solar power, offer a cost-effective alternative to traditional energy storage and standby capacity. 11 Accurate solar power prediction is crucial for large-scale renewable energy plants, and a hybrid model combining machine learning and statistical methods is introduced for this purpose. 12 By utilizing diverse machine learning models, such as LSTM, GRU, and Auto-GRU, the model aims to enhance forecasting accuracy. 12 As renewables play a growing role in the grid, the study introduces the Traditional Encoder Single Deep Learning (TESDL) method to predict distributed renewable energy production using deep learning techniques. TESDL significantly improves accuracy over traditional models, making it a valuable addition to the renewable energy sector. 13
Another research presents a novel hybrid deep-learning neural network for 24-h ahead wind power generation forecasting. This approach combines a Convolutional Neural Network (CNN) with a Radial Basis Function Neural Network (RBFNN) and successfully addresses wind speed fluctuations, offering highly accurate predictions. 14 Machine learning models, particularly artificial neural networks (ANN), are extensively used for energy production prediction and asset reliability analysis. The paper reviews various case studies and approaches, emphasizing the significance of ANN techniques in renewable energy source dependability analysis. 15 A review paper focuses on using deep learning techniques for wind speed and wind power forecasting, given the complexity of these nonlinear problems. The paper highlights data processing, feature extraction, and relationship learning stages of deep learning models. 16 The study underlines the importance of deep learning algorithms in modeling solar and wind energy resources and emphasizes their accuracy and performance evaluation. It also suggests developing hybrid deep-learning techniques to enhance prediction performance. 17 To improve photovoltaic power forecasting, especially for short-term predictions, a study introduces a solar irradiation forecaster based on deep learning techniques, highlighting the Convolutional Neural Network's superior accuracy. 18 Accurate wind power generation prediction for offshore wind systems is addressed, introducing the use of high-frequency SCADA data and deep learning neural networks to improve forecasting accuracy. This approach significantly enhances prediction efficiency for offshore wind power systems. 19 The research in 20 aims to create an effective system for wind power prediction, evaluating the accuracy of different machine learning models, including Long Short-Term Memory (LSTM), Gated Reference Unit (GRU), and Recurrent Neural Network (RNN).
Similarly, 21 highlights the role of machine learning algorithms in wind power generation prediction and their use in areas lacking pre-existing models. A comprehensive survey of advanced machine learning and deep learning techniques assisting in renewable energy generation is provided, underlining the challenges and potential benefits of these methods. 22 Solar power forecasting using deep learning models, such as the convolutional neural network (CNN), long short-term memory network (LSTM), and a hybrid model combining both, is discussed in. 23 To enhance the prediction accuracy of wind and photovoltaic power generation, a research paper introduces an intelligent prediction system combining decomposition algorithms and deep learning, demonstrating its effectiveness. 24 The significance of artificial neural network (ANN), machine learning (ML), and Deep Learning (DL) techniques in forecasting renewable energy and load demand is discussed, offering insights into their strengths and weaknesses for microgrid forecasting. 25 To effectively predict wind power generation, the paper presents deep learning models, particularly Long Short-Term Memory (LSTM) and Residual LSTM, as the most accurate methods for wind power generation prediction. 26 In 27 the a pressing need for reliable short-term predictions of solar PV generation, given its dependence on weather conditions. It explores the use of deep learning, specifically the Long Short Term Memory (LSTM) algorithm, to forecast solar power output. Performance evaluations compared LSTM with the Multi-layer Perceptron (MLP) network using several metrics. The results highlight that the LSTM network outperforms MLP, offering more accurate forecasts for different types of days. The combination of deep learning and energy efficiency holds promise for promoting sustainable energy, decarburization, and digitization in the electricity sector. A wind power generation prediction using deep and machine learning models, such as DNN, KNN regressor, LSTM, and more is present in. 28 An optimization technique based on stochastic fractal search and particle swarm optimization (SFSPSO) is proposed to fine-tune LSTM parameters. The main idea of the article/paper is to propose a data-driven approach for wind power forecasting using deep sequence-to-sequence Long Short-Term Memory (LSTM) methods along with a gated recurrent neural network (RNN). The focus is on improving the accuracy of wind power forecasts, to make more reliable predictions for energy management and grid stability. 29
The study's primary objective is to bolster the efficiency and dependability of smart grids by integrating renewable energy sources (RES) through advanced artificial intelligence (AI) methodologies, which involves several specific aims. Firstly, to gather and scrutinize real-time energy generation data (MWh) from solar and wind plants for a comprehensive understanding of the current energy landscape. Secondly, to employ a diverse array of ML and DL models to enhance prediction accuracy and optimize parameters within smart grids. Thirdly, to execute a parallel fusion approach aimed at collectively amplifying the performance of hybrid ML and DL models to diminish prediction error rates, ensuring more precise and dependable energy generation forecasts. Finally, to validate the efficacy of the proposed AI-driven techniques in augmenting smart grid operations and energy management.
Literature review/related work
The research, despite its thorough survey of deep learning approaches for power forecasting in smart microgrids, recognizes a limitation in its reliance on substantial historical data for DL-based applications. The study could improve by suggesting practical solutions, such as optimizing DL-based forecasting with smaller datasets, utilizing advanced data preprocessing and feature engineering techniques, and exploring technological innovations to mitigate infrastructure demands and reduce big data processing requirements in smart microgrids. 30 The study introduces an enhanced ensemble algorithm for solar energy forecasting, yet neglects the computational resource implications of the proposed DSE-XGB model. It fails to address potential challenges linked to heightened computational demands, hindering practical applicability.
To improve feasibility, the study should explore strategies like parallel processing or model optimization to optimize computational efficiency, ensuring scalability and broader applicability. 31 The study proposes a deep learning-based ensemble approach for energy demand forecasting but lacks model interpretability. The complexity of the model may hinder user understanding. To address this, incorporating techniques like feature importance analysis and enhancing transparency in how the ensemble model captures chronological dependencies can improve user trust and understanding. 32 The study introduces a hybrid variational decomposition model (HVDM) for accurate power production forecasting in microgrid farms but lacks explicit interpretability. The intricate nature of HVDM, coupled with the improvised dynamic group-based cooperative search (IDGC) mechanism, may reduce transparency, hindering user understanding. To enhance interpretability, integrating model-agnostic methods or feature importance analysis and providing insights into HVDM's decision-making process would improve user trust and understanding. 33
The study introduces an efficient hybrid solar irradiance forecasting model but overlooks the computational demands associated with the complex strategies. The use of deep long short-term memory-convolutional neural network (LSTM-CNN) models, optimized by a modified whale optimization algorithm, may hinder practical implementation due to increased computational requirements. To address this, the study should explore strategies for optimizing computational efficiency, considering parallel processing or model optimization techniques to ensure broader applicability and scalability of the proposed hybrid model. 34 The study provides a comprehensive overview of ML applications in manufacturing but lacks consideration for data security challenges in the energy sector. While emphasizing data management's importance, it overlooks specific strategies for ensuring security and privacy. To address this, the study should discuss robust safety measures, and explore encryption and access controls. Emphasizing standardized guidelines for secure data handling would enhance the sustainability of ML applications in non-industrial energy management. 35
The paper extensively reviews learning-based short-term forecasting models for smart grid applications and explores various models for wind speed forecasting. However, it lacks explicit consideration of the interpretability and explainability of the 41 employed models. To address this deficiency, discussions on techniques for model interpretability, such as feature importance analysis or model-agnostic interpretability methods, should be incorporated. Insights into how the models respond to different seasonal effects and features in the input data would enhance user understanding and decision-making in energy storage planning and policy recommendations. 36
Similarly, another study on machine learning in smart grids lacks consideration of challenges with lightweight solutions and high-performance processing for decision-making. While identifying issues, it doesn't explore or offer solutions, limiting implementation. To address this, the study should focus on lightweight machine learning and high-performance data processing, potentially using edge computing. This would improve decision-making efficiency, ensure practical applicability in smart grids, and suggest avenues for further research in these areas. 37 The study introduces a hybrid ensemble machine learning model for energy demand prediction but lacks explicit consideration of model interpretability. Although emphasizing forecasting accuracy, it overlooks explanations for the hybrid ensemble model's decision-making processes, hindering user trust and practical implementation. To overcome this, the thesis should integrate discussions on interpretability techniques like feature importance analysis or model-agnostic methods. Insights into how the model responds to various features in the time-series data would enhance user understanding, fostering practical utility in decision-making scenarios. 38
The article reviews ML and DL techniques for renewable energy forecasting but lacks solutions for improving model interpretability, crucial for trust in sustainable energy transitions. Overcoming this, it should discuss techniques like feature analysis and model-agnostic methods, providing insights into complex relationships. This would enhance decision-making in grid operation. The study could also suggest further research to improve interpretability in renewable energy forecasting. 39 The paper, proficient in short-term load forecasting, overlooks key considerations for deploying deep learning in real-world smart grids. It lacks insights into interpretability, computational efficiency, and model robustness, crucial for practical reliability. To rectify this, the paper should address these challenges, fortifying the proposed models’ applicability in smart grid scenarios. Additionally, it could suggest avenues for research to overcome specific challenges in smart grid applications. 40
Hence under the above discussion, this study contributes to the efficiency and reliability of smart grids by integrating renewable energy sources (RES) using advanced AI techniques. It utilizes real-time energy generation data from solar and wind plants, offering a comprehensive dataset for analysis and prediction. Employing a combination of ML and DL models, including XGBoost, CatBoost, LightGBM, LSTM, BiLSTM, and GRU, the study addresses challenges related to prediction accuracy and parameter optimization in smart grids. Introducing a parallel fusion approach to collectively enhance the results obtained from hybrid ML and DL models, the study achieves notable reductions in error rates. Specifically, it achieves reductions in error rates, with hybrid ML, hybrid DL, and the fusion approach.
Methodology
Figure 1, explains a robust data-driven approach for predicting power generation from renewable energy sources, employing a hybrid model that seamlessly integrates DL and machine learning techniques. Initial data collection focuses on two types of renewable energy plants, gathering information on wind speed for wind plants and Plane of Array (POA) irradiance for solar plants, both essential factors influencing power generation. The time-series nature of the data includes dependent variables, combining power output from solar and wind plants measured in Megawatt-hours (MWh), and independent variables, comprising solar irradiance (POA) and wind speed as predictors. Model development involves a hybrid DL model, combining LSTM, Bi-LSTM, and GRU, as well as a hybrid ML model, incorporating XGBoost, LightGBM, and CatBoost. The fusion approach combines predictions from both hybrid models, employing techniques such as averaging or weighted averages, while graphical visualization facilitates a clear comparison of predicted values against actual data. This comprehensive methodology, encompassing data collection, preprocessing, model training, hybridization, model fusion, and evaluation, harnesses the strengths of advanced predictive modeling techniques to address the intricate task of forecasting renewable energy output. The ultimate goal is to enhance prediction accuracy and reliability, crucial for effective energy resource management and planning in the face of the stochastic nature of wind and solar energy availability.

The overall framework of the proposed methodology.
As depicted in Figure 2, the provided flowchart, the outlined procedure delineates a systematic method for constructing a predictive model through a hybrid approach that seamlessly combines DL and ML techniques. The process commences with Raw Input Data, the initial collection of data for the model, followed by Data Preprocessing steps such as Data Integration and Characteristics Analysis, and Data Feature Processing encompassing normalization, handling missing values, encoding categorical variables, feature selection, and engineering. The subsequent phase involves Data Segregation into an Independent Variable Time Series Data Set and a Dependent Variable Time Series Data Set, further split into a Training Set (20%) and a Testing Set (80%). Model Development ensues, incorporating DL models like LSTM, Bi-LSTM, GRU, and a Hybrid DL Model that integrates their strengths, alongside models such as XGBoost, CatBoost, LightGBM, and a Hybrid ML Model that leverages their capabilities. The Model Fusion and Prediction stage employs a parallel fusion approach (PFA), combining predictions from the hybrid DL and ML models through techniques like model averaging or weighted averaging. The final Output represents the conclusive prediction result derived from this fusion.

Flowchart for final simulation results.
This comprehensive and integrated approach, tailored for intricate time series forecasting tasks, concludes the process, producing a robust predictive model that harnesses the advantages of both DL and ML models.
The flowchart in Figure 3(a) illustrates the sequential processes required to construct a hybrid machine-learning model with sequential layers. The procedure starts with the Hybrid Machine Learning Model, which suggests the use of a hybrid methodology that potentially integrates many machine learning algorithms. The Sequential Layer is the first layer in the sequence, after which the LSTM Layer is responsible for processing sequences and time-series data. The subsequent layer is the Bi-LSTM Layer, which operates on data in both the forward and backward directions, followed by the GRU Layer, which is another kind of recurrent neural network layer. The Dense Layer is a kind of layer in a neural network where each neuron is linked to every neuron in the preceding layer. Next, the model is compiled by defining the optimizer, loss function, and measurements. After the compilation, the model is trained using the training data. The Fit Model stage is responsible for changing the model parameters. Ultimately, the well-trained model utilizes its knowledge to generate accurate predictions on fresh data, so ending the whole process.

(a) hybrid machine learning model layer structure (b) hybrid deep learning model layer structure (c) enhanced prediction accuracy via parallel fusion approach.
Similarly, the flowchart in Figure 3(b) outlines the sequential procedures for constructing a hybrid deep learning model, which includes many iterative processes and phases of model training. The procedure starts with the Hybrid Deep Learning Model, which denotes the amalgamation of several deep learning methodologies. Sequential Iteration consists of consecutive layers, followed by iterations of XGBoost, CAT Boost, and Light GBM, which are all robust machine-learning methods. Next, the model is compiled by defining the optimizer, loss function, and measurements. The compilation process in Figure 3(b) shows the model undergoes training using the training data, whereby the parameters are adjusted during the Fit Model stage. The trained model is used to make predictions, so ending the process.
The flowchart (c) demonstrates a parallel fusion method that combines hybrid machine learning and deep learning models to improve the accuracy of predictions. The method starts with the simultaneous execution of both the Hybrid Machine Learning Model and the Hybrid Deep Learning Model. The Parallel Fusion Approach uses predictions from both models to obtain enhanced prediction accuracy. The purpose of this parallel processing is to enhance the overall accuracy of predictions, therefore ending the process.
These flowcharts in Figure 3 provide a systematic representation of the procedures involved in developing and using hybrid machine learning and deep learning models, as well as adopting a parallel fusion technique to improve prediction accuracy.
A suite of DL models, including Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), and Gated Recurrent Units (GRU), is developed. These models excel in managing sequential data and capturing the temporal dependencies characteristic of time series data. Subsequently, these models are amalgamated into a hybrid framework, aiming to utilize their strengths while compensating for their respective limitations.41,42
Hybrid machine learning model
Parallel to the DL approach, a series of ML models comprising eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), and Categorical Boosting (CatBoost) is trained. Renowned for their efficacy with structured data and resilience against various feature types, these models form the backbone of the MLstrand. Similar to the DL approach, these models are integrated into a hybrid structure to harness the collective advantages of each model. A brief explanation of these models is given below.
Extreme gradient boosting (XGBoost)
The predictive modeling process commences with the utilization of sample data, where the dataset is systematically divided into subsets. The subsequent step involves the construction of decision trees, with each subset serving as a basis for the creation of an individual tree. Once the decision trees are in place, the predictive phase ensues, where each tree provides predictions for the specific data points it was trained on as shown in Figure 4. The culmination of this process lies in the amalgamation of predictions from all the trees, culminating in the generation of the final prediction results. This comprehensive approach ensures a robust and nuanced prediction model, harnessing the power of decision trees and the collaborative strength of their collective insights.

The basic structure of XGboost.
The dataset is divided into n subsets D1, D2,….Dn. This is typically done using techniques like random sampling with replacement (bootstrap sampling). For each subset Di, a decision tree is constructed. Decision trees are built by splitting the data on the feature that results in the largest information gain or the smallest loss, which can be quantified by equation (1).
Each decision tree makes an individual prediction for each data point. These predictions are weighted and summed to update the model. The final prediction for a given data point is the sum of the weighted predictions from all the trees as in equation (2).
The process of combining the trees in XGBoost is done through an iterative refinement, where each new tree is built to correct the error made by the existing ensemble of trees. This is called boosting.
The given Figure 5, offers a visual representation of the architecture of a LightGBM model, denoting Light Gradient Boosting Machine. This framework, grounded in tree-based learning algorithms, distinguishes itself by its emphasis on distribution and efficiency. Notable advantages include reduced memory usage, heightened efficiency, enhanced accuracy, and compatibility with parallel and GPU learning. The diagram unfolds key components and concepts integral to the understanding of LightGBM. The Feature Vector, representing inputs to the model, encapsulates a high-dimensional portrayal of data attributes. Decision Trees, constituting a collection of trained trees, serve as predictive engines. The concept of Majority Voting elucidates the process of aggregating predictions from all trees to determine the final prediction. The conclusive stage involves Final Predictions, representing the model's output after amalgamating the collective votes of individual trees. This comprehensive overview sheds light on the intricacies of LightGBM and its core elements.

Basic structure of light GBM.
The input data is transformed into a feature vector, which is a numerical representation that captures the characteristics of the data point. The dimensionality of this vector corresponds to the number of features in the dataset. LightGBM builds decision trees, which are represented here with red and yellow nodes (typically, nodes are color-blind in diagrams, but let's assume they indicate different splits or decisions). These trees are constructed one at a time, and each tree learns from the mistakes of the previous ones (i.e., boosting). In LightGBM, trees are grown leaf-wise (best-first), rather than level-wise. This means that the algorithm chooses the leaf it believes will yield the best split, rather than balancing the tree by splitting level by level. The loss reduction from the leaf expansion can be represented below in equation (3).
Each tree in the ensemble gives a vote for the class it predicts for the input feature vector. In a binary classification problem, a vote might be for “Class I” or “Class II.” The final prediction is determined by majority voting among all trees. This means that the class with the most votes from different trees is chosen as the final prediction.
Mathematically if there are N trees and to present the prediction of ith tree, the final prediction can be written as in equation (4).
The visual representation is provided in Figure 6, offers insights into the conceptual framework of the CatBoost algorithm, a powerful machine learning tool adept at effectively managing categorical features while maintaining robustness against overfitting, especially in the context of a large number of categories. The breakdown of the diagram reveals essential components and associated concepts integral to understanding the inner workings of CatBoost. Commencing with the Sample Data Set, which encapsulates categorical features, the algorithm employs a Preferred Statistical Approach for Categorical Features. The Feature Combination step follows, wherein categorical features transform numerical values suitable for algorithmic processing. The Construction of N Trees Independently involves building trees using a subset of the data through bootstrap sampling. Weight Increase comes into play, enabling adjustments to the weight of misclassified points. The Mean Weights of All Predictions culminate in the averaging of predictions from all trees, contributing to the generation of the final prediction. The Prediction Results, representing the output after averaging, underscore the algorithm's efficacy in providing accurate and reliable predictions. This elucidation offers a comprehensive understanding of the CatBoost algorithm and its strategic approach to handling categorical features in machine learning. The process begins with a dataset that includes categorical features, which are non-numerical and typically represent discrete groups or classes. CatBoost has a unique way of processing categorical variables through a combination of one-hot encoding and a special form of target statistics. Instead of transforming categorical values into binary vectors (as done in one-hot encoding), CatBoost calculates statistics based on the target variable for each category and uses these as the numerical representation.

CAT boost diagram structure.
Categorical features are combined into a feature vector after being converted into numerical form. The statistics used for conversion are sensitive to overfitting, so CatBoost uses a process called ‘ordered boosting’ and random permutations to prevent this. CatBoost builds multiple trees where each tree is trained on the data left out by the previous trees. This is represented in the diagram as separate panels for each tree's decision boundaries. Each point's contribution to the loss function is weighted, and points that are harder to predict see their weights increased in subsequent iterations. This is depicted by the size of the points increasing in the diagram. The decision boundary shifts to correct misclassifications. The algorithm computes predictions from all trees, and each prediction is weighted. The final model prediction in equation (5) is the average weighted prediction across all trees:
Figure 7, illustrates a parallel fusion approach utilizing two different hybrid-modified models. One is the hybrid machine learning model (HML) which is the combination of three ML techniques i.e., XG Boost, CAT Boost, and Light Boost. The other one is a hybrid deep learning model (HDL) which consists of GRU, LSTM, and Bi-LSTM. Each of these models generates a set of parallel feature vectors (PFVs) for different segments (indicated by indices), which are then fused to produce a final set of fused feature vectors (Fused V).

Mathematical architectural concept of parallel fusion approach.
Based on the diagram, the mathematical representation of the parallel fusion process could be formulated as follows:
Let
The fused feature vector for each segment ‘j’ can then be represented as in equations (6) and (7) below:
The final fused feature set through the parallel fusion approach can be expressed as the collection of all fused feature vectors as in equation (7):
Section 4 explains the analysis of the observed data from both solar and wind plants in this study, followed by a discussion on the losses and the outcome.
The results and discussion section explores the outcomes of the study and discusses their significance in detail.
Actual data analysis for solar and wind plant
Dependent (MWh) and independent (POA) parameter analysis of solar plant
A box plot, often referred to as a box-and-whisker plot, visually represents the distribution of solar plant data by displaying the five-number summary of a dataset given in Table 1 and shown in Figure 8(a), the minimum, the first quartile (Q1), the median (Q2), the third quartile (Q3), and the maximum. The data is divided into four quartiles for the independent variable POA (Plane of Array). The first quartile (Q1) spans from 0.4 to 0.7, the second quartile (Q2) or median ranges from 0.7 to 0.8, the third quartile (Q3) ranges from 0.8 to 0.88, and the fourth quartile (Q4) spans from 0.88 to 1.0. Similarly, for the dependent variable MWh (Megawatt-hours), the data is also segmented into four quartiles. The first quartile (Q1) ranges from 0.39 to 0.67, the second quartile (Q2) or median spans from 0.67 to 0.78, the third quartile (Q3) ranges from 0.78 to 0.88, and the fourth quartile (Q4) spans from 0.88 to 1.0. These box plots offer a comprehensive view of the data distribution for both POA and MWh variables, highlighting central tendencies, variability, and the potential presence of outliers within each dataset.

Actual data analysis of solar plants independent (POA) vs dependent (MWh) data: (a) box plot; (b) histogram; (c) heat map.
Summary of statistical data for solar and wind energy metrics.
Figure 8(b), presents histograms for the dependent variable, MWh from the solar plant, and the independent variable, POA. In the green histogram depicting POA data, the x-axis displays the values of POA, while the y-axis represents the frequency of these values. The histogram peaks at a frequency corresponding to a POA value of 0.8. Conversely, the second histogram, representing MWh data, displays MWh values on the x-axis and their respective frequencies on the y-axis. Here, the histogram peaks at a frequency corresponding to an MWh value of 0.86. Also mention in Table 1. Similarly, the correlation between MWh and POA of the solar plant can be interpreted from Figure 8(c), and Table 1. The correlation coefficient between MWh and itself (MWh) is 1.0, which indicates a perfect positive correlation, meaning that as MWh increases, MWh also increases in a perfectly linear manner. Similarly, the correlation coefficient between MWh and POA is 0.55. This value indicates a moderate positive correlation between the two variables. It suggests that as the MWh values increase, the POA values also tend to increase, but not in a perfectly linear manner as indicated by the coefficient being less than 1.0. The correlation coefficient between POA and itself (POA) is 1.0, indicating a perfect positive correlation, implying that as POA increases, POA also increases linearly.
The box plot in Figure 9(a), illustrates the distribution of the wind speed and MWh data variables for the wind plant. For the wind speed data, the x-axis delineates ranges from 0.0 to 1.0, with the first quartile spanning 0.0 to 0.32, the second from 0.32 to 0.46, the third from 0.46 to 0.63, and the fourth from 0.63 to 1.0. Conversely, the MWh data's x-axis ranges from 0.0 to 1.0, with quartiles divided into 0.0–0.1 for the first, 0.1–0.21 for the second, 0.21–0.58 for the third, and 0.58–1.0 for the fourth. Within the box plot, the box itself represents the interquartile range between the first and third quartiles, with the median marked by a line inside the box. The whiskers extend to the minimum and maximum values within 1.5 times the interquartile range from the quartiles. This visualization aids in comparing the distribution, identifying any outliers, and discerning patterns or discrepancies between the wind speed and MWh data sets as shown in Table 1.

Actual data analysis of wind plants independent (POA) vs dependent (MWh) data: (a) box plot; (b) histogram; (c) heat map.
In Figure 9(b), the histogram displays the distribution of the wind speed and MWh data variables. In the green graph, the x-axis represents the wind speed values, while the y-axis indicates the frequency of these values. The peak frequency occurs at 0.3 and decreases gradually towards 1.0. In contrast, the second graph illustrates the MWh values on the x-axis and their corresponding frequencies on the y-axis. Here, the peak frequency is observed at 0.1, followed by a gradual decline in frequencies as the MWh values increase. The specific information is also given in Table 1.
The correlation of the wind plant data i.e., between MWh and wind speed is significant, as indicated by the correlation coefficient of 0.66 as shown in Figure 9(c). This value suggests a moderate positive correlation between the two variables as shown in Table 1. A correlation coefficient of 1.0 implies a perfect positive correlation, meaning that as one variable increases, the other also increases proportionally. In this case, the correlation coefficient of 0.66 indicates a positive relationship between MWh and wind speed, albeit not as strong as a perfect correlation. This moderate positive correlation suggests that higher wind speeds are associated with higher MWh production, although other factors may also influence the relationship.
Table 2 and Figure 10, provide a thorough comparison of multiple ML models, evaluating their performance measures for both solar and wind energy data. The examined models include individual ML models (XGB, LightGBM, CatBoost), individual DL models (LSTM, Bi-LSTM, GRU), hybrid machine learning model (HML), hybrid deep learning model (HDL), and parallel fusion approach (PFA).

Comparative analysis of errors through machine learning, deep learning, HML, HDL models, and PFA for solar and wind energy plants: (a) mean square error (MSE); (b) root mean square error (RMSE); (c) root mean square logarithmic error (RMSLE); (d) coefficient of determination (R2); (e) medean absolute error (MedAE).
Comparison of solar and wind error values data.
The evaluated metrics include Mean Squared Error (MSE), Root Mean Squared Error (RMSE), coefficient of determination (R2), Root Mean Squared Logarithmic Error (RMSLE), and Median Absolute Error (MedAE). In evaluating the mean squared error (MSE) for solar data, CatBoost has the lowest error rate of 0.1380, followed closely by Fusion and HML with error rates of 0.1352. Regarding wind data MSE, XGB has the largest error rate of 0.3725, while Fusion and HML have the lowest error rate, which is 0.1535. When considering the Root Mean Square Error (RMSE), Fusion has the smallest error of 0.0368 for solar data, whereas HDL has the smallest error of 0.0735 for wind data. The R-squared values indicate the degree to which the models accurately represent the real data. HML has the greatest coefficient of determination (R2) value of 0.9516 for solar data, suggesting its better performance. In the case of wind data, both Fusion and HML lead with an R-squared of 0.8249. RMSLE quantifies the proportion between the anticipated and actual values, where smaller values indicate more accuracy. In this case, the HML model has the lowest Root Mean Squared Logarithmic Error (RMSLE) of 0.0215 for solar data. As for wind data, both the Fusion and HML models have the lowest value of 0.0735. MedAE is the median of the absolute differences between the expected and actual values. HML achieves the lowest error of 0.0159 for solar data, and once again leads with an error of 0.0356 for wind data.
Therefore, while each model has its advantages in various metrics, the Fusion model constantly demonstrates strong performance, particularly in terms of MSE and coefficient of determination (R2) values. This highlights its potential for accurately predicting energy in both solar and wind applications.
Analysis through individual, hybrid machine learning models and parallel fusion approach for a solar plant
The graph in Figure 11(a), compares actual solar power production data to predictions from proposed ML and DL models over 70 days. Each model's predictions are labeled with a color in megawatt-hours (MWh) versus time in days, making them easy to compare to the plotted data. XG Boost (XGB), known for its gradient-boosted decision trees optimized for speed and performance; CAT Boost (CAT), known for its categorical features and overfitting resistance; and Light GBM (Light), a tree-based gradient boosting framework for efficient, distributed training, are among the ML models considered. DL models also use recurrent neural networks (RNNs) like LSTM (Long Short-Term Memory), Bi-LSTM (Bidirectional LSTM), and GRU (Gated Recurrent Unit), which reduce computational complexity by consolidating forget and input gates.

(a) MWh from ML and DL models (b) loss of ML and DL models (c) MWh of hybrid ML model (d) loss of hybrid ML model (e) Mwh of hybrid DL model (f) loss of hybrid DL model (g) MWh of parallel fusion approach (h) loss of parallel fusion approach.
The graph's vertical axis, which shows MWh production relative to the plane of array (POA), emphasizes solar panel positioning relative to the sun. Model fidelity is measured by its closeness to the “Actual Data” line. Notably, although all models follow the overall data trends, tiny variabilities appear as peaks and troughs, perhaps due to meteorological changes, seasonal shifts, and intrinsic irradiance variability. However, Figure 11(b), shows the training and validation loss trajectories of multiple ML and DL models spanning epochs during training, revealing their performance characteristics. Minimize loss, since lower numbers imply higher performance. In the analyzed models, XGBoost shows a quick drop followed by a steady loss, indicating effective learning; CatBoost shows a fast initial improvement and constant, stable loss; and LightGBM shows a declining and stabilizing RMSE trend. In addition, LSTM and BiLSTM models reduce losses. However, validation loss significantly surpasses training loss, suggesting overfitting. Both training and validation losses decrease with GRU, showing good generalization without overfitting. The x-axis depicts epochs, or training dataset cycles, while the y-axis exhibits loss. Models should lose similarly throughout training and validation. A large difference between these markers may indicate overfitting, which reduces the model's ability to generalize to unknown data.
Figure 11(c), compares the actual MWh production of a solar plant's plane of array (POA) to a hybrid machine learning model's predictions using XG Boost, CAT Boost, and Light GBM. Data and predictions are compared graphically across 70 days using distinct lines. The x-axis shows days and the y-axis solar plant MWh production about POA. The two lines aligning show that the hybrid model can nearly simulate observable data. Accurate predictions need to correlate the expected line with the actual data. Figure 10(d), shows the hybrid machine-learning model's training and validation loss across 50 epochs. The training loss indicates how well the model represents the training data. The validation loss measures the model's ability to forecast current and future data by doing well on a distinct dataset. Both lines decline significantly early on, indicating quick learning. After that, the lines plateau, suggesting the model has attained its optimal performance. When the training and validation lines are near the latter phases, the model demonstrates good generalization and low overfitting.
Figure 11(e), compares a solar plant's plane of array (POA) MWh output to a hybrid DL model's forecast. This hybrid model uses LSTM, Bi-LSTM, and GRU deep learning architectures. The x-axis shows days, while the y-axis shows megawatt-hours (MWh) of solar power plant energy. To demonstrate its accuracy, the hybrid DL model should project a line that matches the real line. Figure 11(f), shows the hybrid deep learning model's training and validation loss over training epochs. As the model learns from the data, the training loss compares its predictions to the training dataset's values. This reduction should be decreasing, indicating that the model is improving on the training data. The validation loss measures how successfully the model applies what it's learned to new data. Its performance is evaluated on a different validation dataset. This number indicates overfitting, which happens when the model performs well on training data but badly on unknown data, or underfitting, which occurs when the model performs poorly on both training and validation datasets.
Figure 11(g), compares actual solar power production data from the plane of array (POA) with parallel fusion model predictions. The ‘Actual Data’ line shows verified solar power output data, while the ‘Fusion Predictions’ line shows fusion model projections. The graph's x-axis shows length in days and the y-axis shows standardized solar power output in megawatts compared to the POA. Estimating solar power generation, grid management, and strategic energy planning needs predictive research. ‘Fusion Predictions’ and ‘Actual Data’ lines overlap well, indicating the fusion approach's accurate prediction ability. Across epochs, Figure 10(h), shows the performance of XGBoost, CatBoost, LightGBM, LSTM, BiLSTM, and GRU on solar data. Better accuracy is indicated by lower ‘Training Loss’ values for model fit to training data. Both losses in ‘Validation Loss’ should converge without overfitting to test generalization to fresh data. ‘Actual Data’ tests the model with unseen data from a distinct test set.
‘Fusion Actual’ and Fusion Validation Loss’ curves imply fusing model results. Training losses drop rapidly and stabilize, but validation losses vary, reflecting different generalization abilities.
Figure 12(a) explains a comparison of forecasts from XGBoost, CatBoost, LightGBM, LSTM, BLSTM, and GRU against wind farm power production data in MWh correlated with wind speed over an extended period. While XGBoost reacts to data fluctuations and broadly aligns with trends, LightGBM consistently predicts close to actual data. CatBoost shows minor variances, and GRU smoothes data extremes but may underfit. LSTM reveals temporal inconsistencies, and BLSTM portions align with data, indicating its ability to recognize prior and subsequent data contexts. All models mirror the data trend, yet quantitative error metrics enable more precise comparisons. BLSTM and LightGBM prove accurate, while others risk overfitting or missing sequence dependence. To evaluate their performance, consider the wind plant's operational conditions and each model's training data and settings, analyzing both visual alignment and numerical metrics. Switching gears to Figure 11(b), it showcases the training and validation loss of various ML and DL models designed to forecast wind plant production based on wind speed. Here, XGBoost and CatBoost demonstrate swift and effective learning, converging to low losses without overfitting. LightGBM learns efficiently but shows slight overfitting, evidenced by a gap between training and validation losses. LSTM converges robustly but achieves minor reductions in losses compared to gradient boosting models, suggesting it may not fully grasp the data's complexity. Conversely, BiLSTM and GRU start with higher losses that progressively diminish, with GRU maintaining consistently low losses.

(a) MWh from ML and DL models (b) loss of ML and DL models (c) MWh of hybrid ML model (d) loss of hybrid ML model (e) Mwh of hybrid DL model (f) loss of hybrid DL model (g) MWh of parallel fusion approach (h) loss of parallel fusion approach.
Figure 12(c), brings forth the factual wind power production data as a brown line, contrasting with predictions from a hybrid machine learning model depicted by an pink line. The x-axis spans a period of 60 days, while the y-axis represents wind power production normalized to wind speed. Alignment between the brown and pink lines signifies the hybrid model's accuracy in prediction, offering critical insights into its performance. Shifting the focus to Figure 12(d), it illustrates the training and validation loss of the hybrid machine-learning model applied to wind data. The orange line tracks the training loss across epochs, while the blue line represents validation loss. Both lines consistently decline, indicating successful learning without overfitting. This graph underscores the hybrid model's ability to adapt effectively to unfamiliar data, marking it as a promising method for wind power output forecasting.
Figure 12(e), highlights a comparison between the “Actual Data” of wind generation and the “Hybrid-Predictions of DL Models” over 60 days. The hybrid model closely mirrors the real data, revealing a strong correlation and the model's ability to capture wind speed fluctuations. Occasional discrepancies between projected and actual values suggest areas for improvement. Despite the lack of specific dates and times, the model effectively represents the fundamental cyclical patterns in the data, which is crucial for optimizing grid management and energy distribution planning. In Figure 12(f), we delve into the training and validation loss of hybrid DL models trained on wind data across various epochs. Initial learning is successful, as evidenced by low loss values. Both training and validation loss converge, indicating the model avoids overfitting. The flattening of the lines towards the end implies additional epochs may not enhance the model further. While low loss values indicate superior performance, they don't guarantee precise predictions on unfamiliar data.
Therefore, these graphs offer valuable insights into model learning and performance but should be complemented with predictions on unknown data for a comprehensive evaluation. Figure 12(g), unveils a fusion machine-learning model's predictions compared to actual wind power production data, both plotted against wind speed over several days. The actual data, represented by a blue line, serves as a benchmark for assessing the model's performance. The fusion predictions, shown by an orange line, should closely resemble the actual data, reflecting accurate prediction.
This graph visually evaluates the model's accuracy in forecasting wind power production by examining the alignment of these two lines. Lastly, Figure 12(h), charts the training and validation loss curves for various machine-learning algorithms, including XGBoost, CatBoost, LightGBM, LSTM, BiLSTM, and GRU, alongside real data and fusion loss curves. Training loss generally decreases across epochs as the model learns, while validation loss measures the model's performance on new, unseen data. The parallel fusion strategy integrates predictions from diverse models to potentially enhance accuracy. Analysis of this graph facilitates model performance assessment, identifies the most effective model, and allows hyperparameter refinement to improve wind power production forecasts. Table 3 given below shows a summarized comparison of all the above discussions.
Comparative analysis of prediction performance and losses for solar and wind power plants.
In the following Section 5, a comprehensive summary of the results obtained from the experiments conducted is provided, and conclusions are drawn based on the findings.
The research centered on utilizing real-time operational data from a solar power plant and a wind power facility, focusing on energy generation measured in megawatt-hours (MWh). To enhance prediction accuracy, the study employed three DL models Long short-term memory (LSTM), bidirectional Long short-term memory (bi-LSTM), and Gated Recurrent Unit (GRU) alongside ML methods such as XG Boost, Cat Boost, and Light GBM for data analysis and visualization. Hybrid machine learning (HML) and hybrid deep learning (HDL) models were developed based on independent variables like plane of array (POA) for solar plant MWh generation and wind speed for wind plant MWh generation. Subsequently, a parallel fusion approach (PFA) was implemented to optimize the integration of hybrid model outputs. Comparative evaluation between individual hybrid models (HML and HDL) and the parallel fusion approach (PFA) revealed the latter's superiority in terms of enhanced accuracy and efficiency. Inclusively, this study highlights the effectiveness of combining hybrid modeling outputs through the parallel fusion approach (PFA), leveraging real-time data from diverse renewable energy sources (RES) to achieve minimal losses and low error rates. Specifically, the hybrid ML model demonstrates a 15.05% error rate, while the hybrid DL model shows a slightly higher error rate of 19.18%. Impressively, the parallel fusion approach (PFA) outperforms and achieves the lowest error rate at 8.1432%. Although the LightGBM model exhibits the lowest coefficient of determination (R2) values, with 0.0898 for solar and 0.0733 for wind, the CatBoost model stands out with the highest R2 values of 0.0951 for solar and 0.0716 for wind. This study offers a promising and efficient technique for future research endeavors, complementing other advanced AI models and contributing to the ongoing enhancement and optimization of smart grid systems.
Footnotes
Acknowledgments
This work was supported in part by the Department of the National Natural Science Foundation of China under Grant 51877149 and 2023YFB4204700.
Credit author statement
Muhammad Abubakar: Methodology; Writing – original draft; Writing – review & editing. Yanbo Che: Formal analysis; Funding acquisition; Supervision; Project Administration. Ahsan Zafar: Investigation; Visualization. Muhammad Shoaib Bhutta: Conceptualization; Validation; Resources. Mahmoud Ahmad Al-Khasawneh; Data curation.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement
The datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request.
