
Editorial
Select search scope: search across all journals or within the current journal

Knowledge distillation is a process of weight compression where a complex teacher model trains a simplified student model, making the weights and their distribution crucial. This paper investigates the weight distribution in convolutional and fully connected layers of both teacher and student models. For convolutional layers, it was discovered that both teacher and student models exhibit a piecewise power-law distribution. A verification method based on the piecewise power law distribution of convolutional layers was proposed, and the correctness of this law was confirmed. Detailed analysis of the breakpoints and power exponents reveals that the teacher model has smaller breakpoints than the student model; for weights smaller than the breakpoint, the teacher model’s power exponent is lower than that of the student model, whereas, for weights larger than the breakpoint, the teacher model’s power exponent is higher. Based on these findings, a new weight initialization algorithm for convolutional layers was proposed. For fully connected layers, both models demonstrate a skewed distribution. A verification method based on the skewed distribution of fully connected layers was proposed, and the correctness of this law was confirmed. Analysis of the kurtosis and skewness indicates that the student model exhibits higher kurtosis and skewness than the teacher model. Based on these observations, a new weight initialization algorithm for fully connected layers was proposed. Experimental results show that both initialization methods improve the initial and final accuracy of the student model compared to the He initialization method.
The exponential growth of academic papers necessitates sophisticated classification systems to effectively manage and navigate vast information repositories. Despite the proliferation of such systems, traditional approaches often rely on embeddings that do not allow for easy interpretation of classification decisions, creating a gap in transparency and understanding. To address these challenges, we propose an innovative explainable paper classification system that combines Latent Semantic Analysis (LSA) for topic modeling with explainable artificial intelligence (XAI) techniques. Our objective is to identify which topics significantly influence the classification outcomes, incorporating Shapley additive explanations (SHAP) as a key XAI technique. Our system extracts topic assignments and word assignments from paper abstracts using LSA topic modeling. Topic assignments are then employed as embeddings in a multilayer perceptron (MLP) classification model, with the word assignments further utilized alongside SHAP for interpreting the classification results at the corpus, document, and word levels, enhancing interpretability and providing a clear rationale for each classification decision. We applied our model to a dataset from the Web of Science, specifically focusing on the field of nanomaterials. Our model demonstrates superior classification performance compared to several baseline models. Ultimately, our proposed model offers a significant advancement in both the performance and explainability of the system, validated by case studies that illustrate its effectiveness in real-world applications.
In recent years, sequential recommendation has received widespread attention for its role in enhancing user experience and driving personalized content recommendations. However, it also encounters challenges, including the limitations of modeling information and the variability of user preferences. A novel time-aware Long-Short Term Transformer (TLSTSRec) for sequential recommendation is introduced in this paper to address these challenges. TLSTSRec has two major innovative features. (1) Accurate modeling of users is achieved by fully leveraging temporal information. Time information is modeled by creating a trainable timestamp matrix from both the perspectives of time duration and time spectrum. (2) A novel time-aware Transformer model is proposed. To address the inherent variability of user preferences over time, the model combines long-term and short-term temporal information and adjusts the personalized trade-offs between long-term and short-term sequences using adaptive fusion layers. Subsequently, newly designed encoders and decoders are employed to model timestamps and interaction items. Finally, extensive experiments substantiate the effectiveness of TLSTSRec relative to various state-of-the-art sequential recommendation models based on MC/RNN/GNN/SA across a spectrum of widely used metrics. Furthermore, experiments are conducted to validate the rationality of the TLSTSRec structure.
A recommender system is an information filtering system used to predict a user’s rating or preference for an item. Dietary preferences are often influenced by various etiquettes and culture, such as appetite, the selection of ingredients, menu development, cooking methods, choice of tableware, seating arrangement of diners, order of eating, etc. Food delivery service is a courier service in that delivers food to customers by restaurants, stores, or independent delivery companies. With the continuous advances in information systems and data science, recommender systems are gradually developing towards to intentional and behavioral recommendations. Behavioral recommendation is an extension of peer-to-peer recommendation, where merchants find the people who want to buy the product and deliver it. Intentional recommendation is a mindset that seeks to understand the life of consumers; by continuously collecting information about their actions on the internet and displaying events and information that match the life and purchase preferences of consumers. This study considers that data targeting is a method by which food delivery service platforms can understand consumers’ dietary preferences and individual lifestyles so that the food delivery service platform can effectively recommend food to the consumer. Thus, this study implements two stages data mining analytics, including clustering analysis and association rules, to investigate Taiwanese food consumers (
Long-term time series forecasting (LTSF) has become an urgent requirement in many applications, such as wind power supply planning. This is a highly challenging task because it requires considering both the complex frequency-domain and time-domain information in long-term time series simultaneously. However, existing work only considers potential patterns in a single domain (e.g., time or frequency domain), whereas a large amount of time-frequency domain information exists in real-world LTSFs. In this paper, we propose a multi-scale hierarchical network (MHNet) based on time-frequency decomposition to solve the above problem. MHNet first introduces a multi-scale hierarchical representation, extracting and learning features of time series in the time domain, and gradually builds up a global understanding and representation of the time series at different time scales, enabling the model to process time series over lengthy periods of time with lower computational complexity. Then, the robustness to noise is enhanced by employing a transformer that leverages frequency-enhanced decomposition to model global dependencies and integrates attention mechanisms in the frequency domain. Meanwhile, forecasting accuracy is further improved by designing a periodic trend decomposition module for multiple decompositions to reduce input-output fluctuations. Experiments on five real benchmark datasets show that the forecasting accuracy and computational efficiency of MHNet outperform state-of-the-art methods.
In the era of the digital economy, the exploration of useful knowledge from data streams has garnered significant attention due to its wide-ranging applications. However, the rapid and infinite nature of data streams poses challenges for efficiently mining high utility sequential patterns, including strong spatio-temporal constraints and the combinatorial explosion of sequence data search spaces. To address this and adapt to a variety of application scenarios, this paper delves into the investigation and design of an efficient algorithm for high utility sequential pattern mining over data streams based on the sliding window model (HUSP_DS). This algorithm utilizes a projection mechanism within a sliding window to recursively search for all interesting patterns. Additionally, it introduces a novel structure called the dynamic utility index table, which stores information such as the utility and index positions of data stream sequences. Notably, this structure proves highly effective in recursive search processes and utility updates. Comprehensive experimentation, conducted on both real-world and synthetic datasets, have shown that the superior performance of the HUSP_DS algorithm compared to state-of-the-art algorithms. This superiority is particularly evident in terms of temporal and spatial efficiency. Furthermore, the algorithm demonstrates suitability for mining sliding windows of arbitrary sizes, showcasing stable scalability.
Intelligent data analysis rapidly transforms healthcare care by improving patient care and predicting health outcomes through machine learning (ML) techniques. These advanced analytical methods allow intelligent healthcare systems to process large amounts of health data, improving diagnosis, treatment, and patient monitoring. The success of these systems is highly dependent on the quality and balance of the data they analyze. Class imbalance, a situation where certain classes dominate the dataset, can significantly affect the accuracy and effectiveness of ML models. In healthcare, it is not only crucial, but urgent, to accurately represent all conditions, including rare diseases, to ensure proper diagnosis and treatment. For this analysis, data was gathered from six reputable academic databases: ScienceDirect, IEEE Xplore, Scopus, Web of Science, Google Scholar, and PubMed. This review offers a comprehensive overview of current approaches to handling class imbalance, including data preprocessing methods like oversampling, undersampling, hybrid techniques, and ensemble learning strategies such as bagging, boosting, and AdaBoost. It also addresses the limitations of these methods and the ongoing challenges in effectively managing class imbalance in healthcare data. Furthermore, the review explores innovative and promising strategies that have shown success in overcoming class imbalance, with a particular emphasis on fairness, diversity, and ethical considerations, offering a hopeful outlook for the future of healthcare data analysis. The discussion highlights how class imbalance can impact the accuracy and reliability of intelligent healthcare systems, underscoring its significance in improving patient care, healthcare delivery, and the broader medical community.
Magnetic Resonance Imaging (MRI) is a cornerstone of modern medical diagnosis due to its ability to visualize intricate soft tissues without ionizing radiation. However, noise artifacts significantly degrade image quality, hindering accurate diagnosis. Traditional denoising methods struggle to preserve details while effectively reducing noise. While deep learning approaches show promise, they often focus on local information, neglecting long-range dependencies. To address these limitations, this study proposes the deep and shallow feature fusion denoising network (DAS-FFDNet) for MRI denoising. DAS-FFDNet combines shallow and deep feature extraction with a tailored fusion module, effectively capturing both local and global image information. This approach surpasses existing methods in preserving details and reducing noise, as demonstrated on publicly available T1-weighted and T2-weighted brain image datasets. The proposed model offers a valuable tool for enhancing MRI image quality and subsequent analyses.
With the rapid development and popularization of smart mobile devices, users tend to share their visited points-of-interest (POIs) on the network with attached location information, which forms a location-based social network (LBSN). LBSNs contain a wealth of valuable information, including the geographical coordinates of POIs and the social connections among users. Nowadays, lots of trust-enhanced approaches have fused the trust relationships of users together with other auxiliary information to provide more accurate recommendations. However, in the traditional trust-aware approaches, the embedding processes of the information on different graphs with different properties (e.g., user-user graph is an isomorphic graph, user-POI graph is a heterogeneous graph) are independent of each other and different embedding information is directly fused together without guidance, which limits their performance. More effective information fusion strategies are needed to improve the performance of trust-enhanced recommendation. To this end, we propose a
Mining geographic location of social media users is a crucial technology for realizing the mapping of cyberspace to geographical world, which can provide strong support for wide-ranging location-based services. As a typical approach, user geolocation methods based on relationships rely on the assumption of location homophily between users and their neighbors. However, these methods only utilize the geographic influence between pair-wise relationship, resulting in undesired geolocation performance. In this paper, a social media user geolocation method based on geographically compact social subgraphs (SMUG-GCS) is proposed. Firstly, we analyze the relationship pattern among users in geographic proximity, and find a phenomenon that users who are geographically close tend to have tightly social groups. Based on this finding, a subgraph partitioning algorithm is presented which integrates structure compactness and geographical credibility to identify a set of subgraphs, whose nodes are more tightly connected and geographically proximity. Finally, user locations are inferred using the propagation of user information only based on the geographically compact subgraph. Extensive experiments are conducted on three real-world social media datasets. The results show that, compared with 5 typical relationship-based methods, SMUG-GCS improves the geolocating accuracy while reducing storage costs, leading to a significant reduction in median error distance ranging from 26.7% to 82.9%, as well as decrease in storage requirements by up to 56.5%.
The promising Network-on-Chip (NoC) model replaces the existing system-on-chip (SoC) model for complex VLSI circuits. Testing the embedded cores using NoC incurs additional costs in these SoC models. NoC models consist of network interface controllers, Internet Protocol (IP) data centers, routers, and network connections. Technological advancements enable the production of more complex chips, but longer testing times pose a potential problem. NoC packet switching networks provide high-performance interconnection, a significant benefit for IP cores. A multi-objective approach is created by integrating the benefits of the Whale Optimization Algorithm (WOA) and Grey Wolf Optimization (GWO). In order to minimize the duration of testing, the approach implements optimization algorithms that are predicated on the behavior of grey wolves and whales. The P22810 and D695 benchmark circuits are under consideration. We compare the test time with existing optimization techniques. We assess the effectiveness of the suggested hybrid WOA-GWO algorithm using fourteen established benchmark functions and an NP-hard problem. This proposed method minimizes the time needed to test the P22810 benchmark circuit by 69%, 46%, 60%, 19%, and 21% compared to the Modified Ant Colony Optimization, Modified Artificial Bee Colony, WOA, and GWO algorithms. In the same vein, the proposed method reduces the testing time for the d695 benchmark circuit by 72%, 49%, 63%, 21%, and 25% in comparison to the same algorithms. We experimented to determine the time savings achieved by adhering to the suggested procedure throughout the testing process.
Software Defect Prediction plays a crucial role in quality assurance by identifying potential defects early in the software development lifecycle. It is an essential aspect of modern software engineering that significantly contributes to improving software quality and reliability. It utilizes a variety of techniques, including machine learning algorithms like decision trees, support vector machines, and neural networks, to predict defects. A lot of research tried to improve the prediction accuracy but had problems with imbalanced data and hyperparameter tuning of the algorithms. To deal with this, we proposed a novel approach by tuning the hyperparameters of the Synthetic Minority Over-sampling Technique using the Tree-structured Parzen Estimator algorithm within the Optuna framework. Through an analysis of seventeen imbalanced datasets from a different public database, we compare our technique with existing SDP models using K-Nearest Neighbors, Multi-Layer Perceptron, Random Forest, Support Vector Machine, and Extreme Gradient Boosting classifiers. Our findings reveal that optimizing the Synthetic Minority Over-sampling Technique significantly improves the performance of SDP models, resulting in enhanced performance metrics. We have statistically validated our results using Friedman's test.
Artificial Intelligence (AI) is becoming increasingly indispensable across diverse domains as technology rapidly advances. As traditional energy sources dwindle, there's a noticeable pivot towards renewable energy sources (RES). However, to effectively meet energy demands, integrating these RES into smart grids to bolster efficiency is imperative. Despite the transition, ongoing technical challenges persist, specifically in accurately predicting and optimizing smart grid parameters. To tackle these hurdles and enhance smart grid efficiency, various AI techniques are being harnessed. This study leverages real-time energy generation data (MWh) from solar and wind plants over a year, dependent on parameters such as POA and wind speed, respectively. Prediction outcomes are derived using three machine learning (ML) models (XGBoost, CatBoost, and LightGBM) and three deep learning (DL) models (LSTM, BiLSTM, and GRU). From these individual models, two hybrid ML and DL models are developed, yielding promising results. Subsequently, these outcomes are further refined through a parallel fusion approach (PFA), resulting in heightened accuracy and reliability. The implementation of this technique notably reduces error rates to 15.05% for hybrid ML, 19.18% for hybrid DL, and 8.1432% for PFA. This methodology holds substantial potential for future research endeavors, supplementing existing AI models for enhanced efficiency.