Abstract
Music significantly influences dance performances by shaping the overall mood and energy. However, the challenge of selecting appropriate tracks that seamlessly align with diverse dance styles often hampers choreographers’ creative expression and audience engagement. The objective of the study is to develop an intelligent system that utilizes music information retrieval (MIR) and artificial intelligence (AI) techniques to provide music selection and matching suggestions for dance creations. The study collects data from musical tracks, which included various genres and styles tailored to different dance forms. Initially collected data are preprocessed to enhance quality and remove noise. Feature extraction is performed using convolutional neural networks (CNNs), which analyze time-frequency representations of the music to capture relevant musical features. This feature extraction process is integral to the MIR framework, enabling the system to discern patterns and attributes with the music. A suggestion system is then developed, utilizing the extracted features to match music tracks to dance styles. The study proposed a derivative-free optimized refined random forest (DFO-RRF) method that effectively enhances model performance by fine-tuning hyperparameters without gradient calculations and improves accuracy in matching music tracks to dance styles through efficient feature utilization and model performance tuning. The result demonstrated a significant improvement in music-dance matching accuracy compared to traditional methods. The DFO-RRF method, which is the recommended approach, contains the highest peak results in terms of accuracy (96%), recall (96%), maximum load system (2550), and recommendation error (2.5). The integration of music information retrieval (MIR) and derivative-free optimized refined random forest (DFO-RRF) in this study provides a novel and innovative approach to music selection and matching for dance, thereby significantly enhancing the choreographic process. The study enhances dancers’ and choreographers’ creative possibilities by providing music suggestions based on their performances, streamlining the duration process.
Keywords
Introduction
The fusion of music and dance has usually been an essential component of creative expression, with track serving because of the riding pressure that guides the movement, strength, and emotional music of a performance.
1
As choreographers and dancers are trying to push creative obstacles, the technique of selecting and matching music for dance creations has grown increasingly sophisticated. In the digital age, MIR and AI strategies provide transformative opportunities for optimizing this system, enabling greater particular and personalized music picks that align with the inventive, imaginative, and prescient of the dance.
2
Dance is predicated heavily on song to bring its emotions and movements through choreography. It takes a profound comprehension of each of the creative causes underlying the dance and the attributes of the music to provide matched picks for dance productions and track selections.
3
By setting the tone, rhythm, pace, or even the narrative of a dance performance, the proper tune can also grow its standard effect. Dancers and choreographers can beautify the emotional effect and visible splendor in their paintings by way of precisely coordinating music with choreography.
4
MIR is an area of research and era that focuses on studying and extracting meaningful capabilities from the tune, such as rhythm, pace, melody, harmony, timbre, and temper. By using MIR techniques, dancers can gain deeper insights into musical qualities and choose music that matches their dance intentions, techniques, and intended emotional arcs.
5
This advanced evaluation not only facilitates the identification of a track that perfectly enhances choreography in terms of rhythmic synchronization, dynamic contrasts, or thematic storytelling but also directly addresses the practical challenges faced by choreographers in dance settings. AI, in addition, complements this system by introducing machine learning (ML) algorithms that can robotically recommend tracks primarily based on the unique characteristics of a dance.
6
AI-driven structures can analyze a choreographer’s options, the dance fashion, and the supposed emotional tone to curate tailored tune choices. These intelligent systems can learn from large databases of tune and choreography, making an allowance for particularly accurate matching of tune to movement. From classical to modern, AI can accommodate numerous genres and styles, facilitating the invention of unique musical compositions that raise the general dance performance.
7
Through the mixing of MIR and AI, choreographers can refine the tune selection procedure using automating repetitive obligations, providing recommendations that align with favored aesthetics, and exploring new creative opportunities.
8
This generation-driven method also opens doorways to the seamless synchronization of tune and dance in actual time, where AI can dynamically adapt the music because the choreography evolves. Music choice and matching for dance creations can be appreciably more advantageous with the use of MIR and AI strategies. These technologies analyze musical factors including pace, rhythm, melody, and mood, taking into account precise alignment with choreography.
9
By leveraging AI, choreographers can automate the procedure of finding a song that fits unique dance patterns or movements, ensuring a continuing integration of sound and movement. This approach not only streamlines innovative workflows but also opens new opportunities for revolutionary and personalized dance performances.
10
The aim is to develop a shrewd gadget with the usage of MIR and AI to provide correct song selection and matching suggestions for dance creations. Figure 1 depicts the music selection and dance creations. Music selection and dance creations.
The main contributions of this work are as follows: (1) The research challenge goals to create an intelligent machine, which could pick music and endorse suitable music for dance compositions by means of the use of strategies from MIR and AI. (2) The music tracks used with the observation are a lot of genres and patterns specifically proper to distinct dancing types. Preprocessing is completed on the primary set of facts to enhance first-rate and eliminate noise. (3) The outcome showed that, in comparison to conventional techniques, music-dance matching accuracy had significantly improved. A unique method for choosing music for dance is provided by this study’s fusion of MIR and DFO-RRF.
Related works
The purpose of these groups was to investigate how various improvisation forms impact the development of innovative and creative thinking, as well as to identify the best option for fostering creative thinking and innovative thinking linkage ability. 11 The online flipped educational strategy has been demonstrated to improve students’ learning and happiness throughout Shubailan music-making procedures. It also provided suggestions for online music educators. 12 It established a conceptual framework that lessens the warfare between domain-fashionable and domain-unique viewpoints on creativity, as well as between formerly proposed theories of creativity in musical and non-musical domain names. 13 By examining the psychological barrier that stands between mental illness and mental difficulties in dance choreography, the study sought to shed light on this gray area. Using the human body as a material vehicle, improvisational dance was a free dance style in which the dancer’s inner ideas were expressed via dance movements. 14 Given the date and setting of her pieces, a greater discussion of specific ideas and approaches for the music theater pieces was welcomed on this page. A specific hybridism was combined in music theater works that contrast several creative genres, including dance, music, and theater. According to the investigation’s findings, students’ perceptions of the value of music education in incorporating creative education into the classroom were shaped by their music professors. 15 The development of an inclusive design approach, which informed the creation of three wearable musical instruments, was covered. It combined first- and third-person perspectives to examine these case studies by concentrating on the embodied, somatic, and tacit components of movement-based musical engagement. 16 A qualitative method for the case study was used in the investigation, with an emphasis on performance studies in interactions, workflows, and dance practice. The findings demonstrated that some Indonesian traditional dances have not evolved as a result of rigorous presentation guidelines and adherence to movement patterns. 17 The pedagogy of a higher education topic was embedded with features of embodiment, enriched surroundings, tool usage, multimodal integration, and sensor motor integration, which led to the development of unique, integrated artistic works by every student, as evidenced in the preceding section. 18 Understanding the importance of creativity in AI research and how creative practices can be accessed through computers first provides a definition and assessment of creativity in computers and humans. 19 By analyzing haptic-audio creation, an analysis of the concepts and technologies that enable instrumental dynamics with digital musical instruments was provided. 20 An efficient method for creating harmonic musical mashes was presented. The experiment involved synthetic analysis of music’s melody, beat, and lyrics. The similarity ratings for rhythm, melody, and rhyme in the lyrics were used to assess the “harmony” of the mashup transition. 21
Seven expert composers were given a 15-item, open-ended questionnaire as a part of a qualitative examination to investigate the form of perspectives and reports worried inside the introduction of score-primarily based music. Using a grounded concept methodology, human beings categorized the six distinct codes that emerged from information into higher-order classes. 22 The purpose of the project was to investigate how deep learning (DL) technology can support the long-term growth of the music production sector. It examined the views of experts in Taiwanese music creation additionally analyzed and clarified the significance of DL technology in the music production sector using partial least-squares (PLS) 23 regression. The term “music supervision task” referred to the issue. It expanded upon a self-taught algorithm that discovers a relationship between audio and visual material. For music inspection to yield appropriate suggestions, adequate structure was just as important as sufficient substance. A curriculum reform for history education has to be supported in the context of the Internet of Things (IoT) 24 to enhance the effectiveness of teaching music history. The qualities of teaching data that are easy to acquire against the backdrop of the IoT were first discussed in the history course. The growth of the Internet and technology has led to the proliferation of online music platforms and services that stream music. Information overload owing to an excess of digital audio has grown into a typical concern for many consumers. It has been argued that social tags might be useful for music suggestions. 25 It addressed the long-standing problem of explaining why a listener found a certain piece of music acceptable in the work by focusing on harmony, one of the most important but least studied characteristics. 26 Using harmony, it addressed the long-standing problem of explaining why a listener found a certain piece of music acceptable. 27
Methodology
The contest uses artificial intelligence (AI) and music information retrieval (MIR) tools to apply a systematic approach to intelligent music selection and synchronization programming for dance forms. Originally, a collection of songs with different genres and styles is designed especially for dance styles. The collected data have been preprocessed to improve quality and eliminate noise. Convolutional neural networks (CNNs) are then used for feature extraction, where temporal and frequency-related images of music are analyzed to capture suitable features. Following characteristic extraction, a suggestion machine is constructed that leverages these capabilities to fit-tune tracks to corresponding dance patterns. The derivative-free optimized refined random forest (DFO-RRF) method is proposed to improve overall version performance via great-tuning hyperparameters without counting on gradient calculations. This approach drastically improves the accuracy of song-dance matching by efficiently utilizing extracted capabilities and optimizing model performance, thereby supplying treasured hints for dancers and choreographers. Figure 2 indicates the method flow. Flow of the proposed method.
Dataset
The dataset is collected from the open source Kaggle website: https://www.kaggle.com/datasets/ziya07/dance-music-data/data. The intended data for use in research on AI and MIR approaches particularly focus on matching music tracks to suitable dancing genres. The dataset is appropriate for creating intelligent systems that offer music selection and matching recommendations for different dance forms since it includes synthetic music attributes and dance style designations. The information comprises the following features: (1) Tempo (BPM): Beats per minute, which expresses the music track’s tempo. (2) Mel-frequency cepstral coefficients, or MFCCs, are a group of 13 coefficients that describe musical inflection. (3) Chroma attributes: A set of 12 attributes that correspond to the pitch classes and harmonic content of the song. (4) Spectral contrast: Seven characteristics that show how the music spectrum’s peaks and troughs differ from one another. (5) Dance style: The target label representing different dance styles (e.g., Salsa, Ballet, Hip-Hop, and Tango) to which the music track is best suited.
Preprocessing for music selection and dance creations
For preprocessing music selection and dance creation data, it’s crucial to first categorize the music tracks by genre, tempo, and rhythm patterns. Simultaneously, dance sequences should be analyzed for key movements, speed, and style. Applying feature extraction methods like Mel-frequency cepstral coefficients (MFCCs) in music and motion capture analysis for dance will help streamline the data, preparing it for further analysis or ML tasks. Wiener filtering (WF) is applied for noise reduction in music selection, while min-max normalization is utilized for dance creation.
Wiener filtering (WF) for noise reduction in music
Remove background noise, distortions, or unwanted signals from the music tracks. This can be done using spectral subtraction or WF to ensure clean data. The investigated frequency-domain optimization issue of eliminating residual noise while limiting the output voice distortion contains the following mathematics equations (1) and (2).
The maximum permitted local signal distortion is represented by
In the above instance,
A square matrix’s trace is represented by
Min-max normalization for dance creations
Min-max normalization is a technique used to scale facts within a specific variety, usually [0, 1]. In the context of dance creations, this approach may be applied to standardize movement statistics, allowing choreographers to analyze and evaluate exceptional dance sequences efficaciously. Min-max normalization helps when different features have varying units, making it easier to compare them across tracks. It is acknowledged that min-max normalization preserves any relationship found in the data under consideration. The values in the feature under consideration are transferred to a new normalized value using the following equation (7).
Convolutional neural networks (CNNs) for feature extraction
Music selection and dance creation using CNNs involve leveraging CNNs for feature extraction from audio and movement data. CNNs analyze patterns in musical elements such as rhythm, melody, and beats while also capturing motion features from dance sequences. This enables the generation of dance routines aligned with musical selections, automating choreography creation techniques that link audio dynamics with motion characteristics.
Convolutional layers
Convolutional layers build feature maps in response to diverse function detectors, as formerly cited, and compress the entry using removing features of interest. Neurons filter out basic functions consisting of edges with the first convolutional layer. The neurons gather the capacity to acquire records. Neurons are assigned to explore certain sections of the input picture aircraft, but all convolution kernels are capable of extracting features over the duration of the entire input aircraft. This allows them to produce characteristic maps with identical receptive fields. The depth of the stack is determined by the type of neuron quantity. Every type of neuron adds to the stacking by using the equal weight and bias vector.
Convolution
A single method to determine the output size is to the kernel size, map size, and input size are shown individually by equation (8).
Activation
Dance compositions and musical choices should go on even after a skewed and weighted total. Activation characteristic is required, just like in genuine perception neurons, to break the simple linear buildup of effort and allow neural networks to operate as a traditional approximation of continuous feature.
Pooling layer
Preferring to sub-sampling, pooling encompasses a range of techniques, including generic pooling and overlap pooling. It frequently serves as a different convolutional level. The two pooling algorithms that CNNs utilize most frequently are max-pooling and average-pooling. CNNs greatly benefit from sharing. In particular, pooling reduces the dimensionality of the records by concentrating close data into a pooling window, saving the individual from having to become too involved. One way to simplify computations is to decrease the data’s dimension. In addition, following appropriate pooling, a small number of dislocations or scaling have no effect, bringing about invariance in translation, rotation, and scale.
Classification layer
The last complicated information is gathered by the top layer of a network, referred to as the categorizing layer, which then produces a vector of columns with each row pointing in the desired direction of a class music selection and dance creations. More precisely, the probability estimation for each class is represented by each member in the resultant vector, and the sum of all the components is one. While convolutions and pooling transform raw imagine facts into a function area, the classification layer’s effect at the sample area projection produces a clear instance of class. In most circumstances, the result of genuine CNNs is a fully linked layer. The fully linked type layer is really derived from the AI’s understanding of “characteristic extraction.” All high-order information is combined and reweighted, utilizing full input-to-neuron interconnections to achieve the spatial shift.
Derivative-free optimized refined random forest (DFO-RRF)
Music selection and dance creation using derivative-free optimized refined random forest (DFO-RRF) focus on enhancing creative processes through data-driven algorithms. The DFO-RRF model optimizes feature selection by employing a derivative-free optimization strategy that enhances hyperparameter tuning without requiring gradient calculations. This specific approach allows the model to efficiently explore the parameter space, significantly improving its accuracy in correlating music features with suitable dance movements. By contrast, traditional models often depend on gradient descent methods, which may not perform well in high-dimensional or noisy datasets. RRF technique eliminates the need for derivative information while maintaining high accuracy in matching music rhythms to dance styles, leading to more innovative and dynamic performances. This approach leverages ML to blend art and technology seamlessly in choreography and music duration.
Refined random forest (RRF)
Derivative-free optimization (DFO)
The music selection and dance creation using derivative-free optimization (DFO) leverages advanced algorithms to enhance artistic performance by optimizing musical arrangements and dance movements without relying on gradient-based methods. DFO is particularly effective in exploring large, complex search spaces, allowing for the generation of innovative, rhythmically aligned music choices and choreography that adapt to varying performance environments. This approach can drive creative collaborations between technology and the arts, leading to more personalized and dynamic dance compositions that resonate with both performers and audiences. Formally, the following describes a general single-objective optimization problem equation (9):
The DFO-RRF enhances the accuracy and robustness of predictions using RRF, intricate relationships between person options and numerous musical elements or dance patterns. Additionally, its capability to address nonlinear relationships allows for greater nuanced expertise of how different factors affect a person’s pride and engagement. This method no longer reduces the threat of over fitting, however, additionally improves exploration in the answer area, leading to tailored pointers that resonate with individual tastes. Ultimately, integrating DFO-RRF into song choice and dance creation strategies can lead to more personalized and attractive creative experiences.
Additionally, the practical application of this system involves considerations for user interaction. An intuitive user interface is essential for choreographers and dancers to effectively engage with the system. Features such as personalized music recommendations, easy navigation through musical tracks, and visual feedback on dance-music integration would significantly enhance user experience. Future iterations of the system should prioritize user interface design to ensure accessibility and ease of use for non-technical users.
Results
An intelligent system integrating MIR and a DFO-RRF method to enhance music selection for dance performances has been developed. By using CNNs for feature extraction from various musical genres, the system matched tracks to specific dance styles. The DFO-RRF method fine-tuned the model’s hyperparameters without gradient calculations, significantly improving accuracy in music-dance matching. This approach streamlines the creative process for choreographers, offering tailored music suggestions that align with different dance styles and enhancing overall performance quality. The system is powered by an Intel 16 GB of DDR4 RAM, Xeon E3-1230v5 CPU, and an NVIDIA Quadro K420 discrete graphics card. It also has a 1TB SSD, guaranteeing express and successful storage. Comparison of the proposed method with the existing approaches, such as collaborative filtering (CF), 28 matrix decomposition (MD) algorithm, 28 K-nearest neighbor (KNN) algorithm, 28 and deep belief network (DBN), 28 in the metrics of accuracy, sensitivity, specificity, F1 score, MCC, MSE, MAE, and RMSE.
Accuracy
Outcomes of accuracy.

Analysis of accuracy.
In the context of song selection and dance creations, accuracy can relate to how faithfully a dancer performs choreography, adheres to musical rhythms, or fits the intended emotional expression of the music. It can also pertain to the precision in selecting music that enhances the subject, temper, or style of the dance, improving the overall performance first-rate.
Recall
Outcomes of recall.

Analysis of recall.
When someone is asked to recall a specific music selection or a dance creation they participated in, they rely on their reminiscence to retrieve that information. Recall is critical for activities that include music selection and dance creations, in which artists and performers should recall and draw upon their previous studies and information to create and carry out efficiently.
Recommendation error
Outcomes of recommendation error.

Analysis of recommendation error.
This misalignment can lead to user dissatisfaction, reduced engagement, and decreased effectiveness of the advice gadget. Addressing advice mistakes involves refining algorithms to better recognize consumer preferences, incorporating remarks mechanisms, and utilizing diverse datasets to improve accuracy in predicting user wishes in the realms of music and dance.
Maximum system load
Outcomes of maximum system load.

Analysis of maximum system load.
Similarly, for dance creations, it is able to signify the most complex choreography that a performer can execute, thinking about their bodily obstacles and the want for synchronized movements with the music. Understanding and dealing with the most system load is critical for ensuring premier functionality and accomplishing incredible outputs in each creative and technical domain name.
Discussion
Although widely used for suggestions in a variety of fields, such as dance compositions and music choices, CF approaches have several disadvantages. One foremost trouble is the bloodless begin problem, where CF struggles to make accurate hints for new users or gadgets because of insufficient information. CF strategies also face sacristy demanding situations, as user-item interplay matrices often have many lacking values, leading to unreliable guidelines. Additionally, CF tends to have scalability troubles with huge datasets, as computing similarities among all customers or objects can emerge as computationally high-priced. Lastly, CF can suffer from recognition bias, over-recommending popular items, and ignoring niche possibilities. Matrix decomposition algorithms, which include MD, are powerful gear in diverse packages, but they have a few drawbacks. These algorithms can be computationally luxurious, especially for huge matrices, leading to excessive time and reminiscence prices. They will also be sensitive to noise, resulting in reduced accuracy when applied to noisy or imperfect records. Additionally, certain decompositions, like Eigen price-primarily based strategies, particularly in sick-conditioned matrices, proscribing their reliability. KNN set of rules, while simple and powerful, has numerous drawbacks. One main trouble is its excessive computational value at some point of prediction, especially with massive datasets, because it calls for calculating the distance between the center and all data points. Additionally, KNN is sensitive to noisy data and beside-the-point features that could negatively affect accuracy. It also struggles with imbalanced datasets, where the bulk magnificence can dominate predictions. Lastly, the algorithm requires careful selection of the “k” value, as a flawed choice can cause over fitting. DBN is a type of generative model inclusive of more than one layer of latent variables (hidden devices). It is a stack of restricted Boltzmann machines (RBMs), wherein every layer captures complex styles in the input records. DBNs are skilled layer-by-means-of-layer in an unsupervised manner and can be satisfactorily tuned with the usage of supervised learning methods. Dimensionality reduction, and complicated class obligations, including photograph popularity, speech processing, and natural language information. DBNs are effective due to their potential to version high-level abstractions in data. The DFO-RRF offers several advantages. It complements the version’s overall performance by optimizing hyperparameters without counting on gradient-based techniques, making it suitable for problems in which gradients are difficult to compute or non-existent. DFO-RRF improves accuracy and robustness in prediction responsibilities. Additionally, it can take care of high-dimensional datasets and nonlinear relationships efficaciously while lowering the danger of overfitting. Its derivative-free optimization ensures better exploration of the answer space, leading to greater correctness and strong consequences throughout numerous datasets. In comparison to different processes, the proposed method (DFO-RRF) efficaciously addresses these troubles and yields superior outcomes.
Limitations and challenges
While the proposed system demonstrates significant improvements in music-dance matching accuracy, several limitations warrant consideration. Firstly, the representativeness of the dataset may pose challenges, as it primarily includes popular music genres, potentially limiting the model’s applicability to less mainstream or cultural genres. Furthermore, the scalability of the DFO-RRF model across different musical styles and cultural contexts remains to be rigorously tested. Future research should focus on expanding the dataset to include a broader range of musical styles and cultural expressions to enhance the model’s versatility.
Conclusion
The methodologies utilized in this study connect the theoretical aspects of MIR and AI with the nuanced realities choreographers face in the selection of music for dance, establishing a bridge between academia and practical application. A wide variety of musical tracks created for unique dancing patterns are effectively analyzed and categorized using the DFO-RRF approach in MIR techniques. The substantial improvements in music-dance matching accuracy underscore the effectiveness of the proposed DFO-RRF technique, showcasing its capability to satisfactorily tune hyperparameters and optimize model overall performance without requiring gradient calculations. The consequences of these studies not only facilitate more desirable creative expression for choreographers and dancers but also, streamline the regularly hard process of selecting suitable tune tracks. The DFO-RRF method demonstrated exceptional performance metrics, achieving a peak accuracy of 96%, recall of 96%, maximum system load capacity of 2550, and the lowest recommendation error of 2.5, underscoring its effectiveness as a practical tool for choreographers. By providing tailored music suggestions that resonate with the mood and energy of particular dance forms, the system empowers artists to explore new dimensions of their performances and engage audiences more deeply. The device can also lack the functionality for personalized hints primarily based on person options, past picks, or unique overall performance contexts. Incorporating consumer feedback mechanisms could enhance the relevance of the hints supplied. Future studies should involve the collection and evaluation of a more giant and numerous dataset that includes a much wider range of musical genres, styles, and cultural effects. Incorporating conventional and cutting-edge tunes from various international regions ought to improve the device’s potential to cater to extraordinary dance forms and increase its applicability.
