Abstract
In this article, a new complex method is presented that supports decision-making in sports selection. The method consists of two different parts. The first one is the creation of the framework, that is the assessment of the opinions of trainers about the significance of some attributes in special sports. As a result, weights are constructed based on a generalized 7-option Thurstone method. The second part of the process involves assessing athletes’ performance from various perspectives. Finally, the calculated weights are combined with the performance of the athletes. The scores we obtain quantifiably indicate which sport is most likely to be recommended for a given individual. Moreover, the method developed allows us to quantifiably measure the strength of individual athletes and compare them across age groups equally. A computer program has been developed that implements the evaluation process in addition to this the evaluation results are also presented in this article. By extending this method, it can not only be used for sports selection but also for making various complex decisions where there are numerous possible outcomes to choose from. As a result of this research, we present a generalized set of criteria for the evaluation and describe the results obtained with the method. We explain the relationship between coaching decisions and the recommendations of the method and provide a method for determining the direction of future development for individual athletes.
Introduction
Methods of comparisons in pairs are often used in various rankings, decision-making processes, and determining weights for characterizing the importance of the attributes. When decisions are made, and the number of possible choices is relatively small, there may not be a need to rely on decision support methods, as opposed to more complex ones, the support provided by such methods can be valuable, see Cinelli et al. 1 and the plentiful literature in it.
The idea of the application of paired/pairwise comparison methods has appeared in sports evaluations. The results of matches can be considered as the results of the comparisons: the winner (team/player) is better than the defeated one. There might be equal judgments, too, due to ties, depending on the sports. One of the most commonly used methods is the analytic hierarchy process. 2 A pairwise comparison matrix-based method is applied for a chess tournament by Csató, 3 and ranking the best tennis players of all time was determined by Bozóki et al. 4 A method with stochastic background, the Bradley–Terry method is used by Baker and McHale 5 and by McHale and Morton 6 ; in the latter, the authors focus on predictions. A different stochastic method, the Thurstone method is applied for instance, by Orbán-Mihálykó et al. 7 in the case of handball, presenting the excellent quality in group-aggregations. Moreover, by Gyarmati et al. 8 and Temesi et al., 9 there are some comparisons of the different evaluation methods to clarify their properties.
Apart from making evaluations more refined, mathematics, and operations research are gaining increasing importance in sports science due to their proposed recommendations for modifying competition rules as well. 10
The choice of sport is a practical problem for child athletes, which has many different aspects. Obvious influencing factors include body structure, family connections, and the child’s preferences, but there are many other factors; thus, numerous physical and even more psychological factors influence it. The psychological factors should be examined by using paired comparison methods. A detailed analysis can be found by Magnussen et al. 11 The influence of parents is investigated by Knight, 12 and the influence of the trainer is analyzed by Trigueros et al. 13 An expert system based on psychological and physical aspects of sports selection is presented by Vitchenko et al. 14
In this article, we are focusing on the physical abilities of the athletes, and the relevance of the abilities in the special sports is determined by paired comparison methods.
The role of the age in athlete’s sport choice is investigated by Feeley et al. 15 Barynina and Vaitsekhovskii 16 examined the timing of sport selection and its relationship with career success. Based on data collected of about 250 sportsmen, Moesch et al. 17 concluded that it is worth competing in multiple sports before the athlete makes a final decision on their chosen sport. It is asserted that even more remote forms of exercise can improve performance in various sports.18,19
Various tests are used to measure the physical attributes and mental abilities of athletes. “These tests should be reliable and valid” and “the typical error of measurement of each test should be determined so that the results can be interpreted correctly.” 20 Regular and precise measurement of the performances of the athletes is indispensable see Pichardo et al. 19 Similarly, such performance measurements serve as the basis for the methods we have developed. The significance of the different attributes is evaluated by the subjective opinions of the trainers; therefore, a paired comparison method is applied.
There are numerous different scales (a detailed review can be found by Dawis 21 ), but the method of absolute scaling is related to Thurstone. 22 He generalized the method of Thorndike and the method elaborated by Thurstone is generalized from several points of view: hypothetical distribution, number of possible choices, method of parameter estimation, and as well as numerous other aspects. An excellent comparison of the early method mentioned above is contained by Engelhard. 23 It concludes that the 2-option Rasch model (i.e. logistic distribution with maximum likelihood estimation) is preferable to Thurstone’s original method. Nevertheless, these methods can be generalized and combined. Moreover, González-Díaz, 24 concluded that, from axiomatic points of view the maximum likelihood estimation is a reliable choice and, furthermore it involves the possibility of testing hypotheses.
We want to apply such a method which allows for multiple outcome choices (to be flexible) and would like to preserve the possibility of multilevel decision-making. In addition to these, we want to avoid the necessity of making every possible comparison. Such a method has been presented by Orbán-Mihálykó et al.,
26
called generalized Thurstone method. It applies
Combining the significance of the attributes in different sports discipline and the measured performance of the athletes, we have developed a score characterizing the strength of the athletes in the sports disciplines. This score is suitable for making decisions related to the choice of sports, development in the sports and suggestions for further training to improve faster.
The article is structured as follows: “The model” section contains the description of the mathematical model, while in the “Application of the model” section, the detail of its application in our case is described. The “Calculation process” section contains the algorithm of the calculation process supported by examples. In the “Results” section, the results with the directions related to the reliability and usefulness of the method is given. The “Future improvements” section provides some possibilities for further refinements of the mathematical model. Finally, a brief discussion and conclusion closes the article.
The model
The model created by us consists of two parts. In the first part, we put the importance of specific attributes (based on sports disciplines) under evaluation and assign weights to them.
In the second part, we determine the scores for the objects to be evaluated (individuals) corresponding to the previously assessed attributes. We assign these scores and then multiply them by the corresponding weights associated with each attribute. This process yields the overall evaluation score for the individual.
The first part of the model: Assigning weights to the attributes
The generalized Thurstone method allows more than two options.25–27 We consider latent random variables, denoted by
As Thurstone did,

The options and the intervals belonging to them.
Let
Using the likelihood function, we can estimate the magnitude of the strengths,
In the case of a 7-option model, the following conditions together assure the existence and the uniqueness of the maximizer, fixing The objects compared are represented as vertices of a graph, with a directed edge pointing towards the better option in the case of an extreme (much better) decision, and a bidirectional edge for non-extreme decisions. The resulting graph has to be strongly connected. Each decision (much better, better, slightly better, equal, etc.) has to exist. Finally, there is a directed cycle in the graph along much better edges.
The first condition is needed for the boundedness of the expected values, the second is needed to be able to speak of a truly 7-option model, and the third is to ensure that
During our work, the compared attributes were the physical aspects, and the trainers’ opinions provided the results of the comparisons. After evaluating the opinions, we form weights from the obtained expected values using the following formula:
The second part of the model: Evaluation of individuals
For each individual, we determine the extent to which they fulfill each attribute, separately, by assigning a number between 0 and 1, that is, by re-scaling. For instance, a specific attribute (paced curl-up test) has a minimum value of 45 and a maximum value of 70 for female (see Table 1). If the individual achieves 45 or below, the value of 0 is assigned. If they achieve 70 or above, the value of 1 is assigned. Intermediate values are determined along a linear function. For instance, if the value is 60, we evaluate that the individual fulfills the attribute to a degree of
Scale for 14- and 15-year-old male and female athletes.
Sometimes smaller values indicate better performance and they deserve higher scores. In these cases the minimum value swaps places with the maximum value, therefore it becomes the upper bound and the maximum value becomes the lower bound of the scale. Note that the sign of the score would not change. This is the situation for 60 m with standing start (obviously) and for body weight (according to the opinions of the trainers).
We use the weighted arithmetic mean to calculate the individual overall strength due to its excellent properties in multi-criteria decision making, that is, “monotonicity, continuity, symmetry, idempotence and stability for linear transformations.” Moreover, “it reflects the possibly different importance of single criteria.” 29
The measures obtained from the individuals, which we want to evaluate using the above procedure, are multiplied by the weights previously assigned to the corresponding attributes, and these products are summed up. The resulting score of the individual is therefore a number between 0 and 1, where a higher value indicates better fulfillment of the necessary requirements.
Application of the model
In this section, the application of the above procedure is presented. We measured the performance of athletes in different sports disciplines using the method. We aimed to quantifiably assess the athletes’ performance. We developed a computer application to facilitate the classification of athletes into three sports disciplines: throwing, running, and jumping/sprinting. This application is capable of quantifying the athletes’ strengths in each sport discipline.
In Hungary, since 2014, the NETFIT survey 30 has been used to assess the fitness of students. With this 11-item survey, a fairly accurate picture can be obtained regarding the fitness of individual students/athletes, whether they exert short or long-term force or force by the upper or lower extremities. Similar measures are used in other countries, the first was the FITNESSGRAM® in 1977. 31
The NETFIT surveys were conducted on young athletes, and based on the results of these surveys, we evaluated the performance of each young athlete with a number ranging from 0 to 1.
The attributes, concerning their importance, were compared in pairs by coaches in the case of each sport discipline. The attributes were the elements of the NETFIT survey. At least two coaches performed the comparisons for each sport discipline, and we calculated the weights for the attributes in each sport discipline based on these comparisons applying the generalized Thurstone method.
The weights for the attributes in each sport discipline, calculated by (3), are shown in Table 2. For the sake to be more illustrative, we also presented the data in percentages in Figure 2. The figure and table clearly show that the attributes most closely associated with each sport discipline dominate. In the throwing discipline, the small ball throwing and standing long jump tests are crucial. It may be surprising at first that the standing long jump is important, but in events like hammer throwing and javelin throwing, it is essential to have the ability to generate large forces quickly with the lower extremities, thereby accelerating the sports equipment to be thrown.

Weights of attributes in surveys regarding disciplines.
Weights of attributes in surveys regarding disciplines.
For runners (including middle and long-distance runners), the 20 m pacer test is the most important, indicating their capacity for long-term exertion of force. In the jumping/sprinting discipline, the explosiveness of the lower extremities is crucial, making the 60 m with standing start and standing long jump tests the paramount.
To get a reliable value for athletes in each discipline, we need accurate min–max value scales for gender and age. To do this, however, we need to know very well the healthy, good limits for each age-related trait. Rodriguez-Negro et al. 32 investigated how horizontal jump ability varies in U8–U16 (6–15 years old) boys. Getting good scales required, on the one hand, numerous coaches’ opinions, as well as athletes’ data. With the help of this information, min–max scales were developed, of which only the U16 girls and boys shown due to lack of space (see Table 1). If the scales are used incorrectly, it is difficult or impossible to decide on the discipline of the athletes. For example, if the male–female scales are reversed, the overall strengths in the disciplines achieved by the female athletes will be quite low, while the overall strengths of male athletes will be close to the maximum. This makes it difficult to differentiate between disciplines for each athlete.
Calculation process
In this section, the calculation process build during the research is presented. We use diagrams and tables to illustrate the calculation, and we give explanations to the previous chapters to make the method clear and easy to use. The diagram of the calculation (see Figure 3) shows that there are two branches of the calculation that intersect. In the upper branch, the weights of the attributes are determined, that is, how significant a given attribute is for each sport discipline. On the lower branch, we determine the score of each individual attribute.
In the upper branch, the calculation per discipline is as follows:
Comparing each attribute with the others and get the comparisons’ data matrix. Evaluation of the resulting data matrix using the generalized Thurstone method and getting expectations on the basis of formula (2). Calculate the weights of the attributes for each sport discipline, based on these expectations by formula (3). On the lower branch, the calculation per individual is as follows:
Perform tests, and determine the results of the tests. From the test results, determine the re-scaled values for each trait for each individual. Finally, the two branches meet. For every individual and sport discipline, we take the weighted sum of the traits.
This process results in a score for each individual for each discipline. This score characterizes the performance of the athletes in the sports discipline.

Calculation process.
With the aim of being clear, we present an example for the calculation process in the case of two athletes presented in Figure 4.

Example of the calculation.
Consider individual 1. We calculate as follows. For each test, column “value” contains their performance. If
Results
Naturally, the reliability and usefulness factors of the method we developed require testing. For the sake of validity, we have identified three main directions. The first direction concerns when we classify each athlete into disciplines, and then look at whether our method could correctly determine the orientation of each athlete where they could be most successful. The second direction is related to the performance-measurement in the discipline, that is, whether our software could recognize the discipline of each athlete and whether the score obtained for their performance really reflects their effectiveness in that discipline. The third main direction concerns improving the future performance of athletes. The question was: what could/should be improved for each athlete, and what additional information could be obtained from the data received.
Participants’ data
We cooperated with the VEDAC (Veszprém University and Student Athletics Club) to develop and test a discipline selection method and software.
The number of coaches participating in the test and the number of discipline comparisons they carried out, broken down by section, are shown in Table 3. Each coach has made 55 comparisons in their discipline. To reduce the number of comparisons, incomplete comparisons can be applied. Fewer comparisons per coach may be enough for us, as long as we can take into account the comparisons of more coaches. Concerning the optimal structures, the research is going on. Principles of the optimal structures of the incomplete comparisons can be found by Szádoczki et al. 33 and Gyarmati et al. 34
The number of coaches and comparisons.
We can measure the quality of coaches’ comparisons through the consistency of their comparisons. By determining the rate of circular beatings, we can get an overview of this. If there are lots of circular beatings, the data are inconsistent. Circular beatings are defined as follows: objects to evaluate (attributes) are formed into triplets (
The distribution of athletes in the test by gender and age is shown in Table 4. U8 is for 6- and 7-year-old children, U10 is for 8- and 9-year-old children, and so on.
Data of the athletes in the survey.
Before choosing a sport discipline
In this first scenario, we conducted tests on athletes who had not been assigned to a specific discipline. Having asked the coaches to indicate which discipline they believed the athletes should be placed in, we compared their responses with the results obtained from our method. A part of these results is contained in the conference paper. 35
A total of 18 athletes participated in this test, exhibiting significant diversity both in terms of age and performance level. It is necessary to emphasize this, as a 12-year-old may complete a 60-meter race much faster than a 6-year-old. Therefore, we had to create different scales to represent the strength of their results. To accommodate for this, we divided the participants into two groups based on their performance (see Table 5).
Results of those awaiting classification into a sports discipline.
Out of the five older athletes, we correctly identified four, and out of the 13 younger athletes 12. Correctly in this context means that the decision of the algorithm coincides with the decision of the trainer. The exceptions were the followings: in one case, the orientation concerned a 6-year-old child. Orientation is particularly challenging and not easily achievable at such a young age. Moreover, this child’s scores are very low in every sport, and there is only a slight difference (0.01) between the scores belonging to throwing and running. Moreover, his/her orientation would be very early. In the case of the other athlete, the given scores in the various disciplines are low compared to their age group, so this person is at the beginning of development in the corresponding discipline.
Sometimes there are such cases (see e.g. athlete 10), where it is very difficult to determine the orientation of certain athletes. The athlete may have similar strength in each discipline, and the weights given by the model for each discipline may differ only slightly. The coaches are then in a similarly difficult position, finding it hard or impossible to say which discipline would be best for the athlete.
In summary, out of the 18 cases, we made recommendations that aligned with the coaches’ opinions in 16 cases, resulting in an 89% agreement rate, which can be considered sufficiently good. The method reflects well the opinion of the trainers. We believe that this algorithm is able to make numerous early recommendations based on their NETFIT survey results without the personal contacts between the trainers and the person.
Measuring the performance in the sport disciplines
Another way of checking the reliability of the method is the investigation of the relationship between the assigned scores (overall strength) for the athletes categorized before and their official rankings. If these rankings correlate well, then the overall strength properly measures the strength of the athlete in the given sports discipline.
We refer to our ranking as the “NETFIT ranking” from now on. In the case of male athletes (see the top of Table 6), we had two participants (long jump), and their rankings in both systems matched. Furthermore, we observed that the athlete we rated significantly higher in our system indeed achieved much better results compared to the other athletes in their respective discipline.
Results of male and female competitors.
In the case of female athletes (see at the bottom of Table 6), the two rankings were very similar to each other (100 m race). There was only one instance where a swap occurred between the third and fourth positions. In this case, the scores we gave were close to each other. To be more convincing, we will evaluate further athletes in the future.
We also attempted to test the validity of our method more simply, by also carrying out these tests on those who are classified in a specific category, but who have not had the competition results yet, so it is difficult or impossible to compare them. Our method offers the possibility to compare their performance or growth through their scores.
For this measurement, we obtained results from the jumping/sprinting discipline, so we present the results obtained there (see Table 7). As there were larger age differences and they have been training and developing with the help of the sports organization for some time, we applied more coaching min–max scales for U14 (12- and 13-year-old girls and boys), U16 (14- and 15-year-old girls and boys), and U18 (16- and 17-year-old girls and boys). In the light of the results obtained, we can see that in two cases the method was not flawed, one for an injury and the other for an athlete who excels in several disciplines. Overall, it can be said that the method performed as expected. In 12 out of 13 cases, it was able to recommend the same discipline as the coaches (if we do not count the injury, which is why sudden loading is not recommended for the athlete, which is important in his discipline).
Results of athletes classified in the disciplines.
Direction for the performance improvement of athletes
As it was mentioned previously, this method allows us to identify the strengths of each athlete in a given sport and to identify an area where improvement can best influence their subsequent performance as measured by our method and therefore in competitions. In this subsection, we will show how this can be calculated (see Figure 5) and we illustrate it with an example (see Table 8).

Calculation of the attributes lack.
Example of calculation of attributes lack.
The athlete has his or her assessment results and the points assigned to them. From the maximum of 1 point, we subtract his or her score for that attribute to get the maximum improvement in that area. This value is multiplied by the weight previously given to the trait, and the attribute with the highest multiplication is the attribute where the athlete should improve to achieve an even better performance in that particular discipline (see Figure 5).
In the following section, the calculation and the result of one athlete is presented (see Table 8). This athlete is a sprinter, so he belongs to the jumping/sprinting discipline. From Table 8, it is clear that in this discipline for this athlete, the greatest improvement is achieved in the 60 m with a standing start or in the standing long jump test. In other cases, it can be seen that the fitness level of the person assessed is adequate or, if there is a deficit, the weight of this deficit does not have a significant negative impact on his performance in his discipline. Therefore, the optimal training direction can be identified and designated by the elaborated method.
In the remainder of this subsection, we present how some additional information can be extracted from the collected data using the method together with additional coaching information. It is paramount to rely on accurate coaching information, without which the high success rate of the method cannot be ensured. The results of paired comparisons and weights have been presented earlier; conversely, the scales that have been defined by experts follow. Scales for the U16 age group (female and male aged 14 and 15) were presented earlier (see Table 1), as these will be used to calculate the scales of the athletes whose progress was followed by the software.
At this age, boys are still gaining a lot of strength, while girls are in the last stages of gaining strength and their growth decelerates. We have followed the growth of three females (Table 9) and three males (Table 10) over eight months. Each athlete’s discipline is jumping/sprinting. What is promising to see from the tables is that everyone has improved significantly in their discipline, especially the boys demonstrated a considerable improvement, and this is in line with their growth. Additionally, it is also worth mentioning that everyone has the highest score in their discipline thanks to the training, except one boy (male I), who does not. Having identified that in this case an injury transpired, short heavy workloads are not recommended for him, but his endurance remains at a similar level. Everyone was able to essentially maintain or even improve their 60 m with the standing start test, as well as their performance in the standing long jump, which are important in their discipline.
Results of old and new female surveys.
Results of old and new male surveys.
What is also easy to ascertain is that the boys’ back-saver sit and reach test results showed a significant improvement, which shows the presence of effective coaching, because hard competitive training can cause muscles to become shortened and less flexible. As the chart demonstrably illustrates for six of the six subjects with the 20 m pacer test, a weakening can be detected in five of them. This test measures stamina, so it is clear that this trait is diminishing as in their discipline this trait is not needed as much as high sudden force (e.g. 100 m flat running).
There is a strong correlation between competition results and the scores, so an improvement in score indicates a likely improvement in competition results. Consequently, NETFIT results can be useful for determining the direction of improvement for individual athletes, ultimately leading to better competition results. This is particularly true when dealing with multi-faceted events that require multiple strengths, such as decathlons or other events with diverse skill requirements.
Future improvements
To increase the accuracy of the method even more, we can proceed in two directions. One way is to increase the accuracy and reliability of the weight of each survey. In the other direction, we need to refine and improve the scores from the survey results.
As usually in statistics, the higher the sample size, the more accurate the results are. The accuracy and reliability of the method can be further increased by involving more coaches to complete the comparison matrix of the attributes (survey numbers) for each discipline. This might increase the accuracy of the weights of the attributes. With many coaches and lots of data, a breakdown by discipline should also be considered to determine more accurate specific weights and strengths.
The accuracy of the method for determining the athletes’ strengths can be further enhanced by using a function that is more complex than a linear function within the minimum and maximum values of the scale we applied. This function can be, for example, the sigmoid function. We present such a sigmoid function that we can use in our work and hopefully use to more accurately determine the discipline orientation and strength of each athlete. For our sigmoid function, we also need min and max values below or above which we consider the value of the function to be 0 or 1. Because of the complexity of the formula, we introduce some notations to facilitate the calculation:
These make the function looks like this on the interval [min, max]:
In addition, significant further improvements might be achieved by taking into account the opinions of several coaches for each age–gender in the min–max scale, so that the scores obtained for each athlete are more accurate. Sensitivity analysis might be performed to see the extent of the changes in the function of the min–max values of the scales.
The amount of data that can be obtained from the method and its justification can be further increased by involving more athletes in the surveys and evaluations.
Discussion
Many readers may ask why such a complicated system is needed if coaches can decide together on the fate and development of individual athletes. The use of this method can be justified by the desire of individual sports clubs to become more accessible to more people, to accept more athletes for training, to train with a more diverse range of athletes. To do this, it is necessary to assess the performance of a large number of young athletes, and such an automated decision system can help coaches to do this. In addition, the quantification of sporting strengths by this method provides sports administrators with a more objective control than relying solely on the opinions of individual coaches. Moreover, it is possible to store the individual opinions and professional knowledge of each coach, which can be used in the future when new coaches are recruited, thus helping to maintain quality. It also makes it possible to assess the effectiveness of coaching by taking the results into account. The validity of the method is confirmed by the fact that it can provide a fairly accurate and quantified reflection of coaches’ opinions and insights, even with very little data. It can consequently be of great help in selection and monitoring progress. We certainly have to admit that good coaches are still needed to provide the training and develop the athletes, and their knowledge, without that the application would not be able to correctly help the development and selection of athletes. The method is flexible, the results reflect the local peculiarity, as the opinion of the trainers are included in the weight, but the weights are balanced, due to the evaluation method.
Conclusions
Over the course of our research, we successfully developed a method which is suitable for making complex decisions if multiple factors are involved. We created a method that ranks these factors, assigns weights to them, and evaluates their results to assess the “goodness” of complex decisions.
We have developed a software which is capable of quantifiably measuring the athletes’ performance in their respective disciplines. Not only did this method prove to be useful in determining the appropriate discipline for individual athletes but also for ranking the athletes within their disciplines without relying solely on the competition results. This can greatly assist coaches in quantifiably assessing the progress of their athletes and identifying the specific areas of development that can contribute to better rankings.
The developed method can be applied to other situations as well, such as assisting in career choices for students or evaluating various tenders, and providing decision support in diverse scenarios.
Footnotes
Acknowledgements
The research was partly supported by the ÚNKP-22-3 New National Excellence Program of the Ministry for Culture and Innovation from the source of the National Research, Development and Innovation Fund. L Gyarmati thanks the support. The authors express their gratitude to the coaches who assisted in data collection.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship and/or publication of this article.
