Abstract
Since test performance is increasingly relevant in educational and occupational circles, the assessment of test anxiety—the phenomenological, physiological, and behavioral responses to the negative consequences that often emerge in evaluative situations—has become increasingly important to scholars and practitioners. One of the most widely employed scales to measure test anxiety in adolescents is the German Test Anxiety Inventory (in German: Prufungsangstfragebogen, PAF). The current study investigated the psychometric properties of the PAF when administered to Italian students. Our research found evidence of validity, supported the five-factor structure, and demonstrated the test’s good internal consistency. Moreover, the invariance of the dimensional structure across genders was examined. Overall, this study provides evidence for the reliability and validity of the PAF among Italian students.
Introduction
Test anxiety, which is also known as evaluative anxiety or performance anxiety, “refers to the set of phenomenological, physiological and behavioural responses that accompany concern about possible negative consequences or failure of an exam or a similar evaluative situation” (Sieber, O’Neil, & Tobias, 1977, p. 23). Although test anxiety has been treated as one of the specifiers of a generalized anxiety disorder in the fifth edition of the Diagnostic and Statistical Manual of Mental Disorders (American Psychiatric Association, 2013), several researchers (Blöte, Kint, Miers, & Westenberg, 2009; Hofmann, Heinrichs, & Moscovitch, 2004; Hook & Valentiner, 2002) have claimed that it is qualitatively distinct from other types of anxiety disorders. Indeed, this kind of anxiety seems to have a phobic quality, greater than the features of general anxiety or other types of social anxiety disorders.
Test anxiety is a situation-specific anxiogenic response that derives from trait anxiety and affects evaluative situations (Spielberger & Vagg, 1995; Zeidner, 2007). Test anxious people perceive performance evaluations as threatening in nature, which in turn leads them to respond with heightened state anxiety (Ringeisen, Raufelder, Schnell, & Rohrmann, 2016). Considering that tests and other performance indicators are frequently used in both educational and occupational contexts (Roediger & Karpicke, 2006), test anxiety is seen as a serious issue among scholars and practitioners alike. Indeed, test anxiety is a relatively widespread condition, as research has found that between 14% and 50% of college students (Cizek & Burg, 2006), and more than 33% of school-age children and adolescents, experience some degree of test anxiety (Methia, 2004).
Gender difference is an important factor to consider, as female students across all age groups report higher levels of test anxiety than male students (Chapell et al., 2005; Goetz, Bieg, Lüdtke, Pekrun, & Hall, 2013; Hodapp, Rohrmann, & Ringeisen, 2011; A. S. McDonald, 2001; Ringeisen et al., 2016; Schnell, Tibubos, Rohrmann, & Hodapp, 2013; Wren & Benson, 2004; Zeidner, 2007). For instance, in their work on noncognitive predictors of grades among freshman students, Fischer, Schult, and Hell (2013) found that females showed significantly higher test anxiety than males, and this difference displayed an effect size of .36. Chin, Williams, Taylor, and Harvey (2017) similarly found that female high school students showed higher levels of test anxiety than boys. Finally, even in a sample involving younger children (average age of 10 years), girls reported higher levels of test anxiety than boys (Lohbeck, Nitkowski, & Petermann, 2016); this difference had an effect size of .32.
Test anxiety can negatively affect education, as test-anxious students “do not approach a task such as a test with a positive outlook or expectation of success, but with dread regarding the potential for negative evaluation or failure” (Cizek & Burg, 2006, p. 17), and often feel tense and worried in evaluative situations (Gierl & Rogers, 1996). For this reason, it is very likely that students with high levels of test anxiety exhibit diminished or poor performance (Cassady & Johnson, 2002; Chapell et al., 2005; Segool, Carlson, Goforth, von der Embse, & Barterian, 2013) and, in some cases, school attrition (Cizek & Burg, 2006). Several models have focused on the impact of test anxiety in performance evaluations. For instance, according to supporters of control-value theory (Pekrun, 2006; Pekrun, Frenzel, Goetz, & Perry, 2007), low dispositional control beliefs in performance situations and high personal relevance for academic success increase the anticipated risk of failure via test anxiety. Meanwhile, proponents of self-regulation theory (self-regulation theory (SRT); Schwarzer, 2001) postulate that low levels of self-efficacy can negatively affect performance through goal setting and goal significance, initiation and perseverance of learning, and test anxiety. Furthermore, research based on the tripartite model of emotions (TME; Chin et al., 2017) suggests that a general tendency to experience negative affect and physiological hyperarousal often predicts higher levels of test anxiety in adolescents, while positive affect did not influence test anxiety. Moreover, TME shows that negative emotions have a detrimental effect on educational outcomes through test anxiety. Thus, test anxiety occupies a central role in mediating the interplay between dispositional characteristics and performance (Schnell, Ringeisen, Raufelder, & Rohrmann, 2015).
Obviously, in order to more precisely understand the learning and performance capacities of students and screen youths with high levels of test anxiety, it is necessary to have access to valid and reliable measures. For this reason, different measures have been developed to assess test anxiety among high school and college students. In line with a two-dimensional conceptualization of test anxiety into a cognitive and an emotional declination (Liebert & Morris, 1967; Schwarzer, 1984; Wacker, Jaunzeme, & Jaksztat, 2008; Ware, Galassi, & Dew, 1990), Spielberger (1980) developed the Test Anxiety Inventory (TAI), which regards test anxiety as a construct defined by a cognitive component (worry—thoughts of failure regarding performance) and an affective component (emotionality—autonomous excitation and emotional reaction toward the evaluative moment). Subsequently, Sarason (1984) developed a new version of the TAI—the Reactions to Test Scale—composed of four dimensions that split the emotionality component into test and bodily symptoms and the cognitive component into worry and test-irrelevant thinking. In order to better investigate the different facets of test anxiety, Hodapp (1991) then revised this measure, creating a new German version of the scale—the TAI-G—for secondary school and college students. This scale was composed of 30 items and four subscales: Worry, which assesses concerns about individual performance and the consequences of failure; Emotionality, which examines emotional and physical tension; Interference, which measures distraction from the task via irrelevant thoughts; and Lack of Confidence, which evaluates the impact of confidence levels on mastering academic challenges. This new version was eventually shortened and revised, resulting in an easy-to-administer measure—the German Test Anxiety Inventory (in German: Prufungsangstfragebogen, PAF; Hodapp et al., 2011). It consists of 20 items, with five items per subscale. This test has been used predominantly in German studies that investigate test anxiety among adolescents (Hodapp et al., 2011; Ringeisen et al., 2016; Schnell et al., 2013) and young adults (Fischer et al., 2013; Reiss et al., 2017). Several other studies (Ringeisen et al., 2016; Schnell et al., 2013) have also found it to be a reliable and valid measure, and an English version for 11- to 16-year-old students has recently been published (PAF-E; Hoferichter, Raufelder, Ringeisen, Rohrmann, & Bukoski, 2016).
Internal consistency of the PAF has been found to be satisfactory, with Cronbach’s alphas that range from .71 to .85 (Hoferichter et al., 2016; Schnell et al., 2013). In terms of construct validity, an exploratory factor analysis and a confirmatory factor analysis (CFA) supported the four-factor model: Worry explained 26.51% of the total variance, Emotionality 17.51%, Interference 7.63%, and Lack of Confidence 5.60%. Regarding criterion validity, Hoferichter et al. (2016) found that the four scores associated with the PAF subscales were all significantly and positively correlated with educational self-efficacy, with the highest coefficient value found in the Lack of Confidence score. Again, scores pertaining to the Worry and Lack of Confidence subscales were significantly and negatively correlated with performance motivation, and significant and negative correlations between the PAF subscales and grades were found, particularly math grades. Consistent with these findings, Schnell et al.’s (2013) research on high school students found that the Interference and Lack of Confidence subscale scores were significantly negatively correlated with both math and language performance. Additionally, they found that all of the PAF subscale scores, with the exception of Worry, were significantly positively related to math anxiety. PAF subscores have been also found to be correlated with personality variables, including conscientiousness, neuroticism, and agreeableness. As for convergent validity, the four subscores associated with the PAF positively correlated with the inhibitory test anxiety measure (Hoferichter et al., 2016), with correlation coefficients of .59, .51, .30, and .18 with Worry, Emotionality, Interference, and Lack of Confidence subscale scores, respectively. Finally, as support for incremental validity, the total score of the PAF was a significant predictor, together with gender, of overall school performance in a regression model that also included gender and math anxiety (Schnell et al., 2013).
Despite its usefulness, few psychometric studies of the PAF have been published (Hoferichter et al., 2016; Schnell et al., 2013), and psychometric data remain underinvestigated. For instance, CFA of the factorial structure has been tested in only one study (Hoferichter et al., 2016), and measurement invariance—the ability of a test to uniformly measure a specific construct among different groups of respondents—has not been assessed. Measurement invariance represents a fundamental property of a test; if a test does not uniformly measure a construct among different groups of respondents, the comparison of test scores between different groups obviously suffers (Waiyavutti, Johnson, & Deary, 2012). By contrast, invariance testing ensures the fairness and thus the validity of a scale (Kane, 2013). Since gender-related differences in test anxiety have been found (Fischer et al., 2013; Hodapp et al., 2011; Ringeisen et al., 2016; Schnell et al., 2013), and the PAF has been used to examine gender invariance between various test anxiety predictors among adolescents (Ringeisen et al., 2016), it becomes necessary to verify measurement invariance across genders. Also, since test anxiety has been found to be an obstacle to successful performance among both preadolescents and adolescents (Chin, Williams, Taylor, & Harvey, 2017; Lohbeck et al., 2016), we deemed it necessary to analyze the psychometric properties of the PAF among youth who were between 11 and 16 years of age, which is in line with the English validation study (Hoferichter et al., 2016).
Following these premises, this study analyzed the psychometric properties of the PAF in an Italian context and tested the adequacy of the four-factor model (Hodapp, Rohrmann, & Ringeisen, 2011; Hoferichter et al., 2016). However, since Hoferichter et al. (2016) have shown that the covariance between Worry and Lack of Confidence (r = .02) was not significant, and the coefficient value of the covariation between Worry and Interference was .08, we hypothesized that both Worry and Lack of Confidence and Worry and Interference were unrelated dimensions. We were also interested in studying gender invariance via a multigroup confirmatory analysis in order to determine whether the detected factors were replicable in boys and girls.
We also examined the internal consistency of the subscales by calculating Cronbach’s alpha (α) coefficients (Cronbach & Shavelson, 2004), all of which have been found to be satisfactory in previous studies (Hoferichter et al., 2016; Schnell et al., 2013). Additionally, we analyzed the internal consistency by calculating R. P. McDonald’s (1999) omega (Ω), which corresponds to the ratio of true-score variance and total variance, calculated using the item’s coefficient on each factor.
We evaluated the criterion validity of the PAF by examining the correlation between test anxiety and math anxiety and by analyzing the relationship between the total and subscale scores at the Abbreviated Math Anxiety Scale (AMAS, Hopko, Mahadevan, Bare, & Hunt, 2003; Italian version: Primi, Busdraghi, Tomasetto, Morsanyi, & Chiesi, 2014). Test and math anxiety are increasingly recognized as two different, yet partially overlapping, constructs (Schnell et al., 2013). In fact, previous studies have found moderate to strong correlations between test anxiety and math anxiety among high school and college students (Ashcraft, 2002; Dew, Galassi, & Galassi, 1984; Morsanyi, Busdraghi, & Primi, 2014; Primi et al., 2014; Schnell et al., 2013). Additionally, we analyzed the relationship between math performance and test anxiety in order to find a significant negative association with Interference and Lack of Confidence (Schnell et al., 2013). To assess math performance, we used the AC-MT 11–14 standardized mathematics test (Cornoldi & Cazzola, 2004), which is designed for Italian middle school students and measures calculation procedures through written multidigit calculations. Finally, we investigated gender differences in the PAF by emulating previous studies that have found substantially higher levels of test anxiety among female students than male students (Chapell et al., 2005; Chin et al., 2017; Fischer et al., 2013; Goetz et al., 2013; Hodapp et al., 2011; Lohbeck et al., 2016; A. S. McDonald, 2001; Ringeisen et al., 2016; Schnell et al., 2013; Wren & Benson, 2004; Zeidner, 2007).
Methods
Participants
Three hundred and twenty-six youths participated in our study (mean age: 13.36 years, standard deviation (SD): 1.34, range: 11–16 years; 146 males: mean age: 13.34 years, SD: 1.37, range: 11–16 years; 180 females: mean age: 13.38 years, SD: 0.32, range: 11–16 years). All participants attended public secondary schools that were located in both the North-Centre (43%) and South (57%) of Italy. Four schools (two middle schools and two high schools) were randomly selected from each region. The principals were then contacted, and the researchers explained the goals of the study. Once the schools agreed to participate, a detailed study protocol, designed in accordance with the criteria of the Declaration of Helsinki, was approved by the institutional review boards at each school. Two schools—one middle school in the North-Centre of Italy and one high school in the South of Italy—declined to participate because they were already involved in other projects. Written informed consent was requested from the students’ parents, and they were assured that the data would be handled confidentially. The research was conducted during school hours and all students who were invited to participate agreed to do so, which means that sample bias that might have emerged during the recruitment stage was avoided.
Measures
The PAF (Hodapp et al., 2011) measures self-reported anxious emotions and thoughts during test situations. It consists of 20 items with a four-point response scale that ranges from 1 (hardly ever) to 4 (almost always). The four subscales are organized as follows: Lack of Confidence (L) consists of five reversed items (e.g., I trust in my performance), Worry (W) is composed of five items (e.g., I think about how important the test is to me), Emotionality (E) is assessed through five items (e.g., I feel anxious), and Interference (I) includes five items (e.g., Suddenly thoughts cross my mind which inhibit me). The Italian version of the PAF was created by adapting the English version of each item. A forward-translation method was used, whereby two nonprofessional translators worked independently and then compared their translations. A group of five people then read a provisional version of the new scale, revised it, and settled on a final version. In line with previous studies (Reiss et al., 2017; Schnell et al., 2013; Tempel & Neumann, 2016), also a total PAF score was calculated as a measure of test anxiety.
The AMAS (Hopko et al., 2003; Italian version: Primi et al., 2014) is a two-factor measure of math anxiety that assesses Learning Math Anxiety (LMA; five items) and Math Evaluation Anxiety (EMA; four items). A five-point response scale is used, ranging from 1 (low anxiety) to 5 (high anxiety). A sample item is Thinking about an upcoming math test one day before. High scores indicate high levels of math anxiety. In line with previous studies (Cipora, Szczygiel, Willmes, & Nuerk, 2015; Morsanyi et al., 2014; Primi et al., 2014), the two subscale scores and a single composite score were calculated based on how participants rated each statement. Previous studies have found support for its validity and reliability. Concurrent validity was supported by showing good associations between the AMAS subscales and total score, test anxiety, and math attitude (Primi et al., 2014). Convergent and divergent validity were provided in the original study of the AMAS by stressing the strong associations with the Math Anxiety Rating Scale-Revised (Plake & Parker, 1982) and moderate associations with other anxiety measures, including general anxiety, state and trait anxiety, and test anxiety (Hopko et al., 2003). Good internal consistency values were found in previous studies (Hopko et al., 2003; Primi et al., 2014), and the test–retest reliability was found to be excellent (Hopko et al., 2003). Cronbach’s alpha coefficients in the present study were .76 for the LMA subscale, .75 for the EMA subscale, and .79 for the total scale.
In order to assess math performance, the AC-MT 11–14 standardized mathematics test (Cornoldi & Cazzola, 2004) was used. This test assesses calculation procedures and number comprehension by means of a set of paper-and-pencil tasks that can be grouped into two areas: “written calculation” and “number knowledge.” Our research used calculation tasks whereby participants were asked to solve eight written multidigit calculations, including two additions, two subtractions, two multiplications, and two divisions. Total scores were obtained by adding up correct responses. Good reliability of this measure was found in previous studies both in terms of test–retest stability (Caviola, Primi, Chiesi, & Mammarella, 2017) and internal consistency (Cornoldi & Cazzola, 2004). Moreover, adequate concurrent validity (Cornoldi & Cazzola, 2004) and discriminant validity (Passolunghi, Caviola, De Agostini, Perin, & Mammarella, 2016) were confirmed.
Procedure
The measures were administered individually during class time. All participants completed the PAF. In order to analyze criterion validity, a subsample of 138 students (68 males and 70 females, mean age: 13.06 years, SD: 0.62) also completed the AMAS and the AC-MT. The order of administration of the scales was randomized. Participants were provided with a brief introduction to the study and instructions on how to proceed. Answers were collected in a paper-and-pencil format. Data collection was completed in approximately 20 minutes for the PAF, and 40 minutes for the PAF, AMAS, and AC-MT combined.
Statistical analyses
A missing value analysis was performed during the early stages of the data analysis phase. Missing data were not allowed to exceed 10% of the total cases in the sample (Kline, 1998). Indeed, when missing data exceeded 10%, we decided to exclude the case. In total, two cases were excluded, which means that the data of 324 participants were ultimately analyzed. When missing data were deemed acceptable, the score of each omitted item was treated as being equal to the mean of its subscale.
Univariate distributions of PAF items were then examined in order to assess normality, placing special emphasis on skewness and kurtosis indices. Subsequently, the four-factor structure was tested by a CFA that employed the maximum likelihood estimator (AMOS software; Arbuckle, 2003). To verify the model’s fit, the following indices were taken into account: The ratio of chi-square to its degrees of freedom (χ2/df), the comparative fit index (CFI; Bentler, 1990), the Tucker–Lewis index (TLI; Tucker & Lewis, 1973), and the root mean square error of approximation (RMSEA; Steiger & Lind, 1980). In the case of χ2/df, values below or equal to 2 are considered good, while values between 2 and 3 are considered acceptable (Schermelleh-Engel, Moosbrugger, & Müller, 2003). For the TLI and CFI, values above .90 are indicative of acceptable fit, while values above .95 are indicative of excellent fit (Hu & Bentler, 1999). The RMSEA value is considered acceptable when it is below .08 and good when it is below .05 (Kline, 2010).
Gender invariance analyses were conducted by performing hierarchically nested CFAs (see Byrne, 2004, for testing multigroup invariance with AMOS). Invariance was evaluated using not only Δχ2, which is sensitive to sample size, but also ΔCFI, which has been found to be the most sensitive index to detect a lack of invariance (Meade, Johnson, & Braddy, 2008). An absolute value of ΔCFI of less than .01 was used (Cheung & Rensvold, 2002; Dimitrov, 2010).
In order to assess the reliability of the PAF, Cronbach’s alphas and McDonald’s omega values were obtained for the various subscales. To examine validity, correlation coefficients (Pearson’s r) were calculated with the AMAS subscale/total score and the AC-MT score. After testing for gender invariance, we compared the PAF subscale scores of male and female students through an independent samples t test, adopting a hypothesis stipulating that gender difference was evidence of discriminant validity.
Results
Dimensionality
We looked at the PAF items to assess normality. Skewness and kurtosis indices were between −1 and +1 for all but three items, which were slightly outside the range of normality (Table 1). However, this type of deviation is often considered to be negligible (Ghasemi & Zahediasl, 2012).
Means, SDs, skewness, and kurtosis of the 20 items of the PAF.
Note: The PAF uses the following Likert-type scale: 1 = hardly ever, 2 = sometimes, 3 = often, and 4 = nearly always. n = 324. SD: standard deviation.
A CFA was carried out by testing the PAF’s four-factor model. The results showed that the fit indices were not acceptable. Modification indices suggested adding error covariance between items 1 and 4 of the Lack of Confidence subscale, items 7 and 13 of the Worry subscale, items 8 and 10 of the Emotionality subscale, and items 3 and 15 of the Interference subscale. These links were theoretically justified because linked items reported similar contents (Byrne, 2004). The modified model showed a good fit (χ2 = 304.47, df = 162, p < .001, χ2/df = 1.88, CFI = .93, RMSEA = .05, 90% confidence interval (CI) = .04–.06). Each item loaded strongly and significantly on its hypothesized factor, and the hypothesized correlations between the factors were all significant ranging from .32 to .50 (Figure 1).

The four-factor model of the PAF in the global sample (n = 324). Standardized parameters were all significant at .001. L: lack of confidence; W: worry; E: emotionality; I = interference; PAF: German Test Anxiety Inventory.
Gender invariance
In order to test gender invariance according to recommended standards (Dimitrov, 2010; Little, 1997; Vandenberg & Lance, 2000), the independence model was fitted (χ2 = 2493.73, df = 380, p < .001). As Table 2 illustrates, weak or metric factorial invariance and configural invariance was supported, thus confirming that the factor loadings were equal across genders. The factors’ variances and covariances were also found to be invariant. Finally, the equality of the items’ variances and covariances was also tested. We obtained a nonsignificant Δχ2, and the differences in CFI values were less than .01. Thus, the results supported gender invariance.
Goodness-of-fit statistics for each level of structural and measurement invariance across genders.
Note: χ2: chi-square test; df: degrees of freedom; CFI: robust comparative fit index; RMSEA: robust root mean square error of approximation; Δχ2: Satorra–Bentler scaled difference; Δdf: difference in degrees of freedom between nested models; p: probability value of Δχ2 test; ΔCFI: difference between robust CFIs of nested models. n = 324.
Reliability
Cronbach’s alphas showed acceptable internal consistency, producing values equal to .83 (95% CI = .80–.86) for the Lack of Confidence subscale, .72 (95% CI = .69–.77) for the Worry subscale, .79 (95% CI [.76, .83]) for the Emotionality subscale, and .81 (95% CI = .77–.84) for the Interference subscale. These values did not increase when an item was deleted, and all item-corrected total correlations were above .44. Following the standards proposed by the European Federation of Psychologists’ Association (EFPA; Evers et al., 2013), the internal consistency value was good for the Worry and Emotionality subscales and excellent for the Lack of Confidence and Interference subscales.
McDonald’s Ω scores were found to be .86 for the Lack of Confidence subscale, .75 for the Worry subscale, .85 for the Emotionality subscale, and .80 for the Interference subscale. According to standards proposed by the EFPA (Evers et al., 2013), the internal consistency value was good for the Lack of Confidence, Emotionality, and Interference subscales and adequate for the Worry subscale.
Validity
To test the validity, the relationship between the PAF subscales and total score, the AMAS subscales and total score, and the AC-MT total score was assessed (Table 3). The results showed that all of the PAF subscales and the PAF total score were significantly and positively correlated to both AMAS subscales and its total score. Additionally, significant negative relationships were found between math performance, the PAF total score, and the Lack of Confidence and Interference subscale scores. The values of the correlations were either adequate (ranging from .20 to .34), good (ranging from .35 to .49), or excellent (≥.50), all of which support the criterion validity of the PAF, according to EFPA standards (Evers et al., 2013).
Pearson correlations between the PAF subscale scores, the AMAS subscale scores and total score, and the AC-MT score.
Note: L: lack of confidence; W: worry; E: emotionality; I: interference; PAF: German Test Anxiety Inventory; LMA: Learning Math Anxiety; EMA: Math Evaluation Anxiety; AMAS: Abbreviated Math Anxiety Scale; AC-MT: AC-MT 11–14 standardized mathematics test. n = 138; SD: standard deviation.
**p < .01; ***p < .001.
Gender differences
After having verified the measurement equivalence of the scale among both male and female respondents, a comparison by gender was conducted on the PAF subscale scores. No significant differences were found with respect to the Lack of Confidence (t(322) = −.06, p = .950, Cohen’s d = .01) and the Interference (t(322) = 1.57, p = .118, Cohen’s d = .18) subscales. Gender differences were found in the Emotionality subscale (t(322) = −4.14, p < .001), as female students (M: 10.25; SD: 3.67) displayed higher levels of anxiety than male students (M: 8.66; SD: 3.13). This difference had a moderate effect size (Cohen’s d = .47). Additionally, a small (Cohen’s d = .23) difference was found in the Worry dimension (t(322) = −2.10, p = .036), as girls (M: 14.62; SD: 3.12) displayed significantly higher scores than boys (M: 13.88; SD: 3.20).
Discussion
Performance anxiety has been recognized as a variable that can negatively affect not only isolated performance but also broader educational processes (Cassady & Johnson, 2002; Chapell et al., 2005; Cizek & Burg, 2006; Segool et al., 2013). Thus, measuring test anxiety in students should be seen as an important issue. For this reason, this study investigated the psychometric properties of the Italian translation of the PAF (Hodapp et al., 2011), a self-report scale that measures anxious emotions and thoughts during test situations, treating lack of confidence, worry, emotionality, and interference as key factors.
Our research suggests that the Italian version of the PAF displays characteristics comparable to those reported in previous studies involving adolescents (Hoferichter et al., 2016; Schnell et al., 2013). Our factor analysis, in particular, supported the four-factor structure of the scale, and good internal consistency was found for the various subscales. Whereas previous studies have not taken into account measurement invariance, this study found that the four-factor structure was invariant across genders. This finding legitimates the use of the PAF to compare male and female students and analyze gender-related differences in test anxiety and relative predictors.
We also found evidence of the validity of the PAF in the association between test anxiety, math anxiety (Ashcraft, 2002; Dew et al., 1984; Morsanyi et al., 2014; Primi et al., 2014; Schnell et al., 2013), and math performance (Hoferichter et al., 2016; Schnell et al., 2013). Additionally, we found support for gender-related differences in test anxiety, as our data suggest that female students were more anxious toward evaluative situations than were male students, a trend that was especially noticeable in relation to the Emotionality and Worry dimensions.
The PAF can help practitioners identify at-risk male and female adolescents who suffer from high levels of test anxiety and plan appropriate educational interventions and/or therapeutic treatments (see Ergene, 2003, for a meta-analysis on effective interventions to reduce test anxiety). Using previous research by Reiss et al. (2017) as a guide, we concluded that different strategies (i.e., relaxation techniques and imagery rescripting) should be employed in order to reduce test anxiety in students. For instance, reinforcing learning and motivation, investigating dysfunctional beliefs about performance and restructuring them into adaptive thoughts, teaching functional coping strategies, and providing training on how to identify and overcome blackout situations during testing situations could be done with the aim of reducing test anxiety. These activities could be conducted not only in a therapeutic setting but also in a classroom setting by psychologists and counselors. Another potential use of the PAF can be in the field of experimental studies in which there is the necessity of measuring test anxiety to select high-anxious individuals (Tempel & Neumann, 2016). Additionally, assessment of test anxiety is necessary to measure the effects of different type of didactic materials on test anxiety (e.g., Rockinson-Szapkiw, Wendt, & Lunde, 2013) and to evaluate the consequences of test anxiety on learning under different kinds of social conditions (e.g., Cassady, 2004). Finally, the PAF should be used as a self-report measure to control for test anxiety in studies that aim to better understand the effects of math anxiety on math performance and reasoning (Morsanyi et al., 2014).
Some limitations characterize this study. For instance, the participants were all secondary school students, representing a relatively homogeneous population in terms of age. Future studies with a more heterogeneous sample could be useful in detecting measurement invariance. Indeed, it might be useful to compare participants from different age groups who are at different stages of the educational process. Further analyses are also needed to investigate cross-cultural invariance of the instrument. Another limitation is that we used only math anxiety and math performance as criterion validity measures. A broader range of measures are needed in future investigations of validity, including constructs that are correlated with test anxiety—most notably performance motivation, self-efficacy, school grades, and personality dimensions such as conscientiousness, neuroticism, and agreeableness (Hoferichter et al., 2016). Despite these limitations, the PAF maintains good psychometric properties when used with Italian students, which means that it could be a useful instrument for measuring test anxiety in adolescents.
