Abstract
Rasch analysis was conducted to enhance the precision of the widely used 10-item Perceived Stress Scale using two datasets (n = 450 each) randomly selected from samples of the New Zealand general population (n = 1102), New Zealand university students (n = 479) and US university students (n = 396). The best Rasch model fit (χ2(27) = 29.92, p = .36), good person separation reliability (.80) and coverage (98%) of the sample by the scale items were achieved when locally dependent items were combined into subtests. These findings support reliability and internal structural validity of the 10-item Perceived Stress Scale. The instrument precision can be further improved using the ordinal-to-linear conversion tables published here.
Introduction
Stress is described as heightened emotional states associated with physiological changes (Helton and Näswall, 2015; McEwen and Stellar, 1993). Research has shown that extended exposure to stress can lead to aversive health effects (Cohen et al., 1998; Hillhouse et al., 1991). Both effective stress management and stress reduction are critical in reducing negative stress effects on individual’s health. Development of effective methods to manage and reduce stress requires accurate assessment of perceived stress levels to evaluate and compare unique contributions of various predictors, situations and behaviours that may trigger and maintain stress. In particular, precise assessment of perceived stress is critical because it reflects the subjective evaluation of environmental events (Bloch et al., 2004), which directly influences physiological responses responsible for aversive health effects (LeDoux, 2000; Medvedev et al., 2015).
The Perceived Stress Scale (PSS) (Cohen and Williamson, 1988) was constructed as a subjective measure of perceived stress to assess the extent to which a person’s life is perceived as ‘unpredictable, uncontrollable, overloading’ (p. 387), relative to the individual’s coping abilities. The scale is very widely used, exceeding 11,000 citations by the middle of 2016, according to Google Scholar. The PSS has been cross-culturally validated and translated into 25 different languages (Cohen, 2013). The original PSS version contains 14 items and has good internal consistency (Cronbach’s alpha > .80) and satisfactory construct validity (Cohen and Williamson, 1988). The authors subjected the 14-item PSS to principal component analysis (PCA) and found four items that displayed poor loadings on the first principal component in the range of 0.11–0.39, which were removed resulting in the popular 10-item version of the PSS (called PSS-10). A 4-item PSS version was also introduced as a quick assessment tool, but the PSS-10 demonstrated better internal consistency (alpha = .78) compared to the 4-item PSS version (alpha = .60) and was recommended by the authors for use in future research (Cohen and Williamson, 1988). The PSS-10 has been used with different populations for both clinical assessments and empirical investigations including validation studies, all confirming its satisfactory psychometric properties (Mitchell et al., 2008; Roberti et al., 2006; Taylor, 2015). Currently, there is no agreement on the factor structure of the PSS-10, with some studies including the original validation report confirming unidimensionality (Cohen and Williamson, 1988; Cole, 1999), while others argue that a two-factor solution provides a better fit (Barbosa-Leiker et al., 2013; Taylor, 2015; Teh et al., 2015).
There are only few studies that used item response theory (IRT) to investigate functioning of the individual PSS-10 items (Cole, 1999; Sharp et al., 2007; Taylor, 2015). Cole (1999) used IRT to investigate potential item bias of the PSS-10 with a larger US sample (n = 2264). Significant, but very small, differences in performance of several items were reported by gender, ethnicity and education level categories. Sharp et al. (2007) tested the performance of the PSS-10 items in a clinical sample of asthma patients and reported that few items function differently across ethnic and literacy factors. Recently, Taylor (2015) used the graded response IRT model and reported satisfactory functioning of individual items. These inconsistent findings suggest that further investigation is necessary. Moreover, these findings do not provide feasible solutions to improve the overall precision of the PSS-10.
Technically, an ordinal scale such as the PSS should not be used with parametric statistics without violating their fundamental assumptions. However, researchers have used the PSS with parametric statistics (Chavez-Korell and Torres, 2014; Gitchel et al., 2011), which could potentially result in misleading conclusions and implications of the findings linked to individual’s health and well-being. It has been demonstrated that adding ordinal scores of individual items together would not produce an accurate total estimate because each individual item explains different amount of variance due to a latent trait (Allen and Yen, 1979; Stucki et al., 1996). Moreover, usage of the ordinal PSS-10 in research may affect comparisons with neurophysiological interval or ratio-level data (e.g. electroencephalogram, skin conductance and heart rate), an especially important consideration in stress research. Therefore, it is important to improve precision of the PSS-10 up to an interval-level scale, which can be conducted using Rasch analysis – an approach that is particularly suited for this purpose (Rasch 1960; Tennant and Conaghan, 2007).
The dichotomous Rasch model was developed first (Rasch, 1960), which was followed by two models designed for polytomous items, such as used in the PSS-10, including the Partial Credit version (Masters, 1982) and the Rating Scale version (Andrich, 1978). Both models based on assumption that differences between thresholds of individual items vary. However, the Rating Scale model assumes that these variations are similar across all items but the Partial Credit model allows thresholds distances to vary across items, so that every item has individual rating scale parameters. Both polytomous Rasch models are commonly used in health assessment to test and improve psychometric properties of ordinal measures with two or more response options (Hobart and Cano, 2009; Lundgren Nilsson and Tennant, 2011). The decision which polytomous model to use depends on the likelihood-ratio test, which compares threshold distances between individual items prior to main analysis. The unrestricted Partial Credit model will be used if these distances are significantly different (Tennant and Conaghan, 2007). The end product of Rasch analysis is transformation of scores from an ordinal to an interval scale that increases precision of measurement, which was established both theoretically (Rasch, 1960; Tennant and Conaghan, 2007) and empirically (Norquist et al., 2004).
Rasch analysis has demonstrated its distinct advantages over other more traditional psychometric methods, which have been discussed in detail elsewhere (Rasch, 1960; Wilson, 2005). The analysis involves testing several important psychometric parameters including potential item bias, local independence assumptions (e.g. unidimensionality), correct stochastic ordering of items and response option ordering in polytomous items (Tennant and Conaghan, 2007). When data meet these requirements and a fit to the Rasch model is achieved, participants can be located according to their ability on a scale measuring the latent factor (in this case perceived stress). Indeed, in Rasch analysis, both the participants and the items are ordered on an interval scale using the same log-odds metric (Pomplun and Custer, 2005). This results in the graphical representation of an item-person threshold distribution, showing how well the range of the items’ difficulty covers the ability of a sample. Rasch analysis provides a system that can accurately and rigorously investigate this perceived stress measure in terms of potential item bias, differential item functioning (DIF) and unidimensionality.
In sum, it is assumed that the PSS-10 is a measure of perceived stress and has generally accepted psychometric properties that can be applied to a wide range of clinical and non-clinical populations. However, its ability to discriminate precisely between perceived stress levels has not been investigated in sufficient detail. Rasch analysis is a suitable method to investigate the ability of the scale and individual items to discriminate on their overarching latent factor. However, to the best of our knowledge, this advanced technique has not yet been applied to scrutinize and to improve the psychometric properties of the PSS. This study aims to apply Rasch analysis to explore strategies to improve the psychometric properties of the PSS-10 up to an interval level of measurement. This psychometric investigation explicitly focuses on the measurement of the overarching latent factor of perceived stress, rather than its underlying facets already explored in the literature (Roberti et al., 2006; Taylor, 2015). To ensure that results are generalizable to diverse populations, a combined sample of respondents from New Zealand and the United States is used, consisting of university students as well as respondents from the general population. Additionally, the dataset is sufficiently large to allow splitting of the combined sample into two sets, thus allowing replication of the Rasch analysis of one half of the sample with the other half. An ordinal-to-linear conversion table was generated, which can be used to enhance precision of the PSS-10.
Method
Participants
This study combined three independently collected samples. Sample 1 (n = 1102) consisted of the Auckland (New Zealand) general population who participated in a postal survey on noise sensitivity and health (Hill et al., 2014). The mean age was 51 (standard deviation (SD) = 16.42), and 65 per cent were female. Sample 2 (n = 479) contained New Zealand university students enrolled in various health science courses at the University of Auckland and Auckland University of Technology. The majority (76%) were female, and the overall mean age was 19.96 (SD = 4.47). Sample 3 (n = 396) consisted of US university students enrolled in first-year psychology classes at West Chester University in the United States. The mean age was 19.18 (SD = 2.21), and 45 per cent were female. From each of these three samples, 300 participants were randomly selected to create an overall sample of 900 participants from the New Zealand general population, New Zealand university students and US university students. Each subset was randomly divided by half to create two samples of 450 participants, where 150 participants were included from each sample population. The main analyses were conducted using one of these two samples and were then replicated with the second.
As the ethnic profile for the New Zealand and US samples was very different, common categories were created so that DIF by ethnicity could be compared. In the overall sample of 900, 65 per cent were classified as Caucasian, 5 per cent as Polynesian, 9 per cent as Asian and 19 per cent as other.
Procedure
The procedure for the New Zealand general population (Sample 1) is reported in detail elsewhere (Hill et al., 2014). Auckland residents received a questionnaire on noise sensitivity, perceived stress and health, which they subsequently posted back to the researchers using a self-addressed pre-paid envelope. Participants in Sample 2 were university students enrolled in health science courses at two major universities in New Zealand completing a survey on motivation to learn, quality of life and perceived stress. Students at the University of Auckland received an invitation to complete an online survey and thus responded at a time of their convenience. Students at Auckland University of Technology were approached in lecture theatres and completed the questionnaire during the lecture break or after the lecture. For Sample 3, students taking an introductory psychology class at West Chester University in Pennsylvania, United States, completed a research study on motivation to learn, quality of life and perceived stress as an option in fulfilling their class research credit. The study was conducted in compliance with the authors’ institutional ethics committees.
Measures
While the questionnaires in the three above-mentioned studies contained various scales, this study focuses on the PSS-10 only. The PSS-10 is a 10-item self-report questionnaire of perceived stress operationalized as subjective evaluation of lack of control, unpredictability and overload in participants’ daily life (Cohen and Williamson, 1988). The instrument uses a 5-point Likert-scale response format (1 = ‘Never’ to 5 = ‘Very often’), and a total score is calculated after reverse-coding items 4, 5, 7 and 8 and then adding scores of all 10 items together.
Data analysis
Descriptive statistics and reliability of the PSS-10 were computed using IBM SPSS v.22, and Rasch analysis used RUMM2030 (Andrich et al., 2009). A likelihood-ratio test was conducted first and indicated that differences between response options thresholds of individual items vary significantly across items meaning that the unrestricted (Partial Credit) version of the Rasch model should be used for the current dataset. The Rasch analysis is conducted in an iterative way until all individual items show sequential ordering of response categories thresholds, satisfactory overall and individual item fit to the model are achieved and unidimensionality is clearly evident (Siegert et al., 2010). Category threshold is disordered if a probability to choose the closest higher response option is lower than any of subsequent higher response options. The Rasch model fit is evaluated by the mean item and person location, individual item fit residual and the overall item-trait interaction chi-square test/p value using the following criteria (Gustafsson, 1980; Tennant and Conaghan, 2007):
The item location mean is used as a base and set to zero.
A person location mean ±0.5 indicates a good coverage of a sample by a scale.
In case of an overall perfect fit, both item and person fit residual are 0.00 (SD = 1.00).
Individual items fit residuals should be in the range from −2.50 to +2.50.
The item-trait interaction chi-square is an indicator of the overall model fit and should be not significant (p > .05) if data fit the Rasch model.
The person separation reliability, which is similar to a Cronbach’s alpha numerically (range from 0 to 1), is not an index of Rasch model fit but indicates how well individual trait levels are spread out along the scale continuum represented by the items (Fisher, 1992).
Both the overall and the individual item fit to the Rasch model could be affected by local dependency between items, which refers to a situation when two or more items are associated in some way. For example, one item about negative emotions asked about ‘been upset’ and the other about ‘been angered’. If a person is often getting angry, they should also frequently get upset. Such relationship violates local independency assumption, compromises estimation of model parameters and inflates reliability (Wright, 1996). Local dependency between two or more items is reflected by the residual correlations matrix. A residual correlation with magnitude more than 0.20 compared to the mean of all residual correlations is a sign of local dependency (Christensen et al., 2017; Marais and Andrich, 2008). In this study, we considered removing items as the last resort to achieve the model fit. Instead of removing locally dependent items, these items can be simply added together into a subtest or testlet to solve local dependency issues (Lundgren-Nilsson et al., 2013; Wainer and Kiely, 1987). Using the example with negative emotions, the subtest would be similar to one item measuring both aspects (get upset and angered). Subtests in Rasch analysis are analogous to item parcels, and their advantages have been well documented elsewhere (Little et al., 2002; Rushton et al., 1983). Combined items show higher reliability compared to individual items, more scale points that contribute to accuracy of measurement and lower risk of spurious correlations. Moreover, more accurate estimates of latent structures were obtained using item parcels compared to individual items because combining items measuring the same construct reduces measurement error due to an individual item.
The basic assumption of the Rasch model is unidimensionality of a scale (Rasch, 1960, 1961). Smith’s (2002) unidimensionality test is commonly used in Rasch analysis with RUMM2030 software (Andrich et al., 2009) that involves PCA of the residuals and the equating t test. Unidimensionality of the scale is evident if significant t-test comparisons do not exceed 5 per cent or if the lower bound of a binominal confidence interval computed for the number of significant t-tests overlaps 5 per cent cut-off point. Rasch model requires no significant differences (DIF) in item functioning due to personal factors including gender, age, sample (e.g. US university students vs New Zealand university students vs New Zealand general population), ethnic groups and education levels. DIF analysis involves comparing distributions of individual scores aggregated by class interval mean scores between groups for each person factor and for every individual item using analysis of variance (ANOVA) and t-tests for pairwise comparisons (Bonferroni adjusted for the number of tests).
The first sample of 450 was used for the main Rasch analysis, which was then replicated using the second half of the sample. Both samples (datasets a and b) contained a large enough number to satisfy the recommended sample size estimates for the Rasch analysis (Linacre, 1994). To examine if there is any significant difference to participant estimates of the modified version, participant estimates for the original scale and the modified version were compared using a paired-samples t-test. Finally, ordinal-to-interval transformation scores were computed that allow users to transform ordinal data to an interval-level scale.
Results
Cronbach’s alpha for the PSS-10 with the current dataset (N = 900) was .88 indicating good internal consistency with all 10 items having item-to-total correlations in the range from 0.49 to 0.74. Table 1 provides summary of the Rasch model fit statistics for the initial and the final analysis including both the overall and the individual item fit indices for the sample (a) and the overall fit indices for replication of the analysis with sample (b). Overall, the person location mean was within the acceptable range ±0.50 in both samples suggesting good targeting of the sample by the scale items (Table 1). Also, the scale showed good person separation reliability of .88. However, in both samples, the initial overall fit to the Rasch model was affected by significant item-trait interaction indexed by chi-square test: χ2(90) = 129.26 for sample (a) and χ2(90) = 150.62 for sample (b), p < .001 (Table 1). This means that the scale cannot adequately discriminate between respondents at different levels of the latent trait (perceived stress). Although no items with disordered thresholds were identified, item 10 displayed significant misfit to the Rasch model with fit residual below the acceptable cut-off point of −2.50 (Table 1) in both samples. Additionally, item 4 showed deviation from the model expectations with a positive fit residual above 2.50 in sample (b) but not in sample (a).
Summary of the Rasch model fit statistics for the initial (1a) and final (2a) analysis of the PSS-10 (n = 450) and final analysis of its subsequent replication (2b; n = 450).
SD: standard deviation; PSS: Perceived Stress Scale; LB: lower bound of the 95 per cent confidence interval surrounding t-test.
Significant misfit to the Rasch model.
Given the overall poor fit and misfits at individual item level, the residual correlation matrix was examined because local dependency between items affects both discrimination parameters and test information associated with the Rasch model fit (Lundgren Nilsson et al., 2013).
Local dependency
Examination of the residual correlation matrix indicated correlations with magnitude above critical value of 0.20 for items 1, 2 and 9; items 4, 5, 7, and 8; and items 6, 10 and 3; indicating local dependency between those items. To verify this observation, a Spearman’s correlation matrix between all PSS-10 items was generated and confirmed higher correlations between these items in the range of 0.50–0.60. These observations together confirmed three groups of locally dependent items, which were combined into three subtests to address local dependency (Lundgren Nilsson et al., 2013). These minor modifications provided good solution to improve the overall scale function that satisfied the expectations of the Rasch model in both samples (χ2(27) = 29.92, p = .36 (a), χ2(27) = 36.30, p = .11 (b); Table 1). At this stage, the person location mean was close to zero in both samples (Table 1) and indicated better coverage of the sample by the subtest items with acceptable person separation reliability (.80–.81). Additionally, none of the individual subtest items showed misfit to the Rasch model (Table 1, 2a). Therefore, the best model fit was achieved without the need to remove any of the PSS-10 items.
DIF
ANOVA indicated significant DIF effects by sample on the subtest one (F(2, 6) = 6.27, p = .002, Bonferroni adjusted p = .006), but not for the subtests two and three. However, post hoc comparisons between sample groups showed no significant difference with p ranging from .065 to .217. There were no significant DIFs in functioning of subtest items due to other personal factors including gender, age, ethnic groups and education levels (p > .05, Bonferroni adjusted).
Test for unidimensionality
To test the unidimensionality of the final model solution, the person estimates from the subtest with the highest positive loadings on the first principal component were compared with the estimates from the subtest with the highest negative loadings. Unidimensionality was confirmed for both samples with 5.19 per cent significant t-tests overlapping 5 per cent cut-off point on the lower bound of the confidence interval (3.16%) for sample (a) and 4.74 per cent of significant t-tests and cut-off overlap of 2.71 per cent for sample (b) (Table 1).
Item-person threshold distribution
Figure 1 shows the person-item threshold distribution of the modified PSS-10 after combining locally dependent items into three subtests (Table 1, analysis 2a). Here, subtest item thresholds and person ability levels on the latent factor measured by the PSS-10 are plotted using the same metric in logit units. Distribution of persons is close normal, and the modified PSS-10 item thresholds satisfactorily cover 98 per cent of the participants’ abilities on the latent factor (perceived stress).

Distribution of persons and item thresholds set to interval length of 0.20 in logit units making 30 groups (final analysis, sample (a), n = 450).
Equating t-test
The means of person estimates from the original PSS-10 and the modified version were compared by a paired-samples t-test. The difference between the person estimates of the two versions was significant (t(449) = 7.22, p < .01), indicating successful alteration of the ability of the final model to discriminate between individual levels of perceived stress compared to the original version. This confirms that the implemented modifications (subtests) resulted in an improved solution for the PSS.
Ordinal-to-interval conversion table
Table 2 provides interval scores in both logit units and the original PSS-10 scale format that allows researchers to convert ordinal raw scores to interval-level scores.
Converting from a raw PSS-10 score (10–50) to an interval scale in logit units and using the original scale metrics.
PSS: Perceived Stress Scale.
The raw score is calculated by reverse-scoring questions 4, 5, 7 and 8 and then adding scores of all 10 items together. This table cannot be used for respondents with missing data.
Researchers who have already used the PSS-10 to collect data or are planning to use the scale can apply the results of this study as follows: calculate the raw score by reverse-scoring questions 4, 5, 7 and 8 and then adding scores of all 10 items together. Next, use Table 2 to convert these scores to the corresponding interval-scale scores ranging from 10 to 50, identical to the range of scores of the original PSS scoring system. Using the conversion table provided here, users are able to increase the precision of the PSS-10. It should be noted that conversion from the ordinal-to-interval scale proposed here does not require altering the original response format of the PSS-10 scale. This conversion table was independently generated with the second sample (b) (n = 450), showing almost identical results.
Discussion
The PSS-10 is a widely used instrument to measure perceived stress on an ordinal scale; however, its precision has not been yet fully optimized. This study used the subtests methodology analogous to Lundgren Nilsson et al. (2013) to improve the psychometric properties and precision of the 10-item PSS up to an interval-level scale. The best model fit was achieved after combining locally dependent items into three subtests, and no systematic DIF by personal factor such as gender, ethnicity, education and sample population were evident. After these minor modifications, the psychometric properties of the PSS-10 are robust, and transformation from an ordinal to an interval-level scale can be conducted using the conversion algorithm provided in Table 2. This Rasch analysis contributed to the limited number of IRT-based studies (Cole, 1999; Sharp et al., 2007; Taylor, 2015) that focused on the functioning of individual PSS-10 items by increasing the precision of the PSS-10 and addressing local dependency issues.
Local dependency found between items 4, 5, 7 and 8 is consistent with earlier research (Roberti et al., 2006; Taylor, 2015), where the same items were proposed as a second factor in a two-factor PSS-10 solution. All these four items are negatively worded and thus measure coping abilities as opposed to perceived stress, which could explain previous findings of these items loading together as a factor (Roberti et al., 2006; Taylor, 2015). However, after combining locally dependent items into subtests, unidimensionality of the PSS-10 was clearly evident in both the main and the replication samples, suggesting that, after Rasch modifications, the PSS-10 is clearly tapping into one overarching latent factor. Local dependency found between items 1, 2 and 9 is not surprising because these items all contain themes of control over external events. Another subtest included items 3, 6 and 10, which are explicitly related to perceived helplessness (e.g. could not overcome or cope) and hence explain local dependency. These clusters of locally dependent items may have influenced variability of the earlier factor analysis (Roberti et al., 2006; Taylor, 2015) and may even generated spurious factors (Lundgren Nilsson et al., 2013). However, the clear evidence of unidimensionality of the PSS-10 based on three subtests replicated by two random samples together with the large amount of shared variance suggest that a total PSS-10 interval score reflects perceived stress levels of the majority of people (Figure 1).
The main contribution of this study is that the PSS-10 raw score can now be converted from an ordinal scale to an interval scale, which means that parametric statistics can be conducted without violating their fundamental assumptions. The interval-level estimates of the latent factor offer researchers the opportunity to examine the effects of mediators and moderators of perceived stress in various contexts. Rasch interval–transformed PSS-10 scores can reliably be used in such models given that potential item biases (DIF) and local dependency issues are ultimately resolved by Rasch analysis if data fit the model expectations. The improved precision of the instrument is also highly desirable in clinical assessment, where it informs accuracy of diagnosis and treatment of stress-related conditions. It is important to note that these improvements were possible without the need to alter the original PSS-10 response format meaning that existing datasets can easily be re-analyzed to provide interval level of measurement.
The following limitations are acknowledged. The samples may not reflect the full diversity of New Zealand’s or the US’s ethnic groups and no efforts were made to purposively sample under-represented groups. The response rate of 15 per cent was relatively low for the New Zealand general population sample (Hill et al., 2014), although such response rates are not uncommon for research of this nature in New Zealand (Krägeloh et al., 2013). Even though the Rasch-modified PSS-10 covers 98 per cent of the sample abilities, there are individuals uncovered by item thresholds on both upper and lower level of the scale. Future studies may consider to test whether using more response options with extreme categories will provide better coverage for individuals with lower and higher stress levels.
Conclusion
Stress may affect both physical and mental health and its accurate assessment represents an ongoing challenge. This study has demonstrated that after minor modifications, the widely used perceived stress measure PSS-10 satisfies the expectations of the unidimensional Rasch measurement model. The precision of the PSS-10 can be optimized up to an interval-level scale using the ordinal-to-interval conversion tables published here, which can be of benefit for researchers investigating neurophysiological, psychological and environmental correlates of stress. This study is best assessed from the perspective of Rasch analysis.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study is a part of a doctoral work of the first author funded by the Vice-Chancellor’s Scholarship of the Auckland University of Technology.
