Abstract
As online data collection services such as Amazon’s Mechanical Turk (MTurk) gain popularity, the quality and representativeness of such data sources have gained research attention. To date, the majority of existing studies have compared MTurk workers with undergraduate samples, localized community samples, or other Internet-based samples, and thus, there remains little known about the personality and mental health constructs of MTurk workers relative to a national representative sample. The present study addresses these limitations and broadens the scope of existing research through the use of the Personality Assessment Inventory, a multiscale, self-report questionnaire which provides information regarding data validity and personality and psychopathology features standardized against a national U.S. census–matched normative sample. Results indicate that MTurk workers generally provide high-quality data and are reasonably representative of the general population across most psychological dimensions assessed. However, several distinguishing features of MTurk workers emerged that were consistent with prior findings of such individuals, primarily involving somewhat higher negative affect and lower social engagement.
Internet data sources have broadened the scope of data collection in behavioral research, as data sets that once took months to build can now be collected in a single afternoon. Amazon’s Mechanical Turk (MTurk) platform is one commonly used data collection service in which registered users from around the world, or workers, complete surveys and computerized tasks, referred to as Human Intelligence Tasks or HITs, for small financial incentives. Collecting data via services such as MTurk has become so commonplace that a Google Scholar search reveals references to “Mechanical Turk” in more than 20,000 studies published within the past 10 years. However, while such ease of data collection has clear advantages for research efficiency and productivity, concerns about data quality and representativeness have also been raised (Chandler & Shapiro, 2016).
As Internet-based data collection grew in popularity, apprehensions about data quality arose around allowing participants to complete study tasks wherever and whenever they were inclined to do so, leading to concerns about inattentiveness, lack of supervision, and the number of possible distractions (Chandler & Shapiro, 2016). Furthermore, the typically small financial compensation per task (e.g., some might pay US$0.05 for completion of a short survey or task) provides an incentive for workers to complete tasks as quickly as possible in order to maximize earnings, potentially at a cost to accurate responding. MTurk workers also have unconstrained access to external resources that are not afforded to participants in laboratories, and there is some evidence that MTurk workers may be using those resources to answer factual-based research items (Goodman, Cryder, & Cheema, 2013).
Despite such concerns, however, research has generally suggested that MTurk workers provide high-quality data, demonstrating psychometric equivalence to other data collection methods with respect to internal consistency (Arditte, Çek, Shaw, & Timpano, 2016; Shapiro, Chandler, & Mueller, 2013) and test–retest reliability (Buhrmester, Kwang, & Gosling, 2011; Shapiro et al., 2013). Likewise, although “cheating” with external sources does seem to be a valid concern, providing MTurk workers with specific instructions to not use external resources significantly reduces this occurrence (Goodman et al., 2013), suggesting that MTurk workers attend to and comply with study directives. Furthermore, MTurk workers demonstrate passing rates on instructional manipulation checks (IMCs; Oppenheimer, Meyvis, & Davidenko, 2009) that are typically comparable to other data sources, if not better (Hauser & Schwarz, 2016; Paolacci, Chandler, & Ipeirotis, 2010; Kees, Berry, Burton, & Sheehan, 2017), offering further evidence that MTurk workers are attentive to study instructions and tasks. Nevertheless, there are some exceptions to this general finding, with Goodman et al. (2013) reporting that MTurk workers had similar IMC passing rates as a community sample drawn from street passerby but performed significantly worse as compared with an undergraduate subject pool.
Although attentiveness is essential to data quality, attention checks do not address the possibility that response styles may be systematically distorted in other ways, potentially through factors such as social desirability or deviant responding (Shapiro et al., 2013). Thus, in addition to IMCs, the inclusion of positive and negative response distortion indicators represents another important consideration when assessing data quality, particularly for studies pertaining to clinical symptoms. One means by which researchers have attempted to detect response distortion is the inclusion of validity scales, such as the F scale of the Minnesota Multiphasic Personality Inventory–2 (MMPI-2; Butcher, Dahlstron, Graham, Tellegen, & Kaemmer, 1989). In a study of MTurk worker characteristics, Shapiro et al. (2013) reported that per the interpretive guidelines of the MMPI-2, 3% of respondents met the cutoff for “malingering” on the F scale (i.e., five standard deviations above the mean). Furthermore, applying a more conservative cutoff (i.e., three standard deviations above the mean) resulted in 10% of the sample being classified as potential feigners. However, although these data seem to imply prominent levels of response distortion, it is important to note that individuals with elevated F scores were excluded from the normative data of the MMPI-2 (Butcher et al., 1989), and Shapiro et al. (2013) do not report mean F scale scores. Thus, it is unclear the extent to which the results of Shapiro et al. (2013) reflect F scale elevations in the MTurk sample that diverge from what might be found in a normative community sample that was not previously screened using F scores. Additionally, because Shapiro et al. only administered the F scale, possible corresponding elevations on the MMPI-2 clinical scales could not be examined.
While data quality is an important concern, recent discussions about the utility of MTurk data have also focused on data representativeness. Demographic comparisons have suggested that MTurk workers are typically more representative than undergraduate samples in terms of age and ethnical diversity (Buhrmester et al., 2011), but there are a few notable demographic differences between MTurk samples and the U.S. population as a whole. MTurk workers tend to be younger and better educated than the general population, although they are often unemployed or underemployed (Shapiro et al., 2013) and report lower incomes (Corrigan, Bink, Fokuo, & Schmidt, 2015; Paolacci et al., 2010). Furthermore, MTurk workers are less likely to own their own home (Berinsky, Huber, & Lenz, 2012), and nearly 20% report the oldest member of their household as being 20 years older than the respondent (Casey, Chandler, Levine, Proctor, & Strolovitch, 2015), indicating that a large proportion of MTurk workers may be living with parents.
MTurk workers have also been found to differ from traditional samples on psychological dimensions, with differences that seem to be consistent with what is known about frequent Internet users (Goodman & Paolacci, 2017; Paolacci & Chandler, 2014). With respect to five-factor personality trait domains, MTurk workers have been shown to be more neurotic and less extroverted than both student and community samples (Goodman et al., 2013; Kosara & Ziemkiewicz, 2010), and they have also been found to be less agreeable than undergraduate participants (Kosara & Ziemkiewicz, 2010). MTurk workers also report significantly lower self-esteem than students and marginally lower self-esteem as compared with a community sample (Goodman et al., 2013). With regard to psychopathology, the most notable and reliable difference between MTurk workers and the general population is in prevalence rates of social anxiety and withdrawal (Arditte et al., 2016; Shapiro et al., 2013). Arditte et al. (2016) reported that nearly half (49%) of MTurk respondents met clinical criteria for social anxiety disorder, in contrast to the 7% 12-month prevalence rate in the general population. MTurk workers also seem to have higher prevalence rates of autism spectrum disorder than the general population (Mitchell & Locke, 2015) and report more autism–spectrum traits than student samples (Eriksson, 2013). These findings are aligned with further research suggesting that MTurk workers score slightly higher on the personality domain of social detachment (Miller, Crowe, Weiss, Maples-Keller, & Lynam, 2017). Additionally, there is some evidence that MTurk workers demonstrate clinical levels of depression as well as subclinical levels of anxiety (Arditte et al., 2016), although other research has suggested that depression and anxiety prevalence rates among MTurk workers mirror the general population (Shapiro et al., 2013).
Research regarding the personality traits and clinical symptoms of MTurk workers have begun to characterize the population of MTurk workers, but there is little research examining these personality and mental health constructs in comparison with a representative normative sample. To date, the majority of existing studies have compared MTurk workers against undergraduate student samples, localized community samples, or other Internet-based collection services, given that normative data were not available for the measurements used. The present study addresses these limitations and builds on previous research through the use of the Personality Assessment Inventory (PAI; Morey, 1991, 2007), a multiscale, self-report questionnaire which provides information regarding data validity and personality and psychopathology features that was standardized against a national U.S. census–matched normative sample. The present study aims to determine if previously observed trends are corroborated when compared with a large, nationally representative sample, thus potentially offering broad insight into how MTurk workers may differ from the general population.
Method
Participants
A total of 455 participants (57% male) were recruited from Amazon’s MTurk, with the requirement that they be at least 18 years old and located in the United States (as verified by their IP address). Participants ranged from 19 to 70 years of age, with a mean age of 35.5 years (SD = 10.4). The majority of participants identified as European American (76.5%), followed by African American (7.9%), Asian American (7.7%), Hispanic or Latino/a (5.5%), Native American (1.5%), and Other (0.9%). Individuals located outside the United States were restricted from participation given that the study measure was normed using data collected within the United States, and all study materials were in English. Although researchers have the option to recruit only workers with a documented history of good performance, no restrictions on the basis of task approval rate or any other qualification measure were used in the present study. Participants were compensated $7.50 for completion of the study materials, which included additional measures not examined in the present study (548 total items; mean response time = 67 minutes).
Measures
Personality Assessment Inventory
The Personality Assessment Inventory (PAI; Morey, 1991) is a 344-item structured, self-administered questionnaire that assesses a variety of personality and psychopathology domains. The PAI consists of 22 nonoverlapping scales, including 4 validity scales, 11 clinical scales, 5 treatment scales, and 2 interpersonal scales. The PAI scale and subscale raw scores are linearly transformed to T-scores (mean of 50, standard deviation of 10) to provide interpretation relative to a standardization sample composed of 1,000 community-dwelling adults who were selected based on cross-stratification of the variables of gender, race, and age to match U.S. Bureau of the Census projections as of the publication of the PAI. The only specification for inclusion in the census-matched sample was the endorsement of at least 90% of PAI items, that is, no more than 33 items could be left blank. Participants were not removed from the normative sample on the basis of elevated validity indicators, thus allowing for straightforward interpretation of study sample validity scale elevations as compared with the normative sample. The PAI was used in the present study as a means of comparison between the MTurk sample and the community normative sample on indicators of data quality and personality and psychopathology features. Specific scales relevant to the different aspects of the comparison are described below.
Data Quality
The PAI includes four validity scales that were intended to identify response patterns that deviate from accurate and honest responding, including random, careless, or manipulated response sets. There are two scales for the assessment of random response tendencies (Infrequency, INF; Inconsistency, ICN). The INF scale was conceptually derived to include items with very low endorsement rates among both normative and clinical samples, thus indicating a pattern of infrequent responses that is uncorrelated with psychopathology. Conversely, the ICN scale is composed of 10 highly correlated pairs of items, in which for the majority of respondents, the response to one is highly predictive of the response to the other. High scores on the ICN scale are thus reflective of inconsistency in responding to these pairs of items with similar content. The recommended cutoff scores for INF and ICN are 75T and 73T, respectively; a joint disjunctive use of these scales correctly identified 94.1% of random protocols (Morey, 1991). These scales, particularly the INF scale, have been used in numerous studies to assess data quality (e.g., Boccaccini, Murrie, Hawes, Simpler, & Johnson, 2010; Ingersoll, Hopwood, Wainer, & Donnellan, 2011; Wright et al., 2012).
The PAI also contains scales for the assessment of systematic negative (Negative Impression Management, NIM) and positive (Positive Impression Management, PIM) responding. The NIM scale is composed of items that reflect extreme, atypical symptoms and features that reflect a markedly negative response style. Elevated scores on NIM may indicate an exaggerated presentation of symptoms, possibly as a cry for help or a result of purposeful feigning. Conversely, the PIM scale is composed of items that represent unrealistically virtuous characteristics or features and tend to be more often endorsed by those trying to make a favorable impression than in normative samples. Scores above 57T on PIM and above 81T on NIM have demonstrated success in identifying respondents engaging in positive impression management and negative impression management, respectively (Hawes & Boccaccini, 2009; Morey, 2007). All participant data were included in the present analysis, regardless of elevations on the validity scales.
Personality and Psychopathology Features
The PAI includes 11 clinical scales, 5 treatment scales, and 2 interpersonal scales, as well as corresponding subscales, which assess a broad range of psychological dimensions. Several of these scales are of particular relevance given previous findings regarding MTurk worker characteristics, including scales assessing depression, anxiety, anxiety-related disorders, and social detachment.
Data Analysis
To control for demographic differences between the present sample and community normative sample, mean T-scores, standard deviations, and internal consistencies of the PAI scales and subscales were calculated while weighting the variables of gender, age, and ethnicity with respect to the stratification of the census-matched standardization sample (Morey, 1991). Given that the large sample sizes of both the MTurk sample and the community normative sample would render even minor differences statistically significant, differences between the weighted M-Turk mean T-scores and the community normative mean T-scores were considered worthy of note at two thresholds: exceeding one standard error of measurement (SEM) for the given substantive scale or subscale (i.e., the 67% confidence interval), and more conservatively, exceeding 1.96 SEMs, which reflects the 95% confidence interval for PAI scales and is typically considered to reflect a clinically significant difference (e.g., Jacobson & Truax, 1991). Standard errors of measurement for PAI substantive scales range from 2.8 to 5.7 T-points for the full scales, and 3.9 to 5.7 for the subscales (Morey, 1991). In addition, Cohen’s d effect sizes were computed to indicate the magnitude of difference for each scale between the weighted MTurk sample and the PAI community normative sample (Cohen, 1988).
Results
Demographics
MTurk workers were significantly younger (M = 35.5 years, SD = 10.41) than the community normative sample (M = 44.22 years, SD = 17.23), t(1453) = 9.99, p < .01. Ethnic composition of the MTurk workers differed from the PAI census–matched sample, χ2(5, N = 1,455) = 81.37, p < .01; the MTurk sample was primarily Caucasian/European American (76.5%), and overrepresented Asian Americans (7.7%), while underrepresenting African American (7.9%) and Hispanic/Latinx respondents (5.5%). The MTurk sample was predominantly male (56.9%), which was a significantly higher proportion than the PAI community normative sample (48%), χ2(1, N = 1,455) = 9.96, p < .01.
Data Quality
Table 1 presents the weighted and unweighted mean T-scores and standard deviations of the MTurk sample on the PAI scales and subscales, as well as the Cohen’s d effect sizes comparing the weighted T-scores with the standardized T-scores of the community normative sample (M = 50, SD = 10). MTurk workers did not obtain higher scores on the indicators of careless or random responding (i.e., ICN and INF), and even demonstrated a small negative effect on INF as compared with the community normative sample, signifying that MTurk respondents were similarly attentive to study materials. The MTurk sample also did not differ from the community normative sample on indicators of positive and negative response distortion (i.e., PIM and NIM), suggesting that MTurk workers were not over- or underreporting clinical symptomatology to a greater extent than might be expected in the general population.
Mean T-Scores, Standard Deviations, Effect Sizes, and Internal Consistency of MTurk Data.
Note. Cohen’s d effect sizes reflect the comparison between the mean T-scores of the weighted MTurk sample data and the Personality Assessment Inventory community normative sample (M = 50, SD = 10; Morey, 1991).
ICN and INF do not represent substantive constructs and thus internal consistency analyses of these scales are not applicable. bScores with differences from the community normative sample exceeding one standard error of measurement (i.e., the 67% confidence interval). cScores with differences from the community normative sample exceeding the reliable change index threshold of clinical significance (i.e., the 95% confidence interval).
Psychological Dimensions
In general, the MTurk sample showed negligible differences from the community normative sample across most personality and psychopathology dimensions assessed. The few clinically significant deviations that were observed tended to include scales and subscales pertaining to issues with negative affect and social engagement. Only two scores reached the threshold of clinical significance (i.e., 1.96 SEMs). The most prominent feature of MTurk workers relative to the community sample was higher scores on the social detachment scale (SCZ-S), assessing social isolation and aversion. Additionally, Depression (DEP) scale scores were moderately higher than the normative sample, particularly pertaining to cognitive symptoms (DEP-C). It should be noted that although SCZ-S is a subscale of the Schizophrenia (SCZ) scale, low scores were observed on the Psychotic Experiences (SCZ-P) subscale in the MTurk sample, suggesting that the observed mean SCZ-S score is unlikely to be related to schizophrenia spectrum phenomena.
Sample differences on several additional scores reached the one SEM threshold, reflecting deviations from the community normative sample. MTurk workers in the present sample reported moderately less social support from friends and family (Nonsupport; NON) and slightly more resentment toward others (Resentment; PAR-R) than the community normative sample and appeared to be moderately lower on the Warmth (WRM) interpersonal scale. These observed issues in social functioning were further corroborated by higher scores on the Phobias (ARD-P) subscale of the Anxiety-Related Disorders (ARD) scale; notably, the ARD-P scale often elevates in the presence of social anxiety (Morey, 2007). MTurk workers also demonstrated slightly higher scores on indicators of negative affect relative to community norms, specifically cognitive (ANX-C) anxiety symptoms, and higher scores on the Suicidal Ideation (SUI) scale and the Traumatic Stress (ARD-T) subscale.
Discussion
The results of this study corroborate and extend existing research findings regarding the quality and representativeness of MTurk research samples through comparison to a large, national representative sample. MTurk data collection is clearly efficient, and there is obvious potential for such data collection platforms to positively influence the ease and productivity of empirical research. However, it is important to understand that MTurk samples are not perfectly representative of the broader population. Appropriate interpretation of MTurk-derived data requires that researchers are cognizant of the unique features that differentiate MTurk workers from the general population and consider the potential implications of these differences.
MTurk workers in the present sample were demographically similar to MTurk workers in previous samples, and demographic differences between the MTurk sample and the community normative sample supported some of the previous comparisons between MTurk workers and the general population, with a few exceptions. As consistently demonstrated, MTurk workers were relatively young (Paolacci et al., 2010; Shapiro et al., 2013), and the sample was ethnically diverse but not representative of the general population (Chandler & Shapiro, 2016). The sample was also predominantly male and demonstrated a greater ratio of male to female respondents than might be expected in the general population.
MTurk workers did not differ meaningfully from the community normative sample on any of the four validity scales, offering further support for the general consensus that MTurk workers provide high-quality data and are generally attentive to study tasks (Hauser & Schwarz, 2016; Kees et al., 2017; Paolacci et al., 2010). Likewise, despite some concerns regarding MTurk workers potentially feigning symptoms (Shapiro et al., 2013), the presence of overreporting or underreporting of clinical symptomatology does not appear to vary from what might be expected from a large, representative sample. These findings are particularly promising given that no special qualification requirements were imposed on the present sample, suggesting that researchers can obtain high-quality data from MTurk workers even without very stringent task approval rate restrictions. It is important to note, however, that this does not imply that all MTurk workers will provide valid data, nor does it eliminate the need to use attention checks or validity scales to ensure attentive and accurate responding. In the present sample, there was a small percentage of cases for each validity indicator that exceeded the empirically derived cutoff scores, a similar finding to previous results using the F scale of the MMPI-2 (Shapiro et al., 2013). However, given that the present data closely resemble the mean T-scores of the community normative sample, the recommended empirical cutoffs for data exclusion reported by the PAI manual (Morey, 1991, 2007) seem equally appropriate for use with MTurk samples.
On average, MTurk workers seem to be representative of the general population across most psychological constructs, with a few exceptions. Notably, the differences observed in the present study mirror the findings of previous studies regarding the distinguishing psychological features of MTurk workers, and the suggestion that MTurk workers resemble the stereotypical frequent Internet user (Goodman & Paolacci, 2017; Paolacci & Chandler, 2014) seems to be supported by the present data. In general, MTurk workers seem to have difficulty with social connectedness, reflected by reports of social isolation, limited social support, and interpersonal coldness and resentment. This finding is aligned with previous work citing MTurk workers as more neurotic and less agreeable (Goodman et al., 2013; Kosara & Ziemkiewicz, 2010) and more socially detached (Miller et al., 2017) than participants from other data sources. Social detachment represents an underlying feature of both autism spectrum disorder and social anxiety, both of which appear to be more prevalent among MTurk workers than in the general population (Arditte et al., 2016; Mitchell & Locke, 2015; Shapiro et al., 2013). The hypothesis of increased social discomfort is further corroborated by lower scores on WRM and a medium effect size difference on ARD-P, suggesting enhanced fears around social situations.
MTurk workers also seem to have higher scores on scales related to negative affect, consistent with the higher neuroticism scores in such samples suggested by the literature (Goodman et al., 2013; Kosara & Ziemkiewicz, 2010). Although previous research regarding the depression and general anxiety symptoms of MTurk workers has been equivocal (Arditte et al., 2016; Shapiro et al., 2013), MTurk workers in the present study had moderately higher scores in depressive symptomatology, particularly pertaining to affective and cognitive aspects of depression, and slightly higher scores in general anxiety. However, there are methodological differences between the current study and the ways in which depression and anxiety symptoms of MTurk workers have previously been assessed. Shapiro et al. (2013) used categorical cutoffs of the Beck Depression Inventory (BDI; Beck, Ward, Mendelson, Mock, & Erbaugh, 1961) and the Beck Anxiety Index (BAI; Beck, Epstein, Brown, & Steer, 1988) to classify individuals as clinically depressed or anxious, respectively, and then compared the prevalence rate of individuals meeting clinical depression or anxiety criteria per the BDI and BAI with the corresponding prevalence rates of major depressive disorder and general anxiety in the general population—prevalence rates that were determined by epidemiological interviews that are not directly comparable to BDI and BAI scores. Conversely, Arditte et al. (2016) compared scores of MTurk workers on the Depression, Anxiety, Stress Scales (DASS-21; Henry & Crawford, 2005) to reported nonclinical and clinical means from two other independent studies, determining that MTurk workers were demonstrating levels of depression similar to the outpatient clinical population referenced and just slightly lower levels of anxiety. These two approaches differ from the current approach (i.e., using national norms), which obtained results that appear to reflect a combination of these two previous findings. Specifically, the present data suggest that MTurk workers do seem to report somewhat greater (but subclinical) levels of depression and anxiety symptomatology relative to the general population; however, the MTurk sample obtained scores on these scales that were appreciably lower than is typically observed in clinical samples (Morey, 2007). Although the present study only details the MTurk sample as compared with the PAI community normative sample, interested researchers could conduct similar comparisons for undergraduate, clinical, or other available norms in the PAI manual (Morey, 1991, 2007).
It should be noted that there are possible alternative explanations for these observed differences, including differences in the administration format of the test (i.e., online versus paper-and-pencil) as well as the age of the normative data, given that the PAI was developed more than 25 years ago. Nonetheless, in light of the consistencies between the present findings and existing studies of MTurk workers, it appears most likely that the differences demonstrated are a function of the unique features of the MTurk sample, rather than the administration format or passage of time. Thus, the generally small MTurk/community differences and the predictability of those differences also support the contemporary applicability of the PAI community norms, in that those norms, although collected decades ago, appear to be broadly representative of the general population in the present day, and applicable to online as well as paper-and-pencil administrations of the test.
One possible limitation to the current findings is that the sample recruited may not be fully representative of the MTurk population as a whole, given the relatively high compensation as compared with other available HITs. The compensation in the present study was designed to be fair to participants while also not being unduly coercive, as some have posed ethical concerns regarding the extent to which MTurk workers are paid fair wages and afforded the same rights as traditional research participants (Gleibs, 2017). Although evidence regarding the effects of compensation level on MTurk data quality appears to be mixed (Aker, El-Haj, Albakour, & Kruschwitz, 2012; Mason & Watts, 2010; Rogstadius et al., 2011), it is nevertheless important that researchers apply the same ethical research standards to MTurk studies as in other more traditional forms of research and consider the potential impact of underpayment on data integrity.
The results of this study support the contention that personality and clinical data provided by MTurk workers is generally of high quality and appears to be representative of the U.S. population as a whole. Some exceptions, particularly around social engagement and negative affect, were noted, and researchers should be attuned to these differences when drawing conclusions from MTurk-derived data. Although this study offers further clarity to the characterization of the MTurk worker population, future research is still needed to delineate the myriad other ways in which this group might differ from the general population, for example, with respect to political beliefs (Berinsky et al., 2012). Such research will be important to further evaluate how any observed differences might affect the findings of research conducted in these samples. Because the use of MTurk and other online data collection platforms will no doubt continue to grow, careful documentation of these samples will enable the field to better understand the potential implications of using this unique research resource.
Footnotes
Declaration of Conflicting Interests
The authors declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: Dr. Morey is the author of the Personality Assessment Inventory and receives royalties from its sale.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
