Abstract
In the organizational sciences, scholars are increasingly using experience sampling methods (ESM) to answer questions tied to intraindividual, dynamic phenomenon. However, employing this method to answer organizational research questions comes with a number of complex—and often difficult—decisions surrounding: (1) how the implementation of ESM can advance or elucidate prior between-person theorizing at the within-person level of analysis, (2) how scholars should effectively and efficiently assess within-person constructs, and (3) analytic concerns regarding the proper modeling of interdependent assessments and trends while controlling for potentially confounding factors. The current paper addresses these challenges via a panel of seven researchers who are familiar not only with implementing this methodology but also related theoretical and analytic challenges in this domain. The current paper provides timely, actionable insights aimed toward addressing several complex issues that scholars often face when implementing ESM in their research.
Keywords
Although studies of intraindividual phenomena in management have been around since the 1930s (e.g., Hersey, 1932), organizational research using experience sampling methods (ESM) has substantially increased in recent years. 1 Indeed, a review of the ESM literature (for overviews of the basics of ESM, please see Beal, 2015; Beal & Gabriel, in press; C. D. Fisher & To, 2012; Ohly, Sonnentag, Niessen, & Zapf, 2010) indicates that over half of the more than 300 articles reported in management-related outlets have been published in just the past six years (Clarivate Analytics, 2018). As reviewed by Beal (2015), ESM studies obtain repeated measures (daily, multiple times per day) of employees’ perceptions of various constructs, with the intent of obtaining the lived, day-to-day experience of employees (Weiss & Rupp, 2011) and minimizing retrospective biases that plague single-timepoint, between-person assessments (Beal & Weiss, 2003). In light of the benefits of examining organizational phenomena from a within-person perspective, scholars have utilized ESM to study a variety of topics, including but not limited to job attitudes (Ilies, Scott, & Judge, 2006); affect, mood, and emotions (Beal, Trougakos, Weiss, & Green, 2006; Gabriel, Diefendorff, & Erickson, 2011; Scott & Barnes, 2011); engagement (Bakker & Xanthopoulou, 2009; Bledow, Schmitt, Frese, & Kühnel, 2011); job stressors, recovery, and well-being (Butts, Becker, & Boswell, 2015; Matta, Scott, Colquitt, Koopman, & Passantino, 2017; Scott, Garza, Conlon, & Kim, 2014; Sonnentag & Zjilstra, 2006); and performance-related behaviors (Gabriel et al., 2011; Sonnentag, 2003; Trougakos, Beal, Cheng, Hideg, & Zweig, 2015). Thus, ESM represents an emerging and rapidly expanding method.
Yet, as with many methodological trends, the increasing popularity of this method has brought about a number of complexities and challenges with which scholars must grapple. For instance, how can ESM be used to advance theory? That is, can ESM be used in a way beyond just testing organizational phenomena at a within-person level of analysis (in comparison to the between-person level of analysis) such that ESM can be used to challenge or construct within-person specific theories? Moreover, how can ESM researchers balance concerns over participant burden with those surrounding the quality of the data obtained from multiple, repeated in situ measurements? Finally, at a more pragmatic level, how should researchers provide evidence for the psychometric properties of within-person measures, and what is the “right” sample size at Level 1 (within-person) and Level 2 (between-person) for an ESM study? Actionable answers and ideas tied to such questions are sorely needed yet currently absent from ESM overviews.
To address these issues, the current paper provides a discussion from a panel of ESM experts—individuals who frequently conduct ESM research and/or review/edit research using ESM techniques in top journals (in alphabetical order: Daniel Beal, Marcus Butts, Allison Gabriel, Nathan Podsakoff, Brent Scott, Sabine Sonnentag, and John Trougakos). 2 These experts are responsible for co-authoring 75 unique ESM studies, 217 cumulative years of editorial board experience, and 32 cumulative years as Associate Editors at journals including Academy of Management Journal, Journal of Applied Psychology, Journal of Management, and Personnel Psychology, in addition to holding officer positions in the Research Methods Division of the Academy of Management. The paper presents 10 questions regarding theory, methods, and analyses in ESM studies, each of which are addressed by two to three panelists. The audience of this manuscript is scholars who have a basic understanding of the fundamental features of ESM but want to further their understanding of complexities associated with the method. When relevant, we also speak to reviewers of ESM studies. Although we do provide references for researchers looking for information on basic issues pertaining to ESM, this is not our focus. The goal of each response is to provide insights into best practices and—when possible—specific and immediately actionable recommendations. We also highlight the benefits and consequences associated with the suggestions being made. Importantly, the panelists did not necessarily agree on all issues, highlighting the complexities that surround the issues being discussed in this paper.
Question 1: How Can We Use ESM Studies in a Way to Build New Organizational Theories That Are Within-Person in Nature and/or Provide Evidence That Challenges Existing Theories?
Beal and Gabriel
Building and challenging theories are topics for which much has been written (e.g., Colquitt & Zapata-Phelan, 2007; Hambrick, 2007; Sutton & Staw, 1995; Weick, 1995). We recently added to this literature by considering how different types of ESM may be used in complementary ways to address within-person theory (Beal & Gabriel, in press). Our goal here is to discuss what might be different for within-person research questions by considering three issues: adopting a within-person perspective, conceptualizing constructs as experiential versus abstract/summarized, and understanding dynamics of experiential constructs.
The first consideration is to ensure that a truly within-person approach is used when conceptualizing the relevant constructs and their connections. Although within-person approaches apply to all longitudinal designs, with respect to ESM, the core of this approach emphasizes the following: For a given individual, change in the status of an event, perception, state, or behavior subsequently influences the status of other events, perceptions, states, or behaviors as experienced by that same individual within a short period of time (e.g., within a workday or episode). Although the core of this approach begins with a purely intraindividual process, the within-person perspective rarely considers these processes in isolation from other individuals within the environment. Thus, this within-person approach to research is inherently multilevel such that basic processes of action and reaction reside within a particular individual’s universe of experiences but can easily be influenced by differences between individuals such that one’s standing relative to others might modulate within-person processes.
This perspective can be compared to theories emphasizing intraindividual differences, in which scholars examine how people differ in their stable aggregations of experiences. Scholars may ask employees to assess their motivation over the last six months and examine how motivation in general affects ratings of citizenship over a similar timeframe. Here, rank orders of motivation are compared to the rank orders of citizenship—a between-person comparison. Yet, many theories describe causal connections or elaborate processes in within-person terms. Indeed, theories are often specified in terms of how an event, perception, state, or behavior yields subsequent reactions (within-person phenomena) but are evaluated between-person (cross-sectional surveys). A truly between-person theory that explains causal connections among constructs is difficult to come by as stable individual differences are unlikely to have a causal influence on other stable factors. As a result, researchers who do not use within-person research methods often are using designs incapable of providing evidence relevant to their theory.
For example, theories that relate personality to work outcomes (e.g., Barrick, Mount, & Li, 2013) rarely state simply that a trait causes scores on an outcome to be higher or lower and leave it at that (such a theory would not be very illuminating). Instead, they suggest that a trait gives rise to a more complex set of processes—processes that unfold over time and capture change relative to each individual’s baseline level of each construct. Despite initial levels of these processes being determined by a stable, between-person variable, they are inherently within-person and should require designs that conform to that theoretical specification. To do otherwise is to commit the same ecological fallacy that sparked several multilevel discussions in the groups and teams literature (e.g., Klein, Dansereau, & Hall, 1994; Klein & Kozlowski, 2000; Rousseau, 1985). Thus, when poorly matched designs have been used to test largely within-person theories, the conclusions that result from these studies may or may not be accurate.
The upshot of this critique is that building organizational theories that are within-person is not the problem, as we often do that. Instead, ESM researchers are faced with uncertainty about whether the existing literature (based largely on between-person designs) accurately captures underlying processes specified by a within-person theory. There is a strong temptation to try to verify whether the current literature has it “right.” As useful as it may be to follow that temptation, rebuilding existing theories and testing them based on evidence that is more accurate is often not seen as “building new theory” and is subject to the common—but often inaccurate—“we already knew that” criticism. This will not always be the case. Certainly, research on existing theories that has relied on mismatched methods (i.e., between-person methods to test within-person theories) can produce effects homologous with those obtained from within-person designs. In such cases, the effort involved in documenting such homology may not be worth what is likely to be seen as a meager theoretical contribution (see Question 2 for this dilemma). In other cases, however, existing theories that have relied on poorly matched methods may be unlikely to produce homologous effects with a more appropriate within-person method. In these cases, where prior research “got it wrong” is an opportunity to challenge existing theory and build new theory. The success of such endeavors depends on the ability of the researcher to specify and test a more accurate within-person process that generates effects unlike what is suggested in current work (e.g., the relation between self-efficacy and performance being negative within-person and positive between-person; Vancouver, Thompson, & Williams, 2001).
Despite many existing theories being based on a within-person perspective (without fully realizing it), it is necessary for current within-person researchers to contemplate theories that move beyond what has come before. There are two insights we can offer to achieve this goal. The first can be gained from understanding the difference between within-person methods that examine experiences versus those that examine abstractions. Experience is the concrete, episodic end of a continuum, which focuses on what it is like for employees to perform, withdraw, strive, interact, collaborate, deviate, or engage with work each day. The other end of the continuum is focused on the same issues, but rather than obtaining an immediate sense of the experience, the goal is to understand how abstractions, summaries, or amalgamations of these phenomena are connected (cf. Weiss & Rupp, 2011). These abstractions fluctuate over time but necessarily do so over longer periods (Robinson & Clore, 2002). Importantly, more traditional longitudinal panel designs are appropriate methods to use to examine constructs that are more abstract.
For example, consider job attitudes. Although attitudes are influenced by more immediate experiences (Weiss, Nicholas, & Daus, 1999), there are large components that are based on more stable beliefs about one’s job. In addition, as summary evaluations, job attitudes typically ask individuals to aggregate their experiences over a time period (e.g., “in the past three months”). There is a great deal of evidence suggesting that such aggregations differ in important ways from reports of more immediate experiences (i.e., the “experiencing self” vs. the “remembering self”; Kahneman & Riis, 2005). Indeed, one area of theory building in organizational research that could be informed by ESM would be to examine the differences between these two perceptions and how more immediate experiences combine to create different abstracted experiences. Some of our work on episodes (e.g., Beal, Weiss, Barros, & MacDermid, 2005) has suggested the possibility that goals help frame immediate experiences into thematically coherent episodes, which are themselves a form of experiential aggregation. New theories that can help link fundamental organizational constructs like job attitudes with raw, immediate experiences could help build a more comprehensive understanding of life at work.
A final insight that can help researchers develop new within-person theories, particularly those that emphasize an experiential approach, is a more thorough understanding of temporal dynamics. Because experiential phenomena are inherently dynamic and emphasize deviations from an individual’s baseline experience, one must consider the temporal patterns that form that individual’s baseline experience. As has been discussed by several authors (George & Jones, 2000; Mitchell & James, 2001; Monge, 1990), many important temporal dynamics are associated with constructs that change over time, including the frequency and magnitude of change, presence of trends and cycles, and duration and continuity of changes, to name a few. In most cases, these issues will be difficult, if not impossible, to determine prior to conducting a study (a point discussed in Question 7). As an example, consider the construct of positive social interaction. If one were to examine this construct from a within-person experiential approach, it would clearly be important to determine how often social interactions occur, whether they are relatively even in their dispersal throughout the day or week, and whether there are any repeating patterns (i.e., trends or cycles) present. In addition, if ESM are to provide the necessary experiential context for which it was designed (Csikszentmihalyi, Larson, & Prescott, 1977; Larson & Csikszentmihalyi, 1983), then we would wish to know over what timeframe a representative range of social interactions could be expected to occur. Before a cogent theory is advanced that links this dynamic construct to another, a firm understanding of temporal dynamics must be developed. A theory without this insight cannot delineate natural fluctuations in a criterion construct from changes that are a function of the social interaction construct.
Question 2: What Is Isomorphism and Homology, and What Role Do They Play in ESM Studies? What Is the Value of ESM Studies That Test Isomorphism or Homology?
Scott and Podsakoff
Isomorphism is the “similarity or one-to-one correspondence between two or more elements” (Bliese, Chan, & Ployhart, 2007, p. 553), and “cross-level isomorphism implies that higher-level constructs have similar meanings and properties as their lower-level counterparts” (Tay, Woo, & Vermunt, 2014, p. 78). Homology refers to “similar relationships between parallel constructs across levels of analysis (Kozlowski & Klein, 2000)” (Chen, Bliese, & Mathieu, 2005, p. 376). Within ESM studies, homology exists if a relationship between within-person constructs is similar in significance, sign, and/or magnitude to a relationship observed between the same constructs at a between-person level of analysis. For example, research indicating that the positive relationship between recovery and affective well-being is consistent across within- and between-individual studies would lead to the conclusion that the relationship is homologous. Conversely, researchers reporting that negative affect and organizational citizenship behavior are negatively associated between persons (e.g., Shockley, Ispas, Rossi, & Levine, 2012) but positively associated within persons (e.g., Scott, Matta, & Koopman, in press) would conclude that the relationship is nonhomologous across levels.
Examining and demonstrating psychometric isomorphism is critical for our ability to draw inferences about whether relationships between constructs are similar or different across levels of analysis because “measurement equivalence needs to be established before predictive equivalence” (Tay et al., 2014, p. 80; see also Vandenberg & Lance, 2000). Certain aspects of ESM are conducive to establishing equivalence between constructs at the within- and between-person levels of analyses. First, ESM research does not run the risk of anthropomorphism. Although describing a work team as being “neurotic” may be anthropomorphic, the same cannot be said when comparing an individual’s neuroticism at a given moment to his or her general tendency to be neurotic. Second, although ESM research can be used to capture compilation processes (Chan, 1998), whereby the aggregate variable assesses a phenomenon distinct from its lower level parts (e.g., within-person measures of emotion transformed into a between-person measure of affect spin; Beal, Trougakos, Weiss, & Dalal, 2013), in many cases, the definition of the underlying construct is isomorphic (e.g., job satisfaction). If scholars conducted multilevel construct validation studies using ESM and existing scales, we could gain a better understanding of which items are more or less psychometrically isomorphic via the estimation of composition models (see Enders & Tofighi, 2007). If there are items that are only relevant when between-persons or over long periods of time (and likewise, when assessed within-persons or over short periods of time), the conceptual definition and content domain of the construct could be altered accordingly at each level. Once an in-depth understanding of psychometric isomorphism is obtained, we can ascertain which within- and between-person relationships are homologous.
Chen et al. (2005) provided several arguments for the importance of examining homologous relationships in the groups and teams literature. However, it is important to note that hypothesizing, testing, and/or reporting homologous relationships may be interpreted as a less meaningful contribution to the ESM literature. Whereas non-isomorphic relationships are highlighted for their novelty and ability to “change the conversation” in a literature, it is more difficult for researchers to make these claims when reporting that a relationship is functionally the same across levels. Thus, contemporary researchers are likely to have a harder time convincing an editorial team that a replicated effect in an experience sampling study, although at a different level of analysis, makes a substantial theoretical contribution. Studies reporting homologous effects may appear to have the characteristics of a Tester (an “empirical article that contains high levels of theory testing but low levels of theory building”) rather than a Qualifier (“empirical articles that contain moderate levels of both theory testing and theory building”) or Expander (“empirical articles that are relatively high in both theory building and theory testing”; Colquitt & Zapata-Phelan, 2007, p. 1286). In turn, this may translate into less scholarly impact in the form of citations for studies hypothesizing, testing, and reporting homologous relationships.
In summary, experience sampling studies that hypothesize or report isomorphic relationships may be evaluated as less capable of making a novel theoretical contribution, which may be interpreted negatively by editorial teams and ultimately result in lower impact when compared to studies reporting non-homologous relationships. Thus, ESM researchers may be better served creating hypotheses and testing intraindividual relationships that do not only mirror relationships between the same constructs observed at the between-person level. Therefore, although we agree with the arguments made by Chen et al. (2005), we believe that the discovery of non-homologous relationships will generally be perceived as more interesting from a theoretical perspective because they contradict existing assumptions and findings (Davis, 1971).
However, it may be more accurate to state that the value placed on isomorphism and homology depends on the purpose attributed to these tests. For example, Chen et al. (2005) argued that “homologous models provide a logical basis from which to start considering multi-level relationships…explicit tests of homology help to highlight domains where inferences of homology are warranted and domains where they are not” (p. 376). We tend to agree with these arguments. Nevertheless, as this domain matures, ESM researchers hypothesizing and reporting homologous relationships will likely bear a larger burden when convincing readers that replicating relationships observed at the between-person level is novel and important. For example, arguments in early experience sampling studies emphasizing the importance of testing known relationships at another level would likely be perceived to have limited value in today’s journals.
We have several recommendations for researchers navigating this challenge. In much of the ESM literature, homology is inferred from comparing the findings of an ESM study to a published between-person study that was likely conducted in a different context with a different sample and possibly different measures. One underrepresented advantage of ESM studies is the ability to simultaneously examine both within- and between-individual empirical relationships with the same participants. This can be accomplished by (a) aggregating within-person measures to a higher level; (b) gathering measures of constructs with a stable component at the beginning, during, or end of the ESM procedure (cf. Gabriel, Diefendorff, Chandler, Moran, & Greguras, 2014); or (c) modeling within- and between-person relationships simultaneously with Level 1 (raw and aggregated) data (cf. Preacher, Zyphur, & Zhang, 2010). This would facilitate tests of isomorphism and homology within the context of a single study as well as allow researchers to test compositional/configural models of their focal variables. Interested researchers should examine approaches by Chen et al. (2005) and Tay et al. (2014) as these efforts could improve empirical understanding and theoretical development of dynamic organizational phenomenon.
Question 3: ESM Studies Often Focus on Maximizing Intraindividual (Level 1) Observations, but Is There a Certain Number of Participants Needed to Ensure That the Level 2 Sample Size Is Sufficient? How Many ESM Scholars Include Power Analyses in Their Work?
Gabriel and Butts
When the issue of statistical power in ESM studies comes up during the review process, reviewers typically ask what the “right” sample size is at Level 1 and Level 2. The best response that we have seen is to provide information of how the sample size in question compares to published ESM work. This response is largely because there are few sample size or power recommendations for ESM studies despite multilevel scholars highlighting this issue (e.g., Aguinis, Gottfredson, & Culpepper, 2013; Mathieu, Aguinis, Culpepper, & Chen, 2012).
To confirm that ESM scholars do not discuss sample size, we coded a set of 107 unique ESM studies reported through 2017 in top-tier journals. 3 Only 2 of the 107 studies referenced using a power analysis to gauge the appropriateness of Level 1 or Level 2 sample sizes. Bakker and Xanthopoulou (2009) conducted a power analysis to determine if their three-level model (days [Level 1; n = 620] nested within people [Level 2; n = 124] nested within dyads [Level 3; n = 62]) would suffice. Although it is not clear whether this power analysis was conducted a priori, it provides a sense of how to interpret a sample size at the highest (i.e., dyadic) level. Huang and Ryan (2011) similarly took what seems to be a post hoc approach, using the Power analysis IN Two-level designs (PINT) program (Snijders & Bosker, 1993; Snijders, Bosker, & Guldemond, 2003) to determine if their two-level model (events [Level 1; n = 998] nested within people [Level 2; n = 56]) had sufficient power. The remaining studies did not discuss sample size decisions. Interestingly, some did talk about why power should—or should not—be a concern with ESM data. For example, Gabriel et al. (2011) noted that “57 nurses is a relatively small sample resulting in low power for Level 2 analyses, there were 342 daily measures, resulting in higher power for Level 1 analyses” (p. 1097). Others shared similar sentiments to Ilies and colleagues (e.g., Ilies et al., 2006; Ilies & Judge, 2002; Ilies, Kenney, & Scott, 2011) by stating that this issue was largely a “Level 2 problem,” meaning that any form of cross-level moderation or mediation would be rather limited and the Level 1 relationships would remain unaffected.
These results underscore that power issues are rarely discussed in ESM research. This is understandable as few studies are underpowered at the focal level of analysis—Level 1. There is an important distinction to make between the primary level of focus in ESM work and discussion of power voiced by groups and teams scholars (i.e., Aguinis et al., 2013; Mathieu et al., 2012; Scherbaum & Ferreter, 2009), the latter of which hinges on ample group (Level 2) observations. ESM analyses hinge on power at Level 1 to test fixed effects or random slopes across lower level effects. Because of this, power analyses to calculate sufficient sample size is relegated to the same degree of relevance as a single-level employee-level study where power analyses are rare.
Keeping in mind the aforementioned points about ESM research, there is an issue of power that does not receive enough attention—small effects sizes in “overpowered samples.” Given the increased resources dedicated by researchers and ease of recruitment provided by social media, sample sizes at Level 1 have grown tremendously. For example, it is not unusual for ESM studies to have a Level 1 sample size nearing 1,000. The possible problem with large Level 1 samples is that a significant p value can be detected even when the variance explained in the outcomes of interest is quite small. In these cases, researchers need to carefully explain why a small statistically significant within-person effect size has theoretical and/or practical value. A classic example of seemingly small within-person effect sizes is the Physicians Health Study (Manson, Buring, Satterfield, & Hennekens, 1991) that concluded aspirin significantly reduced cardiovascular disease using a sample of over 22,000 participants. Although the R2 was merely .001 and equated to a 0.77% risk difference, this 0.77% represents individuals whose lives were saved from heart attacks, making this small effect size incredibly important (Rosnow & Rosenthal, 1989). Therefore, although the burden falls on researchers to explain why small within-person effects sizes are practically important (e.g., Cascio & Boudreau, 2008), we caution reviewers against automatic negative reactions to small within-person effect sizes.
As scholars attempt to maximize Level 1 sample sizes, we encourage these decisions to be grounded in time similar to Question 1: What is the reasonable timeframe the phenomenon of interest will unfold? Is one week sufficient? Or, do you actually need to study the phenomenon for an entire month? Further, it is advantageous to conduct power calculations when Level 2 is of interest (e.g., tests of cross-level main, mediated, and moderated effects). The PINT software (Snijders & Bosker, 1993; Snijders et al., 2003) utilized by Huang and Ryan (2011) makes these power calculations relatively straightforward. 4 Finally, scholars should provide effect sizes. Standardized parameter estimates can be included; however, these coefficients often need to be calculated by hand as parameter estimates are not readily available for random effects models in Mplus—the program many of the panelists use for analyses (see Supplemental Online Material Parts B and C). Scholars can also provide R2 values (often referred to as “pseudo R2”), although standardized effect sizes in multilevel models are not easily determined (for recommendations, see: Aguinis & Culpepper, 2015; LaHuis, Hartman, Hakoyama, & Clark, 2014). 5
These caveats aside, it is fruitful to delineate norms around what is to be expected for sample size at Level 1 and Level 2. Based on our database detailed previously, 90 studies provided Level 2 sample information, and 89 provided Level 1 information (this lack of information is surprising—scholars should clearly report Level 1 and 2 sample size numbers). After removing samples deemed outliers based on z scores above/below 2.0, the mean Level 2 sample size was 83 (SD = 32), and the mean Level 1 sample size was 835 (SD = 475). Further, 70% of studies had a Level 2 sample size of 100 or larger and a Level 1 sample size of 1,055 or larger. Based on this, we recommend that ESM studies aim for a Level 2 sample size of at least 83; for Level 1, 835 is recommended. Of course, sample size recommendations should be treated with a focus on power and effect size sensitivity rather than assuming that larger sample sizes are always better; sample sizes smaller than 835 at Level 1 might be powerful enough to detect expected effects if present, and at some point, little value is gained by substantially increasing Level 1 sample size.
Question 4: The Frequent, Repeated In Situ Measurements Required in ESM Studies Place a Substantial Burden on Participants. In Response, Researchers Are Often Concerned With How to Incentivize Participants to Complete Multiple Surveys in the Specified Time Period. What Incentives Are Employed in ESM Studies, and Which Are the Most Effective?
Podsakoff and Sonnentag
Reviewing the same studies coded in the previous response indicates that many ESM researchers employ some form of incentives. The most common incentives are monetary or quasi-monetary (e.g., gift cards, vouchers, books, iPods, household products). Although 53% of studies reported using some form of monetary or quasi-monetary incentive, there was variability in payouts and procedures for administration. Specifically, researchers reported paying each subject anywhere from $20 to $250 for study participation, translating to between about $1.00 and $12.00 for each (potentially) completed ESM survey. In some studies, participants were paid a flat rate regardless of the number of surveys completed; in other studies, payment was contingent on each completed survey (e.g., Spence, Brown, Keeping, & Lian, 2014) or a specific threshold (e.g., 80% of daily surveys; Bono, Glomb, Shen, Kim, & Kock, 2013; Heller & Watson, 2005). Some researchers offered bonuses for threshold participation rates (e.g., Christian, Eisenkraft, & Kapadia, 2015). In an effort to be frugal and/or motivate interest, other scholars created lotteries in which participants were entered (e.g., Butts et al., 2015; Liu, Song, Li, & Liao, 2017; Matta et al., 2017) and only a portion rewarded. 6
A smaller number of researchers (11%) used nonmonetary incentives. Some delivered developmental feedback to the participants and their organization upon completion of the study (e.g., Barnes, Lucianetti, Bhave, & Christian, 2015; Sonnentag, 2003; Sonnentag & Grant, 2012). In university samples, students were sometimes awarded course credit for participation (e.g., Johnson, Lanaj, & Barnes, 2014; Kammeyer-Mueller, Judge, & Scott, 2009) or recruiting employees to participate (Barnes, Wagner, & Ghumann, 2012). A few research teams explicitly indicated that they did not provide any compensation to participants (e.g., Hülsheger, Alberts, Feinholdt, & Lang, 2013; Hülsheger, Lang, Schewe, & Zijlstra, 2015), but it was surprising that about a third of the ESM studies did not report any information regarding their incentives.
We planned to examine the relationship between incentive type and response rates (which was reported to range between 42% and 99% of all ESM surveys), but unreported data made this prohibitive. It may be important to note that although not a guarantee of participation, the vast majority of studies with high response rates used monetary or quasi-monetary incentives. One interpretation of this is that scholars can pay for a better response rate. However, this option may not be available to all researchers (e.g., doctoral students), and it may unintentionally affect data quality. Indeed, some participants might select arbitrary response options to be eligible for an incentive; accordingly, checking for careless responses is important (Meade & Craig, 2012). In addition, monetary incentives might not be equally attractive for all participants; a modest monetary payment may not be attractive to highly paid, busy managers. Thus, we offer suggestions designed to improve participation rates that do not rely on financial resources and might help convince participants that are not likely to be persuaded by small financial incentives.
Specifically, research in social psychology has demonstrated the efficacy of several influence tactics at obtaining higher compliance (Cialdini, 2009). For example, obtaining a commitment from participants regarding their level of participation before an ESM study begins should increase the likelihood that they will act in a manner consistent with their pledge. Indeed, research on commitments has shown powerful effects on compliance in a variety of contexts and indicates that the effects are typically stronger when commitments are active, voluntary, and communicated to at least one other person (Cialdini, 2009). This suggests that an opt-in survey asking participants to indicate how many of the ESM surveys they will complete, providing a follow-up email indicating that this commitment has been registered, and reinforcing it with reminders could be an effective strategy for increasing response rates.
Considerable research also indicates that observing or hearing about others engaging in a behavior signals that this action is acceptable and expected, increasing the likelihood that people will conform (Cialdini, 2009). Despite the fact that people underweight the influential effect of “social proof” on their own behavior (Nolan, Schultz, Cialdini, Goldstein, & Griskevicius, 2008), the effect of this normative information has been demonstrated across contexts (Cialdini & Goldstein, 2004). Given that most ESM studies we reviewed reported response rates above 50%, participants should be more likely to complete a survey after learning that two-thirds or three-quarters of participants have completed it. The larger the proportion of participants, the stronger the signal that this behavior is normative and socially appropriate, even without monetary incentives. Related to this, researchers can facilitate responding by asking participants to create implementation intentions tied to participation. Implementation intentions are “if…then…” statements that help people develop a plan of action (Gollwitzer, 1999; Gollwitzer & Oettingen, 2011). Research has demonstrated that implementation intentions are a powerful volitional strategy that increases the likelihood that a target will perform a behavior (Gollwitzer & Sheeran, 2006). To improve compliance with ESM protocols that need prompt or contingent responses, participants could be instructed to create implementation intentions such as “If I receive a prompt to answer the ESM survey [or experience a specific event], then I will respond in the next 5 minutes.”
Despite the potential of these tactics and strategies, we could not find one ESM study in the organizational literature that used either. Thus, these techniques may not only improve participation but also provide opportunities for research. Researchers could randomly assign participants to a social proof/commitment, implementation-intention, or control condition and observe the effects of these treatments. Researchers could then explore the effects of these conditions not only on participation rates but also data quality. Thus, research is needed to provide scholars with a better understanding of how incentive techniques increase participation in ESM studies.
Question 5: What Techniques Should Researchers (Not) Use to Provide Evidence for the Psychometric Properties of Within-Person Measures?
Podsakoff, Beal, and Scott
Although considered important in between-subjects research, the concern over psychometric scale properties is magnified when conducting ESM research because measures in these studies are completed repeatedly over relatively short periods of time, often by the same (self-report) source (see Question 10 for more on self-reports). However, the critical underlying question of “validity” does not change from between- to within-person designs—that is, researchers must still ask whether they are measuring what they say they are measuring (Borsboom, Mellenbergh, & Van Heerden, 2004). That said, some techniques used to provide validation evidence do not translate from between-subjects to ESM research designs. For example, ESM research should not use test-retest reliability to evaluate the validity of inferences taken from scores of within-person measures. To test this form of reliability in traditional research, one gathers successive measures of a variable from the same participants (separated by several weeks, months, or years) in an effort to demonstrate that scores on these items are relatively stable (Schwab, 2005). Consistent with classical test theory (Allen & Yen, 2002), high-stability coefficients are often interpreted as indicating that the measures are capturing true score variance as opposed to measurement error. Yet, in the context of ESM, high test-retest reliability is not only unnecessary for Level 1 measures but inconsistent with the assumptions of this method. Indeed, measures with short-term stability are antithetical to the focal purpose of ESM—to capture and predict dynamic variability in important organizational phenomena. This perspective is consistent with experts (e.g., Hox, 2010; Raudenbush & Bryk, 2002) who argue that substantial within-person variability, or low stability, in Level 1 outcomes is a prerequisite for examining ESM data using multilevel analyses. In other words, without considerable within-person variability in Level 1 criterion measures, researchers should neither specify nor test multilevel relationships. ESM researchers typically provide support for this requirement by estimating an intraclass coefficient (ICC1) from the results of a null model to determine the proportion of variance in measures that is attributable to within- versus between-person factors. In sum, ESM researchers should not rely on test-retest reliabilities to assess the validity of inferences made from within-person measures. Instead, the proportion of variance attributable to within-person factors versus the total variance in the measures attributable to within- and between-person factors for all Level 1 focal constructs should be provided.
That being said, it is important to note that several validation techniques used in between-subjects research can be used or adapted to examine the psychometric properties of within-person measures in ESM research. For instance, consider internal consistency reliabilities for multi-item measures. Generally speaking, internal consistency reliability reflects the ratio of a scale’s true score variance (captured through the intercorrelations of scores on multiple items) to its total variance (Geldhof, Preacher, & Zyphur, 2014). Measures demonstrating low reliability cause a variety of conceptual and statistical problems because they bound the upper limit for correlations between measures of two variables and increase the likelihood of Type II errors. Given the long history of internal consistency reliability formulas (Cortina, 1993) and the expectations of editors and reviewers, it is not surprising that over 90% of the ESM studies we reviewed (see Question 2 for sample details) reported some form of internal consistency reliability for Level 1 measures, with the majority meeting standard cutoff values often used.
Despite this seemingly positive evidence for reliability of measures used in ESM studies, there is cause for concern. Specifically, Geldhof et al. (2014) argued that “ignoring hierarchical [multi-level] data structures can bias estimates of inter-item relationships, likewise biasing reliability estimates for a desired level of analysis” (p. 72). This occurs because within-person measures contain true score and measurement error variance at the within- and between-person levels and single-level reliability estimates will not reflect the actual reliability of a set of items unless reliability is identical at both levels of analyses (Geldhof et al., 2014). In response to this concern, Geldhof et al. (2014) introduced a multilevel confirmatory factor analysis (MCFA) procedure for calculating multilevel internal consistency reliability estimates. Based on Monte Carlo simulations examining potential biases in reliability estimates, the authors concluded that “level-specific reliability estimates (i.e., α and ω) are generally preferable to single-level estimates whenever ICCs are non-trivial (i.e., ≥ .05)” (p. 89). Our review of the ESM literature verified Geldhof et al.’s claim that “the need to account for multi-level variability…has been largely ignored in the context of estimating a scale’s reliability” (p. 72): Out of over 100 studies, we did not identify one study that reported multilevel reliabilities. Thus, we encourage researchers to employ the procedures described by Geldhof et al. in their article (Mplus syntax is provided) to estimate and report multilevel reliabilities for their Level 1 measures, identifying the specific type of reliability calculated (see also Cho, 2016).
Because reliability is necessary but not sufficient to establish validity, ESM researchers should use the MCFA techniques highlighted by Geldhof et al. (2014) to also establish the fit of their hypothesized model and the factor structure of and empirical relationships between their latent variable Level 1 measures. Although 45% of the ESM studies included in our review did report results of a CFA, less than one-fifth (n = 20) clearly indicated that they conducted an MCFA. Thus, the majority of these ESM studies did not (a) include strong evidence in support for their latent variable measurement models and (b) effectively account for the nesting of Level 1 measures within participants, indicating that many of the reported CFAs potentially violated assumptions of independence (Li, Duncan, Harmer, Acock, & Stoolmiller, 1998; B. O. Muthén, 1994; Wu, Lee, & Lin, 2018). MCFA allows researchers to specify measures of latent variables at Level 1 and Level 2 and separately analyze within- and between-person covariance matrices. We strongly encourage ESM researchers to conduct MCFA on their latent variable measures and report multiple goodness-of-fit indices (e.g., chi-square statistics, Comparative Fit Index [CFI], goodness-of-fit index [GFI], standardized root mean square residual [SRMR; both within and between], and root mean square error of approximation [RMSEA]), multilevel reliability estimates, factor loadings and significance tests, and latent variable correlations. Researchers conducting MCFA on their hypothesized latent variable factor structure should also estimate plausible alternative models, providing assessments of comparative model fit (Jackson, Gillaspy, & Purc-Stephenson, 2009).
Question 6: A Challenge in Designing ESM Studies Is Balancing the Burden Placed on the Participants and Retaining the Psychometric Properties of the Measures Being Included. How Should Researchers Modify or Adapt Scales to Be Applicable in ESM Contexts? What Evidence Can Researchers Provide for the Psychometric Properties of Short-Form Scales? Should Researchers Use Single-Item Measures of Constructs in ESM Research?
Beal, Scott, and Podsakoff
The enterprise of reducing the length of existing scales has preoccupied psychometricians and behavioral researchers for more than a century (Doll, 1917; Spearman, 1910). Arguably, ESM has more of a need in this respect than other types of designs employing psychological measures. Because ESM require repeated assessment over very brief intervals, participant burden is a concern, and assessing multiple constructs in a given study necessitates the use of extremely efficient measurement approaches. Consequently, when designing an ESM study, researchers often try to identify or create short-forms of existing measures. However, simply stating that a short-form measure was used to ease participant burden is insufficient. Researchers should provide the specific rationale for removing/retaining specific items (e.g., random selection, based on daily variation or lack thereof), detail the technique(s) used to identify items for removal, and indicate which items were removed.
Furthermore, researchers should think carefully about the tradeoffs involved in using short-form scales. For example, in addition to potential reductions in reliability (see Question 5 for more discussion), researchers using shortened measures in ESM studies also must contend with the possibility that the content coverage of the short-form is deficient compared to the original measure as fewer items are less likely to cover the full construct domain. Indeed, many longer scales used in between-persons research are multidimensional or superordinate in nature (i.e., specify a higher-order construct and multiple lower-order constructs; Edwards, 2001; Edwards & Bagozzi, 2000). If an ESM researcher seeks a shorter version of the measure, then the task not only involves careful selection of items from each lower-order construct, but it also must ensure that the higher-order construct is adequately represented as well.
Unfortunately, there are few resources providing formal guidance on how the goal of shortening measures should be achieved. Nevertheless, assuming that a short-form of a given measure that is reliable and yields valid inferences is not readily available in the literature, ESM researchers will find themselves in need of guidance. Of some help are recommended practices for shortening measures in traditional survey research (e.g., Smith, McCarthy, & Anderson, 2000; Stanton, Sinar, Balzer, & Smith, 2002). Yet, there are additional challenges that arise in ESM efforts. Some of these make the process simpler; others make it more complex. For example, because ESM studies capture only a select window of time, researchers often choose to focus on a fairly specific experience that occurs with some frequency (e.g., interactional justice experienced today; Judge, Scott, & Ilies, 2006) rather than a comprehensive array of experiences that might be captured by broadly worded items from traditional survey measures (e.g., numerous forms of justice experience in general; Colquitt, 2001). As a result, shorter scales can more easily be constructed by narrowing the domain from which the items are selected.
At times, however, ESM researchers may choose more broadly worded items in an attempt to reduce the scale. Many measures that are designed to assess stable traits or individual differences have numerous items referring to specific states, behaviors, or experiences (e.g., “obeys company rules and regulations even when no one is watching” from a citizenship behavior measure; P. M. Podsakoff, MacKenzie, Moorman, & Fetter, 1990). Because these specific behaviors may not occur with great frequency in the timeframe of an ESM study, selecting a more broadly framed item can achieve greater coverage of a construct domain with fewer items (e.g., “was one of my most conscientious employees today”). The critical elements to this process are ensuring that the items selected still provide adequate coverage for the domain that is selected—sometimes this can be achieved by choosing a few items from a narrowed domain; sometimes this can be achieved by using a few broadly framed items from a larger domain.
A special case of eliminating items to reduce participant burden arises when researchers use a single item. Single items are not uncommon in ESM research; Conway, Rogelberg, and Pitts (2009), for instance, assessed helping by asking: “Since the last signal, did you voluntarily help someone else [in a way that was not an assigned duty]?” (p. 328). In between-person research, single-item measures are discouraged and viewed as a “fatal flaw” because they are presumed to be unreliable and construct deficient (Cronbach & Meehl, 1955; Nunnally, 1978). However, single-item measures have a place in ESM research depending on the circumstances. As noted by Wanous, Reichers, and Hudy (1997), single-item measures can be categorized based on whether they assess self-reported facts (e.g., age, job tenure) or psychological constructs (e.g., job satisfaction, workplace deviance). There are numerous self-reported facts that may be of theoretical interest to ESM researchers, particularly considering that ESM often capture affective events that impact employee well-being and performance on an episodic basis (e.g., Weiss & Cropanzano, 1996). Thus, the occurrence and duration of events such as work or lunch breaks (cf. Trougakos, Hideg, Cheng, & Beal, 2014), working inside versus outside of the office, listening to music, and attending meetings could all be reasonably assessed with single items.
For psychological constructs, the situation is more complicated and, as discussed previously, depends on the breadth of the construct’s content domain. Beginning with narrower constructs, ESM is frequently used to assess discrete emotions, and single items may not only be easier for participants to understand but can also better reflect the intended content domain. Indeed, a potential criticism of multi-item measures is that they may be contaminated (Scarpello & Campbell, 1983). This contamination is especially apparent with multi-item measures of discrete emotion. For example, shame and guilt differ with respect to their attributions and action tendencies (Lazarus, 1991; Lewis, 2000), yet multi-item measures of guilt include items such as “ashamed” (Watson & Clark, 1994). Thus, if researchers want to assess how guilty an employee is feeling at a given moment, it may be better to present a single item to respondents.
With broader constructs, single items are often deficient. When constructing single-item scales outside of ESM research, scholars often take the item with the highest factor loading from an existing scale. However, measurement items may not be psychometrically isomorphic across levels of analysis; thus, the item with the highest between-person factor loading may not be the item with the highest within-person factor loading (Chen, Mathieu, & Bliese, 2004; G. G. Fisher, Matthews, & Gibbons, 2016). Furthermore, the item with the highest factor loading may be less relevant to within-person episodes. For instance, “reported others for breaking rules or policies” is the highest loading item in Lehman and Simpson’s (1992) measure of antagonistic work behaviors, but the item with one of the lowest loadings—“argued with co-workers”— likely occurs with greater frequency and better captures the construct at the within-person level.
If a single item (or a small set of items) is desired to capture a broad construct domain, there are some novel approaches that might assist in the selection and construction of items. One approach implemented recently in research not using ESM involved embedding items from multi-item measures of these constructs in the description of the scale and asking participants to respond to a single item with that description in mind (e.g., Lyons & Scott, 2012). Thus, this approach first defines the domain of interest for the participants and then asks for a global judgment of this domain. This option has utility, particularly for obtaining independent (e.g., supervisor) reports of key outcomes to avoid same-source bias (see Question 10). Another approach is to consider graphical or information-rich items that might convey more information than a standard item. For example, the affect grid (i.e., Russell, Weiss, & Mendelsohn, 1989) assesses both dimensions of the affect circumplex from a single item. A single-item pictorial measure of Other in Self (Aron, Aron, & Smollan, 1992) effectively assesses identity-related constructs, and subjective socioeconomic status has been frequently and effectively measured using a pictorial “ladder” item (Ostrove, Adler, Kupperman, & Washington, 2000).
Regardless of the approach taken to construct a short-form or single-item scale for use in an ESM design, there is likely a need to obtain evidence that the scale or scales that have been modified can generate valid inferences. At times, this need may not rise to the level of a separate validation study (e.g., if the domain being assessed is extremely narrow and time-bound, such as immediate levels of shame or other discrete emotional states). However, as the domain being assessed broadens, efforts to create short-forms will be best served by an accompanying validation. Fortunately, there are relatively straightforward approaches for this sort of effort. For example, content analyses of the reduced set of items can be achieved with relative ease using established procedures (Hinkin & Tracey, 1999; MacKenzie, Podsakoff, & Podsakoff, 2011) to determine if and how the reduced measures are deficient. In addition, a validation study of the reduced set of items could be conducted using an accessible sample (e.g., subject pools, MTurk, online panels). In sum, efforts to ensure that the revised scale yields valid inferences are needed to simultaneously advance ESM research and ensure that participants are not overly burdened.
Question 7: If, When, and/or How Should a Researcher Go About Testing Data for Possible Trends, and What Trends Should Be Tested (e.g., Linear/Quadratic Trends, Between-Day and Within-Week Trends, Within-Day Trends)? Relatedly, How Does One’s Research Question Shape the Answers to These Questions, and How Can It be Used to Improve the Design of the Study?
Beal and Trougakos
The examination of trends in traditional longitudinal research reflects a theoretical emphasis on growth, learning, or other developmental processes (Singer & Willett, 2003). Despite the differing lengths of time over which development takes place in traditional longitudinal studies, most studies are designed to run the length of a single process as it gradually unfolds. As such, modeling a linear or quadratic trend is well matched to the goals of these studies. Considerable work examining longitudinal research methods offers several options for capturing development over different time lengths (Biesanz, Deep-Sossa, Papadakis, Bollen, & Curran, 2004; B. O. Muthén & Muthén, 2000b; Rogosa & Willett, 1985; Singer & Willett, 2003).
ESM studies approach development differently. First, it is rare that an ESM study is designed to capture a single instance of a developmental process. Because ESM studies emphasize the experiences of individuals as they occur in their natural environment, participants typically are not all experiencing the same type of developmental process. Therefore, a linear or quadratic trend over the entire length of the study will not detect a common developmental process. We speak in generalities as it would be possible to design an ESM study to capture a developmental phase (e.g., the first week of employment; experiences before, during, and after a training program; employees in a critical—yet brief—phase of employment). Indeed, a focused analysis of a particular developmental experience would be a welcome, insightful use of ESM.
That said, if a typical ESM study in which participants had no common developmental experience examined a daily linear trend spanning the entire course of a three-week study (e.g., starting at 0 and increasing each day until the end of the study), the magnitude and significance of the resulting trend would be difficult to interpret. Indeed, significant trends of this sort might often signal something akin to a Hawthorne effect, where the ESM itself was generating a patterned response for all participants. Such effects might be indicative of other problems in the data collection. For example, if participants grew weary of the daily survey, then they might begin to respond in an increasingly careless manner, which could result in a general weakening of effects over time due to increasing unreliability (e.g., Meade & Craig, 2012).
We realize that it may sound as if we believe that development does not occur in a typical ESM study. Quite the opposite is true, however, as we believe that short-term development is very likely to occur for most—if not all—individuals in every ESM study. Trends in ESM studies will often reflect a process that repeats over predictable intervals. For example, energy levels might exhibit a decreasing pattern over the course of a single day, only to repeat itself the next day (Benedetti, Diefendorff, Gabriel, & Chandler, 2015). Trends such as these might be driven by a combination of circadian rhythms (Saper, Lu, Chou, & Gooley, 2005) and the particular work context (e.g., a job where tasks pile up during the day), whereas other trends might be more attuned to broader cultural or normative factors (e.g., a trend that begins on Monday and ends on Friday, repeating each week; Zijlstra & Rook, 2008). The point is that these trends are developmental processes very much like those captured by growth trends and curves in longer-term studies (e.g., B. O. Muthén & Muthén, 2000a).
The perceptive reader may have noticed that we have been referring to various effects in ESM designs as discontinuous-but-repeating linear trends while also mentioning an underlying process that is more cyclical in nature. In truth, we suspect that most of the linear trends that can be observed in ESM studies of workplace experiences are actually abbreviated segments of cyclical processes. Because we often assess participants neither outside of normal working hours nor on non–working days (an issue briefly touched on in Question 8), we miss capturing the remaining segments of the cyclical process. In some cases, such as when a cycle that repeats daily exhibits a monotonic increase (or decrease) during work hours (Benedetti et al., 2015), modeling a linear or quadratic trend during working hours may suffice to capture this portion of the cycle. In other cases, a cycle might peak at midday, producing an inverted U pattern during working hours rather than a linear trend (Debus, Sonnentag, Deutsch, & Nussbeck, 2014).
It can be very important to model these trends or cycles, particularly when theory suggests their presence (Beal & Weiss, 2003). Liu and West (2016) examined this issue specifically within the context of weekly cycles in ESM studies. Using an existing data set, they first demonstrated a not uncommon situation where a particular predictor-criterion effect was detected with no weekly cycle included in the model, only to have the effect disappear when the cycle was included. A more formal investigation used simulated data and determined that omitting cycles from models examining typical predictor-criterion effects frequently resulted in both increased Type I error rates and biased estimates of the effect. Type I error rates increased wildly (larger than .40) when cycles were omitted and the underlying actual cycle was large in magnitude. Bias in the estimated effect exhibited a more complex pattern arising from the combination of several factors, including the magnitude of the predictor-criterion effect, whether the cycles of the two variables were synchronized or not, and the magnitude of the cycles themselves. Of note were findings indicating that the bias could, in some cases, result in large underestimation or large overestimation of the actual effect even when the true effect was zero. Perhaps most importantly, including the cyclical patterns in the models accurately recovered the true effect and maintained acceptable Type I error rates in all conditions. Thus, based on this initial research, if a cyclical process exists, leaving it out of the model can create severe interpretational problems, whereas including it in the model appears to have no downsides. Of course, this recommendation is based on a single simulation study and is perhaps premature. Future research might examine, for example, if the inclusion of a trend or cyclical pattern when no true trend or cycle exists might create its own interpretational problems.
Modeling trends and cycles is a common practice in econometric time series analysis (Box, Jenkins, Reinsel, & Ljung, 2015). Typically, there is a single unit generating the time series in these analyses (e.g., a nation’s economic indicators over time). In ESM, individuals are the units, and so there are as many time series as there are people in the sample. Consequently, modeling these effects in ESM data reflects only the average trend or cycle. It is often the case that the coefficients capturing the effects of trends and cycles in ESM data vary significantly across individuals (i.e., exhibit significant random effects). This variability might reflect random variations in the patterns of development for people, but it is also plausible that this variability is systematic, meaningful, and predictable from other individual-level variables (e.g., Beal & Ghandour, 2011). Our recommendation is that researchers think carefully about how patterns of development might vary across individuals and examine this possibility statistically (see the following).
It should be clear by now that trends and cycles represent complex patterns of short-term development in ESM studies. We see four possibilities that researchers should consider. First, trends and cycles could reflect interesting and important processes directly related to theories being examined in the ESM study (Shipp & Cole, 2015). Second, trends and cycles could reflect confounding effects that should be taken into account to provide a clearer interpretation of the effects of interest (Liu & West, 2016). Third, in either of these two cases, trends and cycles might reflect averages around which exist important individual differences in developmental patterns. Fourth, despite the repeated measurements obtained in an ESM study, trends and cycles might simply not exist for the particular predictor or criterion variables included in the model being tested. Although formal testing of effects for trends and/or cycles would be necessary for the first three possibilities, the work by Liu and West (2016) provides initial evidence, with respect to the fourth possibility, that testing for their presence is beneficial. If such tests suggest that trends and cycles (or random effects around the trends and cycles) are not present or their presence has no impact on the effects of other variables in the model, then leaving them out of the final model would be a more parsimonious approach. We should also note the possibility that a fixed effect for a trend or cycle may not exist even though a significant random effect does (i.e., significant variation around a nonsignificant average effect). For example, if some people exhibit a pattern of positive daily growth and others exhibit a similar but inverse pattern of daily growth, then the average trend may be zero. As one can hopefully see from this example, in these cases, it would be useful to include fixed and random effects for the trend or cycle in the model.
The question remains as to how you should test for such effects. Linear or quadratic trends—much like their counterparts in growth curve modeling—involve regressing the outcome variable on a predictor that linearly (or quadratically) increases or decreases over the period of development (e.g., each day, each week; see Question 8). As Liu and West (2016) note, there are several ways to model cyclical effects. Their preference and our own (Beal & Weiss, 2003; West & Hepworth, 1991) has been to use a predictor variable that produces a full cycle over the period in question (e.g., each day, each week), along with a second predictor variable that captures a similar cycle that is opposite in its phase. Thus, a sine wave that repeats each week (calculated by
Question 8: What Sources of Common Method Biases Are Typical ESM Studies Subject to? Does Centering Scores on Level 1 Variables Affect These Biases, and If So, How? Are Other Remedies Available, and How Effective Are They at Addressing Sources of Potential Common Method Biases?
Scott, Sonnentag, and Podsakoff
Because the majority of ESM studies gather both predictors and outcomes from the same (self-report) source, reviewers and readers often voice concerns about potential common method biases (CMB). CMB refers to distortion in empirical estimates of bivariate relationships due to the fact that the variables share a method or a source provides ratings for both variables (P. M. Podsakoff, MacKenzie, Lee, & Podsakoff, 2003; P. M. Podsakoff, MacKenzie, & Podsakoff, 2012). Most often, ESM researchers indicate that concern over CMB for Level 1 relationships is addressed by person-mean centering (referred to as group-mean centering in the broader multilevel literature; Hox, 2010) the scores on these variables. As several authors have pointed out (e.g., Enders & Tofighi, 2007; Hofmann, Griffin, & Gavin, 2000; Raudenbush & Bryk, 2002), centering around an individual person’s mean in ESM yields an unbiased estimate of the within-individual (Level 1) regression coefficient, whereas grand-mean centering or leaving predictors uncentered produces Level 1 coefficients that blend within- and between-person variance. Consequently, it is recommended that researchers use person-mean centering when utilizing ESM to analyze within-person relationships, either explicitly or implicitly by specifying a two-level model with (1) the model of interest at the within-person level and variances at the between-person level or (2) identical paths being modeled at the between- and within-person levels (L. K. Muthén & Muthén, 1998-2015). This decision depends on the research question.
The use of person-mean centering has important implications for the influence of between-person characteristics that are often identified as potential sources of CMB (P. M. Podsakoff et al., 2003). When predictors have been centered within cluster, each person’s (i.e., cluster’s) mean on that predictor is zero. If every person’s mean is zero, there is no variance between persons and differences between people are eliminated. Given that confounds are variables that have a causal impact on both the predictor and outcome (Cohen, Cohen, West, & Aiken, 2003), all forms of between-person differences (e.g., demographics, personality, response tendencies, social desirability) are effectively controlled because they cannot correlate with predictors that have been person-mean centered. Consequently, reviewer criticisms that factors such as trait negative affectivity or social desirability are responsible for within-person relationships are invalid so long as group-mean centering was used. However, this is true only with respect to between-person factors that have mean-based effects on Level 1 outcomes; this is not true of factors that have variance-based effects (see extreme response style as a potential example; e.g., Weijters, Geuens, & Schilleweart, 2010; Weijters, Schilleweart, & Geuens, 2008). It is also important to note that person-mean centering is not a panacea for all types of hypothesis tests or sources of potential CMB. For example, if a researcher’s goal is to examine the influence of a Level 2 predictor on a Level 1 outcome (i.e., an intercepts-as-outcomes model) and wants to control for the influence of a Level 1 predictor that has been group-mean centered, then the Level 1 predictor’s cluster average must be introduced as a Level 2 control variable (e.g., Enders & Tofighi, 2007).
With respect to other sources of potential CMB, researchers should recognize that person-mean centering does not control for potential biases due to transient mood states, affect, or discrete emotions (P. M. Podsakoff et al., 2003). When designing ESM studies, researchers should consider the inclusion of control variables that capture these states (see Question 7 for trends/cycles as controls). Given that many ESM studies focus on affective processes and are based on self-reports, researchers should consider controlling for day-specific mood (e.g., mood in the morning). Because mood can have a profound impact on perceptions and appraisals (Schwarz, 2012), it may be the driver of empirical associations between measures typically assessed in ESM studies. Ideally, researchers should take both the valence (i.e., positive vs. negative mood) and arousal dimension (i.e., high activation vs. low activation) of mood into account (Russell, 1980). Depending on the specific research question, researchers should control for positive activated mood and/or negative activated mood, for instance, by using items from the Positive and Negative Affect Schedule (PANAS; Watson, Clark, & Tellegen, 1988). In addition, researchers may need to control for low-arousal states such as fatigue or serenity. Because low-arousal states are not represented among PANAS items (Mossholder, Kemery, Harris, Armenakis, & McGrath, 1994), researchers may want to use other measures, such as measures from the PANAS-X (Watson & Clark, 1994) or the Profile of Mood States (POMS; McNair, Lorr, & Droppelman, 1971). The inclusion of fatigue might actually be more important than serenity because it can capture exhaustion processes due to respondents’ circadian rhythms or day of the week (Hülsheger, 2016). In sum, when all ESM data come from the same source, controlling for emotions or mood can help address concerns about state-based factors as sources of CMB (P. M. Podsakoff et al., 2012).
However, there are at least two caveats to consider when using this remedy. First, assessing various mood states increases survey length and, in turn, participant burden. Particularly in situations where survey length is an issue and mood is only used as a control variable, researchers should assess this variable with shortened scales or single items. Generally, single-item measures are problematic because of unknown reliability and should be used with caution (see the response to Question 6 as to when single-item measures may be acceptable). Future research should strive to gain a sound understanding of which specific single items best capture specific mood states. Second, it is important to note that there has been debate in the between-subjects literature about whether affect or mood should be considered controls or substantively important variables (e.g., Judge, Erez, & Thoreson, 2000; Spector, Fox, & Van Katwyk, 1999; Spector, Zapf, Chen, & Frese, 2000). Although a full account of this debate is beyond the scope of this paper, consistent with recommendations for potential control variables (Becker, 2005; Spector & Brannick, 2011), we encourage ESM researchers to (a) provide a good conceptual argument for the inclusion of mood states as controls and (b) conduct and report anlayses with and without these variables included.
Another option for addressing the potential biasing effects of mood is to analyze lagged relationships between variables, removing the possibility that transient states bias Level 1 relationships altogether. Although ESM data are necessarily longitudinal, a very small portion of the published articles in this domain explicitly test lagged relationships between predictors and outcomes. In many ESM studies, the analyses capture repeated cross-sectional surveys. Testing lagged relationships has the added benefits of addressing concerns over one form of measurement characteristic as a source of CMB (P. M. Podsakoff et al., 2003) and increasing the strength of the causal inferences ESM researchers can make as it helps address temporal precedence. Researchers should also decide whether to control for the previous time-period measure of their outcomes. For instance, when researchers want to predict OCB in the afternoon, should they control for OCB in the morning? This decision depends on the research question: Do researchers want to predict the level of the outcome contingent on specific events, or do they want to predict change (e.g., morning to afternoon) contingent on specific events? As noted previously, predicting change in mood from morning (t – 1) to a later point during the day (t) could be an obvious aspect of a broader research question. However, predicting the level of specific social interaction patterns on specific days, for instance, might appear to be equally interesting as predicting morning to afternoon change in the social interaction patterns, although such an approach limits the possibility to draw conclusions about causality. In addition, researchers must decide if they want to control for the outcome variable on the previous day. Although it makes sense to take autocorrelations into account (Beal, 2015), in many ESM studies, control variables assessed on the same day have more predictive power than control variables from the previous day (Gabriel, Koopman, Rosen, & Johnson, 2018; Sonnentag & Starzyk, 2015).
Although there are exceptions, a typical ESM study consists of obtaining one or more surveys per day for a two- to three-week period. Accordingly, one question that arises is whether observed relationships are influenced by circadian rhythm. Work by Watson (2000) has shown that mood exhibits both diurnal and weekly variation. Specifically, positive affect rises and then declines throughout the day, and it steadily increases as the week progresses from Sunday to Saturday. Although negative affect exhibits little systematic variation within a given day, it steadily decreases as the week progresses. While some research using ESM has controlled for the time of the day that daily surveys were completed and the day of the week (e.g., using six dummy codes to represent each day) to account for potential circadian rhythm effects (e.g., Scott et al., 2014), we do not necessarily agree with the recommendation of Beal and Trougakos (Question 7) to always account for circadian rhythm. Instead, researchers should include moods (e.g., positive and negative affect) either as controls or substantive variables given that such states are proximal outcomes of circadian rhythm effects and would consume fewer degrees of freedom. We acknowledge, however, that researchers should be aware that the outcome of interest in this case is not the “raw” outcome but rather the residual of the outcome controlling for mood. Ultimately, the implications of accounting for factors such as circadian rhythms, cycles, and moods (in isolation or in conjunction) is something future research should address.
Finally, with respect to CMB, it is possible that similarity in item characteristics across Level 1 variables could bias relationships (P. M. Podsakoff et al., 2003). Although N. P. Podsakoff, Whiting, Welsh, and Mai (2013) have demonstrated that similarity in anchor points for between-person measures can inflate empirical relationships, we are not aware of any research that examines this in the context of ESM research. Thus, it is not clear how important item characteristic effects may be in the context of these research designs. Some might argue that wary ESM researchers should vary the response formats and/or anchor points of their focal variables. However, Beal (2015) has cautioned that this “remedy” can increase participant confusion, irritation, and burden. Unfortunately, without empirical evidence from primary ESM studies that explicitly examine the effects of varying item characteristics or other sources of CMB, we cannot make strong recommendations for ESM researchers regarding these biases or potential remedies. Thus, we recommend that researchers create systematic, rigorous tests for the effects of multiple sources of CMB in the context of ESM research and examine the potential effects of different procedural and statistical remedies (e.g., person-mean centering, including direct measures of transient states as control variables, estimating cross-sectional and lagged relationships, and varying item characteristics). This research will help the entire domain of ESM research and address a very common criticism in this domain.
Question 9: Are There Any Consensus Recommendations for if, How, and When Data From Subjects Starting the Study on Different Days Should Be Treated?
Sonnentag and Trougakos
When planning an ESM study, researchers have to decide about the periods of data collection, focusing on which days participants should start and finish responding to surveys. Often, employees’ work schedules dictate this: For jobs with typical workweeks, data collection starts on a Monday and ends on a Friday. These periods fit nicely with the logic of a workweek and can be easily communicated to participants. They can also be easily scheduled and programmed in the case of electronic or online surveys.
Despite these advantages, this approach comes with a notable drawback: Days on which data collection occurs are confounded with days of the week because for all participants, the first day of data collection is a Monday, the second day is a Tuesday, and so on. For instance, when testing if responding to an ESM survey alters the experience or behavior studied, researchers should examine possible trends in the data (Barta, Tennen, & Litt, 2012) related to the day data are collected and if the day is associated with the focal experience or behavior studied. When collecting ESM data for at least two weeks, researchers should tease apart the potential impact of day of data collection from the day of the week by creating a continuous variable for the day of data collection and several dummy variables for the day of the week. This approach also works well when including participants with irregular workweeks. For instance, there might be people who work only three or four days a week, with some working Monday, Tuesday, and Friday; others working Tuesdays, Wednesday, and Thursday; and so forth. Creating separate variables for day of data collection and weekdays allows researchers to examine what the impact will be on the patterns of relationships between substantive variables when participants are not working on one or more days during the week. As an example, the negative impact of a high workload might be lower on a Thursday or Friday when Wednesday was a day off; such a question would build on ESM work that has considered recovery in the evenings predicting well-being the next workday (e.g., Sonnentag, Binnewies, & Mojza, 2008).
Because ESM studies try to capture short-term fluctuations in experiences and behaviors throughout the workday (and sometimes during off-job time as well), it is essential that participants respond to the surveys at specific timepoints or after specific events occur that are relevant to the research questions of interest. Usually, researchers send reminders at specific points of the day that prompt participants to answer the survey. This approach again works well for participants who work regular work hours, for instance a prototypical “9 to 5” day. However, this “one size fits all” approach is not feasible for participants working nonstandard hours or different types of shifts (e.g., morning shifts in one week, night shifts in the second week, and afternoon shifts in a third week). When including participants who work nonstandard hours and the nonstandard hours are not identical for all participants, researchers will need to send individualized reminders that take the individual working times into account. Thus, the reminder for the morning survey will not be sent at 7 a.m. to everyone but at 6 a.m. to one subgroup of the participants and 7 a.m. to another group. In other word, surveys and their scheduling will need to be tailored to each participant’s specific work schedule. In terms of data analysis, this approach offers two options. First, when interested in experiences around a specific event (e.g., start of the workday), data assessed at different times of the day (clock time) can be combined as long as they refer to the same events in the items/instructions. Second, when interested in the time course of specific events, a separate variable could be created that refers to the specific timepoint when the variable was assessed. This new variable could be the clock time of data collection; alternatively, researchers might create a new variable that captures the time since awakening, for instance when interested in chronobiological processes.
Data collection becomes even more complex when sampling shift workers. Accordingly, ESM studies with shift workers are relatively rare (cf. for studies that include shift workers: Kammeyer-Mueller, Simon, & Judge, 2016; Totterdell, Spelten, Barton, Smith, & Folkard, 1995). When sampling shift workers, researchers may want to assess data only during morning shifts (Baethge & Rigotti, 2013; Blanco-Donoso, Garrosa, Demerouti, & Moreno-Jiménez, 2017). However, this approach limits the generalizability of the findings because the shift type might have an effect on outcomes (Korunka, Kubicek, Prem, & Cvitan, 2012) and valuable information from other shifts might not be considered. Alternatively, researchers might want to tell participants to answer the surveys at the beginning and end of the shift irrespective of the time of day (Bass, Linney, Butler, & Grzywacz, 2007; Kammeyer-Mueller et al., 2016). Additionally, when research questions focus on event-specific questions—for example, a specific type of interaction with a leader, the experiences of certain emotions such as anger, or a critical event of theoretical interest—this can provide direction for participants as to when to respond to surveys. This approach reduces the concerns associated with variability in shifts as specific experiences during work shifts (or after) are the focus of data collection and thus serve as the cue to respond rather than the time of day or a time during a shift. It has to be kept in mind, however, that such events might have a low base rate or participants might not be able to or forget to respond during these times (Reis & Gable, 2000). Accordingly, the data collection period likely needs to be longer, and more frequent participant reminders might be necessary. Or, utilizing a combination of these approaches might be necessary. For example, it can be useful to have surveys at the start and end of shifts in combination with surveys related to specific events during shifts (e.g., Beal et al. [2013] examined busy rushes for restaurant servers during shifts while administering surveys at the start and end of shifts). When collecting data multiple times during a work shift such as this, an additional linear control variable should be created, increasing from the beginning to the end of the shift, thus accounting for the impact of duration of shift.
Finally, many of the aforementioned recommendations apply to data collections with employees working irregular schedules, such as those in health care, food, or entertainment industries where employees may work weekends. For instance, many nurses work four 12-hour shifts followed by five days off, with this cycle or some variation repeating. In these instances, controlling for day of the week, day of the data collection, or number of days between providing data helps account for variance related to this type of scheduling. Moreover, when researchers are interested in participant experiences between workdays, a dummy code for work versus non-workday in addition to day of the study can account for these variations. This approach will allow for analyzing complex configurations of workdays and non-workdays.
Question 10: ESM Scholars Are Using a “Standard” Approach to Collecting Intraindividual Data (e.g., Collecting Three Self-Report Surveys a Day for 10 Workdays). However, Scholars Are Starting to Shift Toward Alternative Data Sources (e.g., Other-Reported Measures of Dependent Variables). What Are Considerations in Collecting Data From Other Sources? How Can They Be Used Beyond Having Another Source of Data for an Outcome Variable?
Trougakos and Gabriel
In reviewing the ESM literature, the addition of secondary sources is often focused on decreasing critiques from reviewers who might be uncomfortable with same-source data (which comprises the vast majority of current ESM research) or would prefer a second source of information. Adding a secondary source was especially important in early ESM studies when few organizational scholars (i.e., future reviewers) were familiar with ESM. As such, ESM researchers had legitimate concerns that a reviewer might not appreciate the nuance of a same-source ESM design, increasing the difficulty of an already challenging process. Thus, adding secondary data sources, like in non-ESM research, is a good way to strengthen one’s research if done properly and should primarily be driven by the research questions at hand.
Importantly, our intent is not to suggest that other-ratings are always needed or even better. In fact, as we point out later in this section, self-report data are often not only sufficient but also may represent the best data to answer the research questions in a given study. There are key decisions to determine if data from another source are necessary and what constructs should be assessed. For instance, additional sources might be vital if the phenomena cannot be captured accurately or fully from focal participant self-reports alone. For example, when a model focuses on a focal participant’s observable behaviors toward others, such as citizenship and affective displays, these actions might be captured best by the targets of these behaviors. Additionally, it may be necessary to capture experiences shared by others interacting with the focal employee such as work-family conflicts between spouses/partners or service encounters with customers; in these instances, there are clearly multiple sides to the story that a researcher should examine. It would likely be hard, for instance, for a scholar studying emotional labor to argue that customer performance is a daily outcome of emotion regulation and rely on self-reports from employees of how effective they were in customer interactions. Thus, the chance to gain a comprehensive understanding of one’s model might best come from perspectives other than the focal participant.
Beyond coworkers and supervisors as secondary sources, there is a growing set of options to incorporate technologies, such as wearable physiological monitors, sociometric badges, and movement or affect monitors (for a review, see Chaffin et al., 2017); as an ESM example, Bono et al. (2013) had employees wear blood pressure cuffs during their 15-day study. While these technologies are still developing, they offer a host of opportunities to capture within-person data. In fact, we would argue that pairing physiological data with ESM data is more appropriate than pairing such measures with cross-sectional assessments; in aggregate, it is hard to discern what is affecting variability in physiological data. As an example, consider heart rate. Normal values of heart rate do not mean much at the between-person level of analysis because each individual’s baseline is different. This means that while the normal range for resting heart rate is between 60 to 100 beats per minute, many people—from competitive athletes to someone serious about working out intensely—can have resting heart rates as low as 40 beats per minute. If someone who has an average resting heart rate of 50 is experiencing a heart rate of 100 in an episode, this would be extremely high for that person, but this information could be lost when aggregated across days. Comparatively, for someone with a regular rate of 90 who then has a rate of 100, this would unlikely be notable for that individual. Thus, for a researcher studying work stress, a one-time measure of heart rate where two people are at 100 beats per minute may not show a significant relationship with a one-time measure of burnout. However, when examined at the within-person level, it will be clearer the person with the resting heart rate of 50 experienced higher arousal on a day he or she had a heart rate of 100 while the person with a typical rate of 90 did not. These physiological measures vary greatly within-person and, if measured correctly, provide a rich source of secondary data.
That said, although we are advocating for other sources of data, there are times that the incorporation of these variables may not be as fruitful. Take, for instance, a rating of helping on a given day from a coworker. On a day where helping ratings are low, it is impossible to determine how comprehensive that coworker evaluation is. First, did the coworker have the opportunity to observe helping enactment? Researchers should always ask how much interaction coworkers had with the focal person in the study. Second, even if the coworker did indicate that they observed the focal participant, this is just one rating. It is possible that the focal employee helped others at work and the coworker providing the ratings just happened to not see those interactions or be the direct recipient of the helping. This would make the coworker rating somewhat deficient. As such, while secondary data can be helpful, the data need to be interpreted and explained with some caution. This example does highlight, however, why secondary data in ESM studies are unique compared to such ratings in a cross-sectional design: In an ESM data collection, you likely see variability in ratings one day to the next, whereas in a cross-sectional design, these day-to-day changes are collapsed when coworkers rely on a mental aggregation of their interactions with the focal participant over a longer timeframe (e.g., over the last month, over the last year).
Additionally, the incorporation of other sources can become homogenous. For instance, it is easier to add a coworker survey at the end of the workday given that coworkers were likely recruited by the focal employee and their motivation may be lower (along with the corresponding payment they are receiving) to participate. Because of this, one quick survey at the end of the day may give researchers the most return on investment. Yet, these rating sources are then solely used as outcomes, which seems limiting given that other-ratings of outcomes is what is typically done in between-subject research (e.g., having a measure of supervisor performance at the end of the year). Instead, coworkers would be well suited to rate shared conflict in the office at the start of the day, a demand that was unexpected (e.g., bad news being sent out to a work team about a project they were focusing on), or the mood that employees started the day in. These ideas represent new avenues if the goal is to couch an ESM collection as necessitating secondary data.
Taking this a step further, secondary sources need not only be within-person. Coworker, supervisor, and significant other ratings might serve as cross-level predictors of focal employees’ within-person experiences. It may be easier logistically (and cost-wise) to have these individuals complete a one-time survey at the start of a study. This is not only convenient for participants and researchers but also addresses interesting theoretical questions. This approach could provide insights into daily self-reported behaviors and experiences in relation to other-reported ratings of job characteristics, work relationships, and support or expectations from supervisors, coworkers, or spouses. This could help address reviewer comments about perception versus reality of work.
If incorporating secondary data is not feasible, there has been an increased use of field experiments or interventions in ESM studies (e.g., Foulk, Lanaj, Tu, Erez, & Archambeau, 2018; Song et al., 2018). These can provide powerful tests of hypotheses, and if done well, they can effectively support causal inferences as you can have within- and between-person randomization of your conditions. Of course, standard practices of experimental research should be followed (e.g., random assignment, utilizing a control group). As a caveat, this work can be very challenging as trying to devise an intervention that would work multiple times over the course of a study along with a suitable control condition could be difficult (e.g., If doing a gratitude intervention, will participants generate new things to be grateful for daily?). The repetitive nature of ESM also increases the likelihood that participants figure out the purpose of a manipulation or others in the workplace speak about conditions, contaminating results. Because of attrition rates associated with ESM, there is likely a need for a large sample. Needless to say, the challenges to this approach add to the risk already associated with the time-consuming and costly nature of ESM studies (it is easy to accrue costs of up to $7,000 to 10,000 in payments).
Finally, secondary data do not always make sense. For instance, the use of same-source data is perfectly acceptable when we are interested in the experiences of the focal individuals or if the phenomena of interest are such that only the focal participants would be privy to changes in them. Events, emotional states, thoughts, perceptions, and behaviors are often captured most accurately by directly asking the people experiencing them day-to-day. Because of this, a single source of data is sufficient when examining how someone experiences an event earlier in the day and how this impacts their own attitudes or emotions later in the day (or the next day). The key here is the variability in someone’s own experiences, states, attitudes, or behaviors as compared to himself or herself. For example, several ESM studies exist on daily perceptions leaders have at work (Johnson et al., 2014; Lanaj, Johnson, & Lee, 2016). If the focus is on how leaders experience their environments—from shifts in leader identity, to the behaviors they are enacting—single-source data are likely best. As such, we offer a critical caveat, particularly to people reviewing ESM research: Do not have a knee-jerk reaction that secondary data are needed. This may be a key critique for cross-sectional data, but in a well-designed ESM study in which scholars temporally separate constructs and control for prior assessments to isolate change, we can see the value in all relevant measures being from the same source. Further, as discussed in the response to Question 8, person-mean centering can help reduce concerns associated with source/method biases that typically plague cross-sectional research.
Conclusion
Although ESM research holds much promise for examining organizational phenomena, there are multiple decision points unique to ESM that scholars must navigate to ensure that the data are not only modeled correctly but also gathered in a manner conducive to challenging and expanding within-person theory and appropriately testing hypothesized within-person and cross-level relationships. Table 1 summarizes specific, actionable recommendations (when applicable) and considerations for authors and reviewers associated with the question topics. Importantly, there are also areas warranting extensive research as ESM continues to gain momentum—from the use of different influence tactics to enhance participation rates to research addressing potential biases in ESM data—to move the field forward.
Recommendations Surrounding Experience Sampling Methods (ESM) for Researchers, Reviewers, and Editors.
As a final point, a substantial hurdle for researchers occurs when they must effectively analyze the complex data obtained using ESM. As such, we have provided Supplemental Online Material (Part B) that includes a table of the characteristics, advantages, and drawbacks of statistical packages commonly used in ESM research. In addition, given that most theories tested with ESM involve an underlying process model that necessitates multilevel mediation, we provide a Supplemental Online Document (Part C) with basic multilevel mediation syntax for Mplus (L. K. Muthén & Muthén, 1998-2015)—the program used and recommended by all panelists. However, we note that there are new approaches to modeling ESM data that more completely address the complexities of dynamic processes (dynamic SEMs; Asparouhov, Hamaker, & Muthén, 2018). Because these models are still being developed, however, we felt it premature to suggest them as a default approach. In sum, we hope that these materials and the panelists’ responses to questions surrounding ESM-related theory, methods, and analyses will assist scholars interested in effectively designing, analyzing, and framing their ESM research.
Supplemental Material
Supplemental Material, Supplemental_File_ORM-17-0075.R1-Final - Experience Sampling Methods: A Discussion of Critical Trends and Considerations for Scholarly Advancement
Supplemental Material, Supplemental_File_ORM-17-0075.R1-Final for Experience Sampling Methods: A Discussion of Critical Trends and Considerations for Scholarly Advancement by Allison S. Gabriel, Nathan P. Podsakoff, Daniel J. Beal, Brent A. Scott, Sabine Sonnentag, John P. Trougakos and Marcus M. Butts in Organizational Research Methods
Footnotes
Authors’ Note
The authors would like to thank Nitya Chawla, Daphna Motro, and Trevor Spoelma for their assistance in collecting the data discussed in this paper.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported in part by a research grant from the Social Sciences and Humanities Council of Canada awarded to John Trougakos (435-2014-0693).
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
