Abstract
This is the first paper in a series of two that synthesizes, compares, and extends methods for causal inference with longitudinal panel data in a structural equation modeling (SEM) framework. Starting with a cross-lagged approach, this paper builds a general cross-lagged panel model (GCLM) with parameters to account for stable factors while increasing the range of dynamic processes that can be modeled. We illustrate the GCLM by examining the relationship between national income and subjective well-being (SWB), showing how to examine hypotheses about short-run (via Granger-Sims tests) versus long-run effects (via impulse responses). When controlling for stable factors, we find no short-run or long-run effects among these variables, showing national SWB to be relatively stable, whereas income is less so. Our second paper addresses the differences between the GCLM and other methods. Online Supplementary Materials offer an Excel file automating GCLM input for Mplus (with an example also for Lavaan in R) and analyses using additional data sets and all program input/output. We also offer an introductory GCLM presentation at https://youtu.be/tHnnaRNPbXs. We conclude with a discussion of issues surrounding causal inference.
Keywords
Causal inference is a core part of scientific research and policy formation. There are multiple pathways to causal inference (Cartwright, 2007, 2011), but a popular approach uses longitudinal panel data made up of multiple units measured at multiple occasions. Such data are useful for causal inference by helping to control for confounds and modeling lagged relationships as units of analysis change over time (Hausman & Taylor, 1981; Kessler & Greenberg, 1981; Liker, Augustyniak, & Duncan, 1985). With this approach, organizational researchers regularly use panel data to infer causality, often with cross-lagged panel models.
For example, at an individual level of analysis, Meier and Spector (2013) studied 663 people at five occasions, finding reciprocal effects among counterproductive work behaviors work stressors, inferring “a vicious cycle with negative consequences for all parties involved” (p. 537). At a higher level of analysis, Van Iddekinge et al. (2009) studied 861 locations of an organization at six occasions, showing reciprocal effects for human resources (HR) factors and performance, thus offering the advice that “human capital investments…can yield a high return” (p. 840). At a national level of analysis, Diener, Tay, and Oishi (2013) studied 135 countries at six occasions, finding reciprocal effects for income and subjective well-being (SWB), inferring that in terms of SWB, “people did not adapt to income increases” (p. 275).
By using such observational data, this work has the potential to show real-world evidence of effects that may otherwise be difficult to uncover. As medical researchers note, such evidence may be useful due to “its potential for complementing the knowledge gained from traditional clinical trials, whose well-known limitations make it difficult to generalize findings” (Sherman et al., 2016, p. 2293; see also Booth & Tannock, 2014). However, given this potential, many studies often fail to capitalize on the unique opportunities that panel data offer, including strengthening causal inferences by controlling for stable factors and testing hypotheses about the long-run behavior of the systems being studied. This occurs due to the lack of integration across fields in the tools used for longitudinal data modeling (contrast: Bollen & Curran, 2006; Box, Jenkins, & Reinsel, 2008; Hsiao, 2014; Lütkepohl, 2005; McArdle & Nesselroade, 2014). The result is that organizational researchers often fail to examine a range of theoretically relevant processes and effects when modeling panel data.
For example, many researchers use latent curve models separately from lagged effects models, perhaps due to a belief that modeling curves precludes lagged effects (e.g., Rogosa & Willett, 1985) or that econometric tools “are usually less applicable for the kinds of data psychologists and micro HR/OB scholars have,” often with few measured occasions T (Ployhart & Ward, 2011, p. 414). Yet, accounting for curves (i.e., trends) is crucial for lagged effects models (Box et al., 2008; Lütkepohl, 2005), and many econometric tools are designed specifically for the “small T” case (Arellano, 2003; Baltagi, 2014; Hsiao, 2014).
To help researchers overcome the limitations of current panel data modeling methods, we synthesize, compare, and extend approaches to panel data modeling in two papers. Our primary goals are to: (a) show how panel data can help test hypotheses (or infer processes) in more powerful and useful ways than are typically found in the organizational literature and for this purpose, (b) introduce methods from disciplines that may be foreign to many readers.
We tackle these by starting with a typical cross-lagged panel model to build a more general structural equation model (SEM), which we call a general cross-lagged panel model (GCLM), that controls for stable factors and increases the range of dynamic processes that can be modeled. Our approach is designed for the typical panel data case where T < 20 (and usually T < 10), but most of what we discuss can be applied to larger T cases by using dynamic structural equation modeling (DSEM; see Asparouhov, Hamaker, & Muthén, 2018). Our second paper compares our approach to others, including multilevel panel data models. Across both papers, we offer an integrative overview drawn from multiple traditions, resulting in powerful new conceptual and statistical tools for modeling panel data.
In what follows, we first conceptually treat GCLM parameters. Then, we treat tests of short-run effects as direct effects among variables, versus long-run “impulse responses” that capture all indirect effects of one variable on another over time. We then describe a general SEM for estimation and hypothesis testing. To illustrate a GCLM, we reexamine the income-SWB relationship at the national level (from Diener et al., 2013), failing to support effects among these variables. We also model individual and organizational effects from Meier and Spector (2013) and Van Iddekinge et al. (2009) to exemplify our points—we reanalyze their data and present GCLM findings in Appendix A in the Supplementary Materials (available in the online version of the journal).
All input/output for the Mplus program are available in the online Supplementary Materials, along with an Excel file to automate Mplus input for a GCLM and its variants. We also include an example of the GCLM in Lavaan for R and note that the Mplus2lavaan (2019) program for R can help translate most Mplus input to Lavaan. All Supplementary Material can be cited and is available for download at https://doi.org/10.26188/5c9ec7295fefd. To help the reader, we also offer a presentation on the GCLM and the processes it captures at https://youtu.be/tHnnaRNPbXs. We conclude by discussing issues in causal inference under uncertainty, including threats to causality due to trends and regime changes (i.e., parameter changes over time).
Before proceeding, we emphasize that our goal is to offer a practical framework for modeling panel data based on the idea that “it pays to experiment with the…techniques that panel data make available” (Halaby, 2004, p. 541). In the end, we agree that “there is no such thing as the methodology for analyzing panel data, but a collection of…techniques that have accumulated from a series of heterogeneous motivations” (Arellano, 2003, p. 2). Our goal is to explore these techniques and expand the toolkits of researchers who regularly use panel data to make causal inferences. In this tradition, we seek to improve current practices.
Building a General Cross-Lagged Panel Model
There are many useful introductions to longitudinal data models (e.g., Allison, 2005, 2009; Baltagi, 2013; Bollen & Brand, 2010; Bond, 2002; Cole, 2012; Enders, 2014; Halaby, 2004; Hamaker, Kuiper, & Grasman, 2015; Hsiao, 2007; Lütkepohl, 2006, 2013). We draw on this work to build a GCLM while focusing on its conceptual logic and tools for hypothesis testing that follow from it (see YouTube). Although the GCLM may seem complex, any subset of its parameters (in Table 1) can be used to build a panel data model, and our methods for hypothesis testing will both clarify and simplify causal inference using the GCLM.
Parameters, Their Purposes, and SEM Specifications (for Observed Variables
Note: SEM = structural equation model.
To begin in a familiar way, we first introduce a cross-lagged panel model and treat the conceptual underpinnings of its parameters. With this structure in place, we then offer several ways to extend the model, proposing a GCLM that includes additional parameters to expand the range of dynamic processes that can be modeled and then used for hypothesis testing.
A Cross-Lagged Panel Model
We start with a cross-lagged panel model where all variables are a function of the past (see Figure 1). Throughout, our figures use SEM notation as follows: Observed variables are squares, latent variables are circles, single-headed arrows show dependence, and double-headed arrows are (co)variances (we omit intercepts/means). For simplicity, we use two variables,

An AR(1)CL(1) model.
We start with a cross-lagged panel model using some specialized notation as follows:
Here,
Occasion effects
To model causal effects in panel data, it is important to account for overall changes in a sample across occasions, which may be due to a variety of aggregate factors that are unrelated to lagged effects (i.e., AR and CL terms). We account for these with an occasion effect
Autoregressive (AR) effects
and
A key part of the cross-lagged model are AR effects that link the past and future (see Figure 1). With this approach, a unit’s current state is a function of its past, so
An AR term
We return to AR terms when treating long-run effects, and Online Appendix B treats the special case of AR ≥ 1, but for now, we lay a foundation for seeing CL effects as causal by noting that AR terms help control for some confounds. For example, employees may engage in counterproductive work behaviors as a matter of habit rather than due to increases in work stressors, so controlling for past counterproductive work behaviors with AR terms is relevant. Similarly, organizations may experience high performance due to persistent market forces rather than changes to HR practices, so again, performance AR terms may be useful. Also, nations may experience low well-being that persists for reasons that may be unrelated to decreases in national income. Thus, AR terms reflect persistence, but they also control for a variable’s past levels to help avoid drawing erroneous causal conclusions using CL terms.
This understanding of AR terms motivates a discussion of CL effects, but before this, it is important to note that some processes cannot be modeled by a single AR term, such as lagged effects that take longer than a lag h = 1 to appear or complex processes that can be modeled by both a positive and negative AR term at different lags. As Figure 1 shows, AR terms recursively link the past to the future (e.g.,
Cross-lagged effects
and
By including AR terms, it becomes posible to use the past of one variable to uniquely predict the future of another. Such CL effects imply that each unit’s current state is a function of its past on other variables; so for example,
With this logic, CL terms are used to infer causality, but as Figure 1 shows, they only imply a “short-run” effect as a direct effect of the past on the future. Just like AR terms, these depict a system’s short-run behavior, with implications for CL terms ≥ 1, as noted in Online Appendix B. Yet, investigating long-run behavior requires examining how the past indirectly affects the future along all AR and CL paths simultaneosuly (e.g., the total effect of an initial
We will cover long-run hypothesis tests when we treat impulse responses. For now, we note that just like AR effects, higher-order CL terms may be needed for some processes. For example, work stressors may have delayed or complex effects on counterproductive work behaviors, requiring a second lag c = 2 in a CL(2) model, such as an effect
Impulses
The model also includes a residual term to allow units to differ over time due to random inputs (Denrell, Fang, & Liu, 2014). Although residuals are often taken for granted in regression, in cross-lagged models, they actually have an important substantive role that requires some theoretical prefacing. For example, consider that rules and routines guide social entities but behavior and events are never predictable as novelties emerge over time (Becker, Knudsen, & March, 2006; Weick, 1998). The same is true for larger economic changes (Lütkepohl, 2015), which are typically unpredictable or even a priori unexplainable (Cochrane, 1994). This is echoed by research efforts in social science that fail to explain substantial variation because of the stochastic nature of many phenomena (Abelson, 1985).
To capture such random inputs for each unit i at a given occasion t, we include a random term
Before this, however, we note that an impulse
The co-movement
Summary and limitations
The cross-lagged model has many useful properties. It controls for occasion effects
However, there are two limitations of this approach that motivate a GCLM. First, all units are treated as if they were the same in the long run—Figure 1 does not reflect any stable between-unit differences. This is anathema to organization research in which individual and organizational differences such as personality or culture are well recognized and units differ systematically over time. By failing to model stable factors, they will be confounded with the system dynamics that should be reflected by AR and CL terms (Hamaker et al., 2015). Thus, a more general model is needed to account for stable factors, which we will call unit effects.
The second limitation is that the dynamic process linking the past and the future via AR and CL terms is assumed to follow a simple, indirect-effects structure. As we noted, AR and CL terms depict persistence (or regression to the mean) of a past impulse, but this might persist (or fade) in complex ways. Thus, a more general model may help to overcome the indirect-effects structures associated with AR and CL terms, which we will treat in the next section using moving average (MA) and cross-lagged moving average (CLMA) terms.
A General Cross-Lagged Model
To generalize the cross-lagged model, we now sequentially introduce unit effects as well as MA and CLMA terms. In doing so and in what follows, we draw on three modeling traditions: (a) vector autoregressive (VAR) models (Canova & Ciccarelli, 2013; Lütkepohl, 2005; Sims, 1980), (b) vector autoregressive moving average (VARMA) models (Box et al., 2008; Browne & Nesselroade, 2005), and (c) dynamic panel data models (Arellano, 2003; Arellano & Bond, 1991; Baltagi, 2014; Hsiao, 2014). From this work, we take the idea that processes and effects may be more complex than AR and CL terms imply. Furthermore, there may be stable factors that differentiate units of analysis over time, to which we now turn.
Unit Effects
Researchers often seek to explain two distinct causes of variation in people, organizations, and larger entities. On the one hand there is variation within units as each one changes relative to itself over time; AR, CL, and impulse terms capture these dynamics as units experience random shocks that persist via AR and CL paths. On the other hand, units may systematically differ from each other, producing variation between units of analysis due to factors that create stability rather than occasion-specific change.
To elaborate, if a unit i is a person, psychological factors can explain stability over time, including stable patterns of embodied cognition (Barsalou, 2008), social roles and norms (Andersen & Chen, 2002; Fournier, Moskowitz, & Zuroff, 2008), cognitive ability (Deary, Pattie, & Starr, 2013), personality or affective traits (Matthews, Deary, & Whiteman, 2003), and habits of thought/action that emerge in stabilized person-environment interactions (Fleeson, 2001; Mischel, & Shoda, 2008; Neal, Wood, & Quinn, 2006). Alternatively, if i is a group, organization, or a nation, substantial scholarship treats how collectives emerge as stable entities, such as by the formation of institutions (March & Olsen, 1989) and collective routines to guide social and material processes (Feldman & Orlikowski, 2011; Winter, 2013).
Such causes of between-unit differences are not the same as causal effects among variables as they change over time (Allison, 2005; Hamaker et al., 2015). Instead, between-unit differences are akin to unit-specific trends (e.g., long-run averages) that systematically differentiate units over time (i.e., between-unit differences). These should not confound the AR, CL, and impulse terms that represent perturbations around any such trends (see Online Appendix B) because stable factors are constant by definition and thus do not have a clear role in models of causality over time. To account for this, we treat each unit i as a function of unit-specific factors that are constant or nearly constant over T, modeled as a unit effect
Unlike

An AR(1)CL(1) model with unit effects.
Although everything changes with time, we model
This can be seen as a
Moving average effects
and
To generalize model dynamics, we now introduce MA and CLMA terms. The idea motivating these is that long-run and short-run dynamics may be different as impulses persist/fade over time, but AR and CL terms imply equivalent long- and short-run dynamics as a single set of parameters linking the past to the future. Because AR and CL terms imply impulse persistence, relying on only them to capture system dynamics is akin to assuming that unexpected changes persist or fade multiplicatively vis-à-vis AR and CL terms. This can be modified by making the future a direct function of past impulses, which is how MA and CLMA terms modify the typical cross-lagged model.
We begin with MA terms, which modify AR paths by making observations a direct function of past impulses (Box et al., 2008; Hamaker, Dolan, & Molenaar, 2002). This allows MA terms to modify the short-run persistence of an impulse, whereas AR (and CL) terms still reflect long-run dynamics. As Figure 3 shows,

An AR(1)CL(1)MA(1) model with unit effects.
By including MA terms, generality is added to the way that dynamic processes can be modeled—specifically, by allowing MA terms to modify the way AR terms imply short-run persistence of impulses. This is seen by path tracing in Figure 3, where short-run persistence of an impulse is a sum of MA and AR terms (i.e., a total effect of
To explain the first case, consider that as MA terms become more positive (
In the second case, adaptation may occur rapidly at first but then slow over time. For example, there may be short-run adaptive responses to changes in counterproductive work behaviors (e.g., management interventions), organizational performance (e.g., increased competition), or national income (e.g., less stringent budget controls), but if these responses fade or become ineffective, then what remains of the initial change may persist. This is made possible by MA terms because as they become more negative (i.e.,
To add additional generality to the model, higher-order MA lags may be included for q MA effects in an MA(q) model. Here, the sum of all MA terms is a shorthand for how MA effects from a single past u modify short-run persistence (
Cross-lagged moving average effect
and
Just as AR and MA terms allow modeling a separate short-run and long-run dynamic structure, Figure 4 shows that the structure associated with CL terms can be extended analogously by making each unit’s standing on an observed variable a direct function of other variables’ past impulses. We call these CLMA terms, which arise when

A full GCLM, AR(1)MA(1)CL(1)CLMA(1) model with unit effects.
By including CLMA terms, the model changes how causal effects can be understood. As noted previously, lagged effects can be seen as implying an effect of past impulses on future observed variables. In turn, just as the short-run persistence for a variable becomes AR+MA, the short-run effect of one variable on another becomes CL+CLMA. As Figure 4 shows,
This type of causality uses an interesting through experiment to ground it: Consider that if impulses are random, then it is as if a natural experiment were done at each occasion by randomly assigning units to a new level on a variable (e.g.,
The issue of causality aside, CLMA terms offer pragmatic value by allowing complex forms of dependence among variables. We treat this by analogizing the two previous MA cases. The first involved delayed adaptation, such that an unexpected change in a variable has an effect on the future of another, but adaptation then limits the duration of effects. This could be a case of short-lived effects of work stressors on counterproductive behaviors, HR practices on performance, or national income on SWB as each system adapts to the change. Here, an impulse
The second case involved a small short-run effect that is highly persistent, such as a reverse-causal case of counterproductive behaviors affecting work stressors, organizational performance affecting HR practices, or national SWB affecting income, but with each effect being small yet long-lived over time. For this, an impulse
The point is that CLMA terms add generality to the kinds of dynamics that can be modeled. For this purpose, researchers may include l higher-order CLMA terms in a CLMA(l) model, such as if l = 2 for a CLMA(2) model. Again, the CLMA effects from a single past u can act as a kind of shorthand indicating how CLMA terms modify short-run effects (
Hypothesis Testing With the GCLM
To facilitate testing hypotheses with the GCLM, there are methods that can be easily implemented even if the models are very complex (e.g., many higher-order lags). As we now describe, short-run effects can be evaluated with Granger-Sims causality tests, whereas long-run effects can be evaluated with impulse responses that indirectly link past impulses to future observed variables over time. We now treat each of these in turn.
Short-Run Effects: Granger-Sims Tests
To facilitate hypothesis testing with the GCLM, we offer a four-step process that is easy to use in SEM software (inspired by Granger, 1969; Sanggyun & Brown, 2010; Sims, 1980, 1986). The method maps onto the Granger-Sims logic that impulses on one variable can be understood as causes of future observations on another. This is a test for short-run effects because our four steps only assess the direct effects of past impulses—short-run effects are direct effects; long-run effects involve indirect effects. For this, null hypothesis significance tests can be used, but we use fit criteria to balance parsimony and statistical fit. Step 1: Estimate a panel data model of interest, such as the full GCLM in Figure 4 and Equations 7 and 8, and obtain model fit information such as information criteria (e.g., Akaike Information Criterion or Bayesian Information Criterion). Step 2: Test an x→y effect by constraining CL and CLMA terms linking x’s impulse Step 3: Test a y→x effect with the same approach, comparing results to Step 1. Step 4: Test x→y and y→x “feedback” or “reciprocal effects” with all constraints from Steps 2 and 3 and compare to Step 1. If feedback exists, then intervening to change
However, these four steps only offer a picture of short-run effects rather than the form effects take over time (Dufour et al., 2006; Dufour & Renault, 1998; Hsiao, 1982; Lütkepohl, 1993). Consider that with more than two variables such as x, m, and y, there may be a direct effect x→y and an indirect effect such as x→m→y over time, but only the former is tested. To tackle these issues, we now treat long-run effects using the logic of impulse responses.
Long-Run Effects: Impulse Responses
Although tests for short-run effects are common, their results may not be useful for planning interventions, which requires predicting the results of actions over time (Cartwright & Hardie, 2012). For this, we use impulse responses, which we treat as total effects of a past impulse on future observations over time, including all indirect effects via AR and CL paths (Lütkepohl, 2005; Sims, 1980; Stock & Watson, 2005). Impulses are the focus because, as Figure 4 shows, “changes in the variables are induced by non-zero residuals, that is, by shocks…. Hence, to study the relations between the variables, the effects of…shocks are traced through the system” (Lütkepohl, 2013, p. 154). Indeed, methods to account for natural or planned experiments can adopt this logic by modeling impulses via time-varying treatment variables (Angrist & Kuersteiner, 2011; Bojinov & Shephard, 2017; Stock & Watson, 2018).
By conceptualizing

(a) Impulse response functions for AR(1)MA(1) model. Note: The y-axis is effect estimates, and the x-axis is the response horizon in years so that the plotted lines indicate the effect of a 1-unit impulse in 2006 over the next 5 years. Solid lines represent effect estimates; dotted lines represent 97.5% and 2.5% confidence intervals obtained using a nonparametric bootstrap with roughly 15,000 replications. Impulse responses begin at the first occasion t = 1 because the highest lag order in the model = 1. (b) Impulse response functions for AR(1)MA(2) model. Note. See the Note for Figure 5a, except the impulse begins in 2007 at t = 2 (rather than 2006 at t = 1) because the highest lag order in the model = 2, so the first occasion is “lost” when estimating effects. Thus, we show the effect of a 1-unit impulse in 2007 over 4 years. (c) Impulse response functions for AR(2)MA(1) model. Note: See the Note for Figure 5b. (d) Impulse response functions for AR(2)MA(2) model. Note: See the Note for Figure 5b.
By estimating and plotting these effects and their confidence intervals (CIs), researchers can test hypotheses that map more directly onto research questions such as if “human capital investments…can yield a high return” (Van Iddekinge et al., 2009, p. 840) or if, in terms of SWB, “people [do] not adapt to income increases” (Diener et al., 2013, p. 275). Impulse responses can show such effects across all paths modeled in a GCLM. Indeed, in the case that effects do not fade due to AR or CL terms = 1 (see Online Appendix B), impulse response analysis offers a simple way to show how all lagged parameters may imply persistent effects in a studied time frame.
This said, impulse response analysis has limitations. Some of these we treat later, but for now we note that the earliest impulse that can be used has a t equal to a model’s highest lag order. This is because higher-order lags involve missing MA and CLMA terms in early occasions (as we note in our next section). Thus, impulse responses must begin at the first impulse with all modeled effects “leaving” the impulse. Also, as is well known for mediation analysis, indirect effect estimates are not normally distributed, so testing can be done using bootstrapped CIs or Bayesian analogues (Dufour et al., 2006; Kilian, 1999; Wright, 2000).
SEM Specification and Estimation
To model panel data, SEMs are useful because of their flexibility (Allison, 2005; Bollen & Brand, 2010). Due to its generality and stable algorithms, we use the approach found in Mplus (see Asparouhov & Muthén, 2014; Muthén, 2002; Muthén & Asparouhov, 2009; L. K. Muthén & Muthén, 1998-2018). As a special case of this, we show an SEM as:
with all terms typically understood as follows:
For concision we assume error-free measures that reduce Equation 9 to
To estimate a GCLM, any unit effect
For the sake of concision, we describe model identification conditions in Online Appendix C but note that many combinations of AR, MA, CL, and CLMA lags are possible (i.e., different p, q, c, and l, respectively) and each will have unique identification conditions. Our online Excel worksheet automates Mplus input for models with different lag orders for different observed variables, but researchers should be aware of constraints on identification as lag orders increase. A basic GCLM with single lag orders is identified with T ≥ 4, but even complex models will often be identified if T ≥ 6 (for general insight, see Bollen, 1989).
Also, there are special considerations for model with higher-order lags, which become interpretable at the first occasion t that is subject to all lagged effects (i.e., when t equals the highest lag order in a model +1; see Online Appendix C). Thus, the highest lag order is equal to the number of early occasions that are “lost” because they cannot be predicted by occasions before t = 1. In these cases, the GCLM includes freely estimated AR and CL terms in early occasions to account for unmodeled effects prior to t = 1 (see Online Appendix C).
Finally, we offer a few comments about
Income and Subjective Well-Being
To illustrate model estimation and interpretation, we reanalyze data from Diener et al. (2013), who used Gallup World Poll data to study the relationship between SWB and income at the national level (other examples are in Online Appendix A). SWB was measured by self-rated life evaluations on a 0 to 10 scale; income was equivalized, log-transformed, and then multiplied by 2 to stabilize model estimation. With N = 135 nations from 2006 to 2011 (T = 6) and roughly 1,000 people responding for each country i at each year t, the data represent about 95% of the world’s adult population. The mean for each country i at each year t was computed to represent average income
Descriptive Statistics.
Note: SWB = average subjective well-being; INC = average income logged;
These data are useful for studying causal effects because income and SWB cannot be easily manipulated and methods with observed proxies for this can have strong assumptions (e.g., Ettner, 1996; Lindahl, 2005; Meer, Miller, & Rosen, 2003). Also, diverse causes can explain covariance in well-being and income. Deaton (2002) notes three possible cases for SWB or health: “[1] Income might cause health, [2] health might cause income, or [3] both might be correlated with other factors; indeed, all three possibilities might be operating” (p. 15; Deaton, 2003; see also Diener & Biswas-Diener, 2002). The GCLM addresses these issues as follows.
First, income may lead to SWB by reducing monetary stressors and increasing access to positive environments (Diener & Biswas-Diener, 2002). Such effects can be understood in relation to life circumstances and the relative comparisons that they allow (Clark, Frijters, & Shields, 2008; Frijters, Haisken-DeNew, & Shields, 2005). Yet, second, some “literature has been skeptical about any causal link from income…and instead tends to emphasize causality in the opposite direction” (Deaton, 2003, pp. 118-119). In terms of well-being, some research shows no lasting effect of income (Easterlin, Morgan, Switek, & Wang, 2012) but an effect of well-being on income via employment and other factors (Binder & Coad, 2010; De Neve & Oswald, 2012; Michaud & Van Soest, 2008; Oswald, Proto, & Sgroi, 2015). Still other studies find bidirectional causality or “feedback” effects (e.g., Chen, Clarke, & Roy, 2014; Devlin & Hansen, 2001; Erdil & Yetkiner, 2009; French, 2012), which many researchers propose should exist for various reasons (e.g., Deaton, 2003; Diener, 2012).
Third, in terms of confounding factors, our model controls for occasion effects (
In sum, a GCLM helps in studying variables like income and SWB or health because researchers want to make causal inferences about them (e.g., Sacks, Stevenson, & Wolfers, 2012). However, weak methods often require admitting that “we shall have little to say about a causal interpretation” (Sacks et al., 2013, p. 8). By way of example, we now explore the process of GCLM specification and checking on the road to causal inference.
Model Specification
Causal inference with the GCLM requires choosing lag orders and some number of unit effects. To make this choice, alternative models can be compared, but this requires first choosing which models to specify for comparison. To guide this, conservative models are typically best for out-of-sample generalizations, wherein conservatism means simpler models that rely on theory, past findings, and contextual information (Allen & Fildes, 2001, 2005; Armstrong et al., 2015). We now motivate four such models for comparison.
Past research shows that SWB (
National income
For the effects among income
Results for all models are in Table 3, with occasion effects omitted for concision and impulse/unit effect covariances standardized as correlations. Impulse responses for all models are in Figures 5a through 5d (generated as indirect effects from an initial impulse to future observed occasions using Mplus’s “MODEL INDIRECT” command), with 95% bootstrapped CIs using 20,000 draws, with missing data in early periods reducing convergence to roughly 15,000. All Mplus input and output is available in our online materials, including an Excel worksheet used to create Mplus input and impulse responses for these four specific models.
Model Results (models are referred to using the lag specification for income
Note: Columns are named after the AR/MA specification for income. SWB = subjective well-being; AR = autoregressive; MA = moving average; CL = cross-lagged; CLMA = cross-lagged moving average; CFI = Confirmatory Fit Index; TLI = Tucker-Lewis Index; RMSEA = root mean squared error of approximation; SRMR = standardized root mean squared residual; AIC = Akaike Information Criterion; BIC = Bayes Information Criterion; aAIC = sample-size adjusted AIC; aBIC = sample-size adjusted BIC.
*p < .05. **p < .01.
Model Selection
Model selection can be done by substantive and statistical checking. We first offer a substantive interpretation of results by checking estimates for consistency with theory and contextual knowledge (see Table 3). For this, we rely on impulse responses because they simplify model comparisons in the presence of varying lag orders (see Figures 5a-5d). We then discuss the use of model fit indices for model selection.
Substantive checking
We first examine the SWB dynamics, with an AR(1)MA(1) structure in all four models. As expected, the persistence of impulses quickly falls (top-left of Figures 5a-5d), with impulses almost entirely faded by the fourth future year. Also, 95% CIs include zero by the second year, so statistical significance exists only for the direct effect of a past impulse. In Table 3, this is the combined AR and MA term
On the other hand, income dynamics tell a different story (see top-right of Figures 5a-5d). The AR(1)MA(1) and AR(2)MA(2) models in Figures 5a and 5d imply mean-reversion, with Table 3 showing the AR(1)MA(1) model’s AR effect
Alternatively, the AR(1)MA(2) and AR(2)MA(1) models in Figures 5b and 5c imply expected persistence in income, with Table 3 showing the AR(1)MA(2) model having an AR effect
For the income→SWB effect (bottom-left of Figures 5a-5d), all impulse responses include zero in 95% CIs. The short-run effect is positive, with a CL+CLMA term
Finally, the SWB→income effect shows 95% CIs include zero at all time horizons (bottom-right of Figures 5a-5d). Yet, unlike the income→SWB effect, the SWB→income effect tends to be negative, with the short-run effect
Statistical checking
Many researchers agree that model selection should use indices balancing statistical fit with model parsimony (Allen & Fildes, 2001, 2005; Armstrong, 2007; Armstrong et al., 2015; Burnham & Anderson, 2004; Hu & Bentler, 1998, 1999; Lütkepohl, 2005). However, different communities use fit indices differently. SEM researchers typically make recommendations based on simulations (e.g., Hu & Bentler, 1998, 1999). This often results in recommending fit index cutoffs that are not specific to panel data or predicting the results of interventions. Researchers from other fields do not always appreciate this approach.
For example, forecasters empirically examine fit index performance for out-of-sample predictions with real data (Fildes & Ord, 2002; Makridakis & Hibon, 2000), showing that accurate prediction can be less a function of fit indices than substantive checking and other factors (see Allen & Fildes, 2001, 2005; Armstrong et al., 2015; Green & Armstrong, 2015). Economists agree, noting that “statistical fit is overemphasized as a criterion…. As a policymaker, I want to use models to help evaluate the effects of out-of-sample changes in policies” (Kocherlakota, 2010, p.17), which requires substantive and contextual reasoning. Therefore, we do not unconditionally endorse the use of cutoff criteria often found in the SEM community—at least until such cutoffs are examined for use with panel data.
Here, we advocate balancing concerns about fit with substantive checking and an interest in parsimony. If SEM fit indices show serious problems, this may be cause for concern, but modest differences in fit or poor fit for a model that accurately depicts a known process seem acceptable. When in doubt, “you should probably aim towards simplicity at the expense of good specification” (Allen & Fildes, 2001, p. 21). However, “[o]f course, any simple model may sometimes be too simple” (Bernanke & Blinder, 1988, p. 1), and therefore theoretical and contextual knowledge of the processes being modeled should always be used.
To illustrate model selection by statistical checking, we use the following fit indices: standardized root mean square error of approximation (RMSEA), Tucker-Lewis index (TLI), Comparative Fit Index (CFI), Akaike Information Criterion (AIC), Bayesian or Schwarz Information Criterion (BIC), and sample-size adjusted version of the AIC and BIC. We also report the standardized root mean square residual (SRMR) but emphasize the former indices for their balance of parsimony and fit. Examining these indices in Table 3 shows no serious problems with any single model and very modest differences in terms of CFI, TLI, RMSEA, and SRMR. The AIC favors the more complex AR(2)MA(2) model, and the BIC favors the more parsimonious AR(1)MA(1) model, which is expected (see Burnham & Anderson, 2004; Lütkepohl, 2005). This is reversed for the sample-size adjusted AIC and BIC.
Importantly, our preferred AR(1)MA(2) model shows acceptable levels of fit using typical SEM indices (e.g., CFI = .978; TLI = .959; SRMR = .026; RMSEA = .094), but it is the worst model in terms of AIC and BIC indices. However, the differences are inconsistent across models and are often minor. Therefore, we favor a AR(1)MA(2) model because of its acceptable fit and because the substantive relationships it shows are consistent with theory.
It is notable that other procedures can be used for model checking, such as for nonlinearity and local misfit using modification indices, covariance residuals, and residual plots (Asparouhov & Muthén, 2014). This is often considered obligatory, so we do not treat it here.
Model Interpretation and Hypothesis Testing
Using the AR(1)MA(2) model for inference (Figure 5b), we do not expect our results to conform to past studies given the sizable unit effects for SWB and the standardized
Granger-Sims Tests.
Note: Parentheses after χ2 values are degrees of freedom. SWB = subjective well-being; CFI = Confirmatory Fit Index; TLI = Tucker-Lewis Index; RMSEA = root mean squared error of approximation; SRMR = standardized root mean squared residual; AIC = Akaike Information Criterion; BIC = Bayes Information Criterion; aAIC = sample-size adjusted AIC; aBIC = sample-size adjusted BIC.
**p < .01.
More interesting is the weak, negative effect of SWB on income, which is opposite of what is often found (Deaton, 2003; Diener et al., 2013). Supporting this effect in the short-run, Table 4 shows that removing it reduces fit via CFI, TLI, and RMSEA. Yet, AIC and BIC terms show improved fit. In terms of the long-run effect, all impulse response CIs encompass zero. In sum, the weak nature of the effect implies it is untrustworthy, but were it present, then it could be explained. For example, some research shows that positive psychological states can negatively affect motivation and resource allocation for goal pursuit (Vancouver, More, & Yoder, 2008). The effect is sensible if increasing SWB demotivates seeking economic welfare or reductions in SWB orient people toward economic welfare.
In sum, by including AR, MA, CL, and CLMA terms, as well as time-varying unit effects and occasion effects, we do not find strong associations between income and SWB, and the SWB→income effect we find is negative, which runs counter to results using typical cross-lagged models (e.g., Diener et al., 2013). As we show in our second paper, this may be due to uncontrolled unit effects and/or a need for MA and CLMA terms in such past studies.
Discussion
This is the first of two papers in which we synthesize, compare, and extend panel data methods using SEM. In this first paper, we proposed a new panel data model, the GCLM, to incorporate stable factors in the form of unit effects while expanding the range of dynamic processes that can be modeled by using MA and CLMA terms. We treated these parameters and their application, covering model specification, checking, and interpretation by studying income-SWB dynamics, which did not support previous findings of positive effects among these variables (e.g., Diener et al., 2013). This suggests reappraising the sign and magnitude of income-SWB effects (Easterlin, 1995, 2001). We now conclude with thoughts on causal inference, starting with threats to this inference; Online Appendix D treats ways to modify the GCLM, including interactions, random slopes, and nonstandard measurement occasions.
Threats to Causal Inference: Trends and Regime Changes
To interpret GCLM results, it is important to address two threats to causal inference (Clements & Mizon, 1991): trends, including seasonal or cyclical effects, and changes in how a system functions, or regime changes (Granger & Newbold, 1974; Lütkepohl, 2005; Sims, Stock, & Watson, 1990). Grappling with these is important because if they exist, they may drive observed relationships rather than the random impulses that are meant to justify causal inference (Hendry, 2004). To raise awareness of these threats, we discuss each in turn.
Concerns over trends have generated substantial work (Harvey, 1985, 1997; Stock & Watson, 1988, 1999), covering unique types of trends: long-run trends due to things like maturation, periodic trends such as seasonal effects or circadian rhythms, cycles that wax and wane unpredictably (e.g., business cycles or depressive states), and random or stochastic trends caused by persistent impulses. Our model accounts for these in five ways: (a) An occasion effect
However, additional tools may be required. For example, periodic trends like seasons or times of day can be modeled with latent variables (similar to “common methods factors”), or latent variables can act as additional unit effects to model unit-specific cyclical trends (e.g., a term
This said, certainty about the existence of trends is often impossible (Heckman, 1991; Stock & Watson, 1999). Although de-trending data is often recommended (e.g., Curran & Bauer, 2011; Curran, Lee, Howard, Lane, & MacCallum, 2012; Hoffman & Stawski, 2009), there is no single way to do this, and tests for trends are often ambiguous (Davidson, 2013; Haldrup, Kruse, Teräsvirta, & Verneskov, 2013). The fact is that the evolution of any system involves mixtures of multiple processes, leading some to say that “no one really understands trends, even though most of us see trends [in] data” (Phillips, 2003, p. C35; Heckman, 1991). Also, visual inspections and detrending methods may be useful for N = 1 cases (see Jebb & Tay, 2016; Jebb, Tay, Wang, & Huang, 2015), but this is impractical with larger N. In the face of uncertainty, unit effects automatically de-trend data, but theoretical and contextual knowledge about a process can also be used (Allen & Fildes, 2001; Armstrong et al., 2015).
Next, regime changes refer to changes in the way a system functions over time—such as when water turns to ice, a person gets a new job, or an organization changes strategy. The idea is that there is a threshold beyond which a system functions differently, complicating prediction and causal inference (see Bak, 1996; D’Souza & Nagler, 2015). To investigate this, increased variances and large AR terms may be observed due to chaotic behavior that occurs during a change (Carpenter et al., 2011; Dakos, van Nes, D’Odorico, & Scheffer, 2012; although see Hastings & Wysham, 2010). This may be part of a “critical slowing” in a system’s ability to recover from impulses (Scheffer, Carpenter, Dakos, & van Nes, 2015; Scheffer et al., 2009). The idea is that feedback mechanisms can become coupled in a system, causing it to become chaotic (Brock & Carpenter, 2010), wherein impulses are amplified or “accelerated” (similar to Bernanke & Mihov, 1998; Kiyotaki & Moore, 1997, 2002).
For example, consider people who experience multiple impulses in succession, such as job loss and a spouse’s death. Variability in emotions may increase as people try to cope, and AR effects may increase as emotions are no longer mean-reverting and people slip into depression (Van de Leemput et al., 2014). Such regime changes complicate causal inference and can be expected in complex systems subjected to random events in the form of impulses (Clements & Hendry, 2001; Hendry & Mizon, 2005; Stock & Watson, 1996).
There are multiple ways to handle regime changes, such as with time-varying AR, MA, CL, and CLMA terms to reflect parameter changes (Bringmann et al., 2016), while keeping in mind that this makes a model sensitive to noise (Boldea & Hall, 2013; Perron, 2006; Stock & Watson, 2009). As with trends, there is no magic bullet for regime changes, and their existence is often uncertain (Badagián, Kaiser, & Peña, 2015). Our model can account for some regime changes with an occasion effect
Causal Inference Under Uncertainty
Even when tackling trends and regime changes, our approach is not without criticism, typically because it does not model the effects of randomly assigned interventions (Holland, 1986; Rubin, 2011). Without this, we theorize impulses as being akin to random assignment (see Lütkepohl, 2013; Sims, 1980, 1992; Stock & Watson, 2005, 2011). Yet, the validity of this theorizing is debatable, as in economics where GDP impulses are said to be due to improved technology (Christiano, Eichenbaum, & Evans, 1999). Also, interpreting impulse responses is complicated by correlated impulses (i.e.,
However, such concerns should not derail using models like the GCLM. Consider that many researchers use cross-sectional regression methods, which in our model can be done by regressions among impulses
For these reasons, we focus on where change seems possible (as in Hamaker, 2012; Molenaar, 2004). For this, we emphasize impulses, which is useful for variables subject to random variation and difficult to be experimented on (Aalen, Røysland, Gran, & Ledergerber, 2012; Dominici, Greenstone, & Sunstein, 2014; Granger, 1980, 1986, 1988, 2003). Of course, this approach has assumptions, but all methods have assumptions that must be balanced with their uses (Cartwright, 2007, 2009; Freedman, 2004; Sekhon, 2009). Even experiments have been criticized because they do not describe how to translate effects into interventions across contexts (Cartwright, 2011, 2012; Cartwright & Munro, 2010; Deaton & Cartwright, 2016).
The problem with all methods for causal inference is that their aim is to guide real-world action, but the consequences of action can never be predicted with certainty (Schön, 1995; Stone, 1989). Thus, even idealized methods such as randomized controlled trials (RCT) cannot enable unconditional inference because there are always gaps between evidence, action, and its consequences (Deaton & Cartwright, 2016). As Cartwright and Hardie (2012) explain: You want evidence that a policy will work here, where you are. Randomized controlled trials do not tell you that. They do not even tell you that a policy works. What they tell you is that a policy worked there, where the trial was carried out…. Our argument is that the changes in tense—from “worked” to “work” to “will work”—are not just a matter of grammatical detail. To move from one to the other requires hard intellectual and practical effort. The fact that it worked there is indeed fact. But for that fact to be evidence that it will work here, it needs to be relevant to that conclusion. To make RCTs relevant you need a lot more information. (p. ix)
Supplemental material
Supplemental Material, From_Data_to_Causes_I_-_Online_Appendices_final_version - From Data to Causes I: Building A General Cross-Lagged Panel Model (GCLM)
Supplemental Material, From_Data_to_Causes_I_-_Online_Appendices_final_version for From Data to Causes I: Building A General Cross-Lagged Panel Model (GCLM) by Michael J. Zyphur, Paul D. Allison, Louis Tay, Manuel C. Voelkle, Kristopher J. Preacher, Zhen Zhang, Ellen L. Hamaker, Ali Shamsollahi, Dean C. Pierides, Peter Koval and Ed Diener in Organizational Research Methods
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by Australian Research Council’s Future Fellowship scheme (project FT140100629).
Supplemental material
Supplemental material for this article is available online at http://journals.sagepub.com/doi/suppl/10.1177/1094428119847278 and
.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
