Abstract
Paul Meehl’s famous critique detailed many of the problematic practices and conceptual confusions that stand in the way of meaningful theoretical progress in psychological science. By integrating many of Meehl’s points, we argue that one of the reasons for the slow progress in psychology is the failure to acknowledge the problem of coordination. This problem arises whenever we attempt to measure quantities that are not directly observable but can be inferred from observable variables. The solution to this problem is far from trivial, as demonstrated by a historical analysis of thermometry. The key challenge is the specification of a functional relationship between theoretical concepts and observations. As we demonstrate, empirical means alone cannot determine this relationship. In the case of psychology, the problem of coordination has dramatic implications in the sense that it severely constrains our ability to make meaningful theoretical claims. We discuss several examples and outline some of the solutions that are currently available.
In 1978, Paul E. Meehl offered a scathing criticism of psychological science. According to Meehl, psychologists were busy occupying themselves with theories that were both “scientifically unimpressive and technologically worthless” (p. 806). The consequence of such an activity is an impediment of cumulative theoretic progress, with entire research communities trapped in vicious cycles in which theories never die but simply fade away (see also Newell, 1973). Behind this unfortunate state of affairs, Meehl argued, was psychologists’ tendency to overlook basic considerations regarding the falsifiability of theories and the inappropriate use of null-hypothesis testing.
The goal of the current article is to relate Meehl’s critique of psychology’s theory-testing practices to the problem of coordination that scientists, historians, and philosophers have discussed for well over a century (e.g., Chang, 2004; Mach, 1896/1986; Reichenbach, 1958; Tal, 2017; Van Fraassen, 2008). 1 We argue that by not addressing this problem, psychological scientists have compromised their ability to assess the relative merits of competing theories, resulting in a proliferation of theoretical concepts or phenomena for which there is little or no actual evidence. Relying on historical and philosophical analyses of thermometry (Chang, 2004; Mach, 1896/1986; Sherry, 2011), we make the case that the answer to the problem of coordination involves a careful and systematic joint development of theoretical models and experimental knowledge. Finally, we discuss readily available testing approaches that sidestep the problem of coordination.
The Falsification of Theories in Psychology
Let T denote the theoretical construct under investigation. For example, T could be a statement about whether a particular activity is governed by a single or dual cognitive process. Let A denote the auxiliary assumptions, such that when considered jointly with T gives rise to a set of predicted outcomes O. The assumptions in A may include both common statistical assumptions (e.g., independence of responses) as well as other elements regarding how constructs in T relate to observations, such as linearity assumptions among independent variables (see Kellen, 2019). The interplay between these concepts lies at the heart of our critique.
The falsifiability of any given theory T, along with auxiliary assumptions A, presupposes the ability to differentiate between the set O of outcomes deemed permissible and the complementary set Oˉ of those that are not. Modus tollens can then be invoked to falsify the conjunction T & A: If T & A is true, then O. We observe Oˉ. Therefore, T & A is false.
The falsifiability of T & A can be low because of the small size of Oˉ relative to O. For example, consider a theory stating that two population means, from a continuous dependent variable, are not equal. Such a theory is vacuous given that Oˉ is a single point on a continuum. Careless consideration of Oˉ can lead to theories that are unlikely (or impossible) to be falsified. However, a relatively large Oˉ does not necessarily mean that T is now easily falsifiable. After all, the falsification of the conjunction T & A can be attributed to a failure of one or more of the auxiliary assumptions in A (Duhem, 1954; Quine, 1963). Borrowing language from Lakatos (1976), A effectively serves as a “protective belt” over T, saving it from falsification. This situation leads researchers to engage in an iterative process in which A is scrutinized and amended before any determination on the merits of T is made (Lakatos, 1976; Meehl, 1990).
Alternatively, one can try to make a case for T by appealing to the falsification of a complementary theory T̅ using a modified logical argument:
2
If T̅ & A is true, then Oˉ. We observe O. Therefore, T̅ & A is false. Therefore, either T is true or A is false.
At the center of Meehl’s (1978) critique is the fact that these important considerations are often ignored or misunderstood by psychological scientists, who merrily entertain vague theories without “sufficient conceptual power (especially mathematical development) to yield the kinds of strong refuters expected by Popperians, Bayesians, and unphilosophical scientists in developed fields like chemistry” (p. 829). To make matters worse, the kind of testing psychological scientists often engage in involves a degenerate form of the modified logical argument given above. Specifically, they test null hypotheses that are trivially false and whose alternatives have little connection with any target theory. Meehl stated that “if you have enough cases and your measures are not totally unreliable, the null hypothesis will always be falsified, regardless of the truth of the substantive theory” (p. 822). “All sorts of competing theories are around,” he continued, “including my grandmother’s common sense, to explain the nonnull statistical difference” (p. 824).
Meehl contrasted this problematic practice with the kind of testing found in the “hard” sciences, in which the alternative hypothesis stands in close relation with a substantive candidate theory: “The logical distance, the difference in meaning or content, so to say, between the alternative hypothesis and substantive theory T is so small that only a logician would be concerned to distinguish them” (p. 824).
The fact that Meehl’s critique is now more than 40 years old presents itself as an opportunity to revisit some of its main points. At first blush, the fact that we encounter theoretical tours de force making a number of precise predictions (e.g., Cox & Shiffrin, 2017) suggests that things have improved considerably. Our point of contention here is that some of the progress in psychology as a whole is apparent only because it is predicated on a misunderstanding of the distinction between theory and auxiliary assumptions. More specifically, some elements of A, whose specific purpose is to bridge the “deductive gap” between theoretical and observational statements, are assumed to belong to T and/or T̅ without proper justification. As a result, these elements will not be scrutinized and refined by researchers, as envisioned by Lakatos (1976). Instead, they will be left untouched, as they are (illegitimately) seen as part of the theories’ “hard cores.”
One consequence of such misunderstandings is the spurious rejection of viable theoretical accounts and the latent-variable structures they propose. For example, Stephens et al. (2018) showed that previous rejections of single-process theories of syllogistic reasoning (i.e., T̅ ), taken as supporting a dual-process account (i.e., T ), hinge on auxiliary assumptions (e.g., a linear relation between latent processes and performance) that are simply taken for granted. When these assumptions are relaxed, it can be shown that the data at large are successfully captured by a single-process account (i.e., the different dependent variables can be described by single latent variable). 3 The problem identified by Stephens et al. (2018) is that previous attempts to test these theories illegitimately considered certain aspects of elements of A as part of T and/or T̅, which in turn results in a minimization of Ō. What this means is that single-process theories are being set up to fail, the end result being the false idea that a successful characterization of the data requires the involvement of two or more processes (i.e., latent variables).
Another consequence is the overstatement of support for certain theories. The empirical success of a conjunction T & A can be quite impressive when O is small. However, it is important to disentangle the contribution of the different elements in T and A to the size of O. Otherwise, one might erroneously attribute the success to the theoretical statements in T when in fact most of the legwork is being accomplished by A. One example of such a situation was identified by Jones and Dzhafarov (2014), who showed that the long-celebrated family of diffusion and ballistic-accumulator models, which are used to obtain precise joint descriptions of response frequencies and latencies, is not falsifiable until auxiliary parametric assumptions are introduced (e.g., the growth-rate variability between trials follows a Gaussian distribution). In other words, the empirical success of T alone is a sure thing.
The Problem of Coordination
To better understand the challenge of establishing and a justifying a precise relationship between theoretical and observational statements, it is useful to frame our discussion within the context of measurement. In a nutshell, measurements are statements in which two quantities are placed in relation to each other according to an established set of rules. The problem of coordination refers to the circular relation that exists between the meaning of an unobservable quantity and its measurement:
Let X be a postulated quantity that is not directly observable.
Let Y be a directly observable quantity that is connected to X by a coordination function f(·), such that Y = f(X).
To measure X through Y, we must know the coordination function f(·) that maps the former onto the latter. The problem is that this function is both unknown and unknowable. It cannot be established empirically because that would require knowing joint instances of Y and X, the latter being the unobservable quantity that we were trying to measure in the first place. Therefore, the coordination function must be defined by a theory instead of being discovered through empirical means.
The problem of coordination is endemic to measurement in all of science, not just psychology. It is therefore instructive to consider the way in which the problem was conceived, approached, and provisionally settled in a context in which measurement seems intuitively easy: the measurement of temperature. According to Mach (1896/1986), earlier efforts in scale construction in thermometry often framed these scales as attempts to approximate some Platonic idea of temperature. In Mach’s view, such a framing is misconceived because it overlooks the fact that any conception of thermal states can exist only by virtue of an arbitrary definition that coordinates them with empirical facts. In other words, the two questions “What counts as a measurement of X?” and “What is X?” cannot be addressed independently of each other (see Chapter 5 in Van Fraassen, 2008).
Although temperature is one of the physical magnitudes that people are most familiar with, it turns out that its measurement is far from trivial. In a body of work that spans over 200 years, we see that the process that led to the thermometers we know today involved a number of development stages, each associated with specific challenges (Chang, 2004; Mach, 1896/1986; Sherry, 2011). Likewise, people are deeply familiar with psychological concepts such as memory, attention, and intelligence because of the role these concepts play in our language (Maraun, 1998). Despite this familiarity, their measurement so far has proven to be frustratingly difficult (e.g., Borsboom et al., 2004; Maraun, 1998; Michell, 1999; Slaney, 2017).
Snapshots from the Development of Temperature Scales
Initial attempts to measure temperature led to the development of ingenious instruments known as thermoscopes. These instruments were based on the observation that most liquids expand with heat, which meant that their registered volume in a sealed container such as a glass tube could be used to determine whether the temperature of A is less than, equal to, or greater than that of B. In other words, the thermoscope provides us with an ordinal scale of temperature. Note that the development of thermoscopes hinges on the assumption that the relationship between volume and temperature is monotonically increasing. 4 A function f(·) is monotonically increasing if Xi ≤ Xj ⇔ f(Xi) ≤ f(Xj), ∀ i,j .
From the use of thermoscopes, we learned that different substances left in the same environment long enough will end up at the same temperature (i.e., they will reach thermal equilibrium), although they may feel differently. This is the so-called zeroth law of thermodynamics (Reif, 1965). The zeroth law enables thermoscopes to become thermometers by being calibrated against each other, as it allows the establishment of fixed points that can be used to set an origin as well as the units of the scale. But determining fixed points turned out to be extremely challenging, as exemplified by the discrepant measurements of the boiling point of water (e.g., 112.20 °C and 233.96 °F) when using different experimental apparatuses. 5 The solution found by a commission appointed by the Royal Society of London in 1776 was to define the boiling point of water as the value recorded when exposing a thermoscope to the steam emerging from the water. The rationale was that measurements obtained under this definition showed little variation across different experimental setups.
Unfortunately, the availability of fixed points did not solve the problem of coordination: To suppose that the points on the scale measure temperature, rather than just volume, is to assume a linear coordination, such that f(X) = αX + β, where α and β are free parameters. In fact, this assumption is rejected by the disagreements observed between thermometers using different liquids (e.g., water, alcohol, mercury, olive oil). These disagreements also showed that different substances have distinct coordination functions. This insight motivated the work of Henri Victor Regnault, who in the mid-1800s evaluated the merits of different thermometers through consistency testing. Regnault’s position was that if there is an attribute that takes on some value and we can specify different ways through which that value could be measured, then these measurements should all agree. Results showed a high degree of consistency between thermometers filled with gases such as hydrogen and air, namely linear relationships. Cases with poor agreement (e.g., sulfuric-acid gas) were deemed unsuitable as thermometric substances. Using the terminology introduced earlier, Regnault attributed any observed inconsistencies to A rather than T̅.
Regnault’s identification of linear relations between the pressures of the gases did not change the fact that their respective relations with temperature remained unknown. In other words, the problem of coordination stood unresolved. However, this should not come as a surprise: Any attempt to measure X requires us to make a statement about what X is, something only a theory can ultimately provide. In the case of thermometry, the provisional theoretical solution came in the form of the kinetic theory of heat, which redefines temperature as the average kinetic energy of the particles of an ideal gas. On the basis of this theory, it became possible to establish a linear relationship between the volume of ideal (or nearly ideal) gases and temperature and subsequently use these results to precisely calibrate thermometers on the basis of other substances such as mercury (e.g., Beattie et al., 1941).
However, although the problem of coordination requires a theoretical response, this requirement does not imply that experimental studies such as the ones described above are in some way secondary. As discussed by Chang (2004, Chapter 5), the circularity inherent to the problem of coordination requires researchers to engage in epistemic iterations, a process in which successive stages of experimental knowledge and theoretical understanding each build on the preceding stages—from noticing that our feelings of warmth are (somewhat imperfectly) tracked by the volume of substances to the stabilization and refinement of experimental procedures. Each development makes the subsequent stage possible, furthering our scientific goals. However, these goals are not necessarily reducible to the pursuit of some kind of realist aspirations (e.g., Kellen, 2019). For instance, note how the definition of temperature offered by kinetic theory is nothing more than an abstraction within the theory’s logical space that happens to characterize the outcomes of a stabilized procedure. The average kinetic energy is not something that exists in the same sense that the “average person” does not exist (see Van Fraassen, 2008, Chapter 5).
The Problem of Coordination in Psychology
As in thermometry, latent variables in psychology give rise to observed variables by means of a coordination function. We can draw an analogy between the measurement of the temperature of a substance with the measurement of a psychological attribute of a person. In measuring temperature, a scientist chooses a procedure P that yields an observation Y that is theoretically related to the unknown temperature X. In the case of a thermometer, Y is the volume of the enclosed thermometric substance. Formally, Y = fP(X). The process is the same for psychological attributes. A scientist chooses a procedure P′ that yields an observation Y ′ that is theoretically related to the value X ′ of the attribute. Again, Y ′ = fP′(X ′). Figure 1 illustrates this analogy. For a concrete example, take the case of memory, a capacity that can defined in broad strokes as the ability to remember. The quantity, or accuracy, of remembering can be easily recorded using a variety of experimental procedures such as free recall or single-item recognition. Drawing a closer analogy, these procedures include a study phase analogous to heating in the sense that better studied items becomes more memorable (hotter). This increase is tracked by the responses given in the test phase (e.g., hit rates or recall rates), the same way that changes in the temperature of a substance are tracked by its volume.

Analogous relationship between thermometry and psychological measurement.
Looking back at our historical discussion of temperature, one might expect to find similar efforts in the development of measures of memory. But one would be disappointed. Nor do we find the kind of virtuous circularities that Chang (2004) alludes to when discussing the process of epistemic iterations. There are three issues that remain unresolved:
Whether different measurement procedures measure the same or different attributes. For example, many researchers might feel that apparently dissimilar procedures, such as cued recall and single-item recognition, may measure different attributes (i.e., types of memory), but fewer might hold that this is the case for more similar procedures, such as recognition memory applied to different classes of items (e.g., different random word lists or pictures of faces vs. pictures of houses). This issue was at the center of the implicit/explicit memory debates that once dominated memory research (e.g., Schacter et al., 1993).
Even if two procedures are taken to measure the same attribute, they may have different coordination functions. This is equivalent to having a thermometer whose measurements hinge on the thermometric substance inside its vessel (e.g., water, mercury, air). In memory research, when very different procedures are compared, such as cued recall and recognition, no one would suppose that a score of say 70% correct responses on each task can be taken as measuring the same strength of memory. However, when more similar procedures are compared, such claims are frequently made (as we demonstrate below).
Even if two procedures are taken to have the same coordination function, this function remains unknown. We encountered this issue earlier when referring to Regnault’s consistency tests. In psychology, this issue leads to the interpretability problems discussed by Loftus (1978) and more recently by Wagenmakers et al. (2012) and Garcia-Marques et al. (2014), in which the understanding of data in terms of interactions and main effects obtained with analysis of variance (ANOVA)-type decompositions is shown to collapse under alternative coordination functions.
Operationalism about coordination
One of the reasons why the development of measures of psychological attributes differs from what is found in other domains such as thermometry is the fact that psychological scientists have by and large adopted (even if tacitly) a peculiar view of measurement known as operationalism (Bridgman, 1927), according to which measurement is simply defined as “any precisely specified operation that yields a number” (Dingle, 1950, p. 11). Operationalism was popularized in psychology by S. S. Stevens when he distinguished four different types of scales (nominal, ordinal, interval, and ratio) on the basis of the operations that produced them (Stevens, 1946). It is this operationalist view of psychological measurement that justifies the assertion that a sum score of numerical ratings constitutes some measure of “something” (e.g., of attitudes, intelligence; for critiques, see Green, 2001; Koch, 1992; Leahey, 1980; Michell, 1999).
It is also this operationalist view that drives researchers to make strong theoretical claims on the basis of the fact that two different measures do not change in exactly the same way across experimental conditions. For instance, the observation that performance in one memory task is affected by an experimental manipulation—whereas performance in a second task is not—has been interpreted by many as evidence for the existence of separate memory systems (for a critical review, see Newell et al., 2011). The argument is that if the measures behave differently it must be because they measure different things. Note how this line of reasoning runs completely counter to Regnault’s: In this particular example, we are dealing with the conjunction T̅ & A, where T̅ is the theory that there is a single attribute measured by both tasks and part of A is the assumption that both tasks have the same linear-coordination function. When both T̅ and A hold, we should expect both measures to register exactly the same changes in performance across conditions. But if they do not (i.e., an interaction is observed), it follows that either T̅ or A (or both) are at fault. Whereas Regnault would interpret such a result as a failure of A, a psychologist operating under the tenets of operationalism would choose to reject T̅.
The reason behind this disagreement is that operationalism ignores the problem of coordination and effectively enforces a complete disconnect with natural reality: We are no longer concerned with the ability of Y capturing some unobservable X—quantity Y has become its own measure. It also follows that any concept of validity has to be discarded because we can no longer evaluate the consistency of measurements coming from different instruments: They measure different things simply because they involve different operations—a vicious circularity. 6 Under operationalism, many of the achievements observed in thermometry would have not been possible. In fact, operationalism would legitimize absurd claims, such as that temperature is a multivalued magnitude, on the grounds that the application of distinct thermometric instruments (e.g., pyrometers, thermocouples, thermometers) yield different numbers before calibration.
Case study: face-inversion effect
For a more concrete example of how the choice of coordination affects theory testing, consider the case of the face-inversion effect (Yin, 1969), which is illustrated in Figure 2. According to the face-inversion effect, people’s ability to recognize faces is more affected by inversion than other pictures of mono-oriented objects such as houses. One common interpretation of the face-inversion effect is that there is “something special” about the way we process facial stimuli (for discussions, see Dunn & Kalish, 2018, Chapter 2; Loftus et al., 2004).

Illustration of the paradigm used by Yin (1969). Pictures of faces and houses were studied upright or inverted and later tested in a two-alternative forced-choice recognition task (top). The observed accuracy (bottom left) and a characterization in which inverted/upright faces and houses have the same memory strength (bottom right) are also shown.
In a typical experiment, participants study items that may be either pictures of faces or houses and either upright or inverted. This conforms to a 2 × 2 factorial design. It is argued that if there is nothing special about faces then the effect of inversion should be the same for both faces and houses. Using the analogy with temperature, whereas upright faces and houses may have different “temperatures” after being “heated” by study (because faces are easier to study, i.e., absorb “heat” more readily, than houses), the “cooling” that comes from inverting them should decrease their final temperatures by the same amount. Let X be the memory strength of upright faces, Δhouse the difference in memory strength for houses, and Δinverted the (negative) change in strength due to inversion, respectively. Then
It follows that
Now, let f(·) be a coordination function. If it is linear, then the equality above holds. That is,
The accuracy data in Figure 2 show that
is larger than
indicating that the effect of inversion on memory is greater for faces. This result, which would be captured by interaction effects in a linear model (e.g., via ANOVA or regression), is used to support the claim that there is something special about the way we process and remember faces.
However, as discussed above, assuming a single linear coordination for both faces and houses lacks justification. One possible remedy is to retreat to the more modest idea that the relationship between accuracy and memory strength is monotonically increasing. This move is equivalent to admitting that we have only a “memory thermoscope” available. To begin with, we may suppose that faces and houses have the same monotonically increasing coordination function, f(·). In this case, the equality of differences is replaced by the following two implications, which comprise inequalities:
and
Because the second of these inequalities is violated by the data shown in Figure 2, we could conclude that there is something special about faces. However, this conclusion depends crucially on the assumption of a common coordination function. This too may be relaxed by assuming that faces and houses have different coordination functions. In this case, it is perfectly possible for the memory strength of studied houses and faces to be the same and for them to be equally affected by inversion. Under this assumption, the potential violation of T̅ is attributed to A, and we can no longer stand by our earlier statement that there is something special about our memory for faces simply because they show a larger inversion effect in a recognition task. It is analogous to saying that there is something special about the temperature of water relative to that of mercury simply because these substances expand differently when exposed to the same heat, effectively ignoring the problem of thermoscope/thermometer calibration (we also know of no attempts to establish fixed points). Finally, note that a similar case can be made regarding the comparison of different groups of individuals: For instance, Laguesse et al. (2012) compared the size of the face-inversion effect in two groups of participants and found a larger effect for one group than the other, which they interpreted as demonstrating that the groups differed in their sensitivity to inversion. Although such an interpretation is not illegitimate, it once again depends on an unjustified assumption of a linear-coordination function (for a detailed discussion, see Dunn & Kalish, 2018).
The face-inversion effect illustrates how the problem of coordination can severely affect the conclusions that can be drawn from observed patterns of data. A theory T̅ is proposed with the goal of serving as a kind of null hypothesis—a statement about the world that the researcher seeks evidence against. This may take the form of proposing that there is only one kind of memory, or that the effects of inverting an image on memory for that image are the same for faces and houses, or that the effect differs across groups of people (e.g., younger and older adults). As noted earlier, such a point hypothesis is vacuous because it rules out almost nothing—Oˉ corresponds to a single point. By coupling T̅ with a linear-coordination function, the minimal size of Oˉ is thereby maintained, and, as more data are collected and the point null is rejected, the theory of interest T is apparently supported. However, by considering more general and more plausible candidates for coordination functions, the outcome predicted by Oˉ is no longer a single point and so is less readily falsified. This response is nothing more than the textbook prescription that researchers examine all aspects of their experimental setups (included in A) before drawing theoretically significant conclusions from the observed data. Finally, the face-inversion effect demonstrates that the way psychological scientists engage with attributes such as memory is completely at odds with the experimental and theoretical developments found in the case of thermometry, a situation that partly results from the persisting influence of operationalism. It also is at odds with Meehl’s (1978) call for a more careful consideration of the link between testing outcomes and theoretical statements.
What Can Be Done?
The testability of a theory is codetermined by the structure of the latent variables that it postulates (e.g., their involvement in different dependent variables) and the coordination functions that are imposed (Dunn & Anderson, 2018). The latter determine which transformations can and cannot affect the mapping of the latent structure onto observations. This relationship between latent structure and coordinations shows that the problem of coordination is not limited to measurement—it is also a problem for theory testing. At this point, it is not clear whether we can overcome the many problems of coordination found in psychology. On the one hand, the replication of some of the achievements found in thermometry—the establishment of fixed points and calibration of scales—seems extremely unlikely. 7 On the other hand, some of the attributes that psychological scientists are interested in (e.g., intelligence, anxiety, dominance) have extremely complicated “grammars” that can seriously compromise their measurability (for a discussion, see Maraun, 1998). But if the history of psychological measurement tells us anything, it is that powerful advances can be achieved when rigorous thinkers are willing to put some intellectual muscle into the enterprise (Rozeboom, 1966).
In any case, it is the responsibility of researchers to specify the nature of their assumed coordination functions. If the interpretation of the empirical results offered by the researcher depends on a specific coordination function, such as one that is linear, then it is only reasonable to expect that this be made explicit, as any assumption would be. However, in so doing, it should become apparent that some hyperspecific coordinations (such as those that are linear) are only rarely justifiable. 8 For this reason, researchers may propose coordinations with fewer unjustified commitments in alternative or complementary methods that assume different coordinations (for a recent example, see Kellen et al., 2020). Depending on the proposed coordination function(s), the outcome Oˉ associated with T̅ will change accordingly. Point hypotheses involving relationships of equality or additivity will not survive any departure from linearity—other implications will have to be worked out. 9
Restricting ourselves to the assumption that coordination functions are monotonic seems to be a reasonable option in most contexts. 10 Fortunately, there are a number of readily available methods that require only the assumption of monotonicity. For instance, signed-difference analysis (Dunn & Anderson, 2018; Dunn & James, 2003) can be used to identify the structural properties of a theory’s latent variables that hold when assuming monotonic coordinations. These structural properties are observable in the directions (+, −, and 0) in which the observable variables can jointly change across conditions in a given experimental design. These differences are described by sign vectors. Note that one special case of signed-difference analysis is state-trace analysis (Bamber, 1979; Dunn & Kalish, 2018), which focuses on whether two dependent variables can be described by a single latent variable with monotonic coordinations. Stephens et al. (2018) applied signed-difference analysis to a corpus of studies used to compare single- and dual-process models of syllogistic reasoning. For example, they requested participants—under deductive or inductive instructions—to judge syllogisms that varied dichotomously in terms of both their validity and causal consistency. The endorsement rates that came out of these studies can be boiled down to 81 sign vectors. Each element of the vector corresponds to the sign of the difference in endorsement rates between causal-consistency conditions. For instance, consider the sign vector
Under monotonic coordinations, different single-process and dual-process theories create different partitions of sign vectors into O and Oˉ. For instance, the sign vector described above cannot be captured by any single-process theory and most dual-process theories. The inconsistencies between the observed differences and each theory’s partition of sign vectors were tested statistically using the order-constrained inference method proposed by Kalish et al. (2016). 11 Results showed that most theories are rejected, including all testable dual-process theories.
One concern often raised when discussing the use of weaker coordination assumptions is that theories become harder to test. Our immediate response to such concerns is to point out that the “increased testability” that one might be reluctant to forfeit was obtained through illegitimate means. Having said that, it is incorrect to assume that one cannot devise strong tests in the absence of stronger coordination assumptions. Strong tests can be devised by using richer experimental designs. 12 For example, Kellen et al. (2019) and McCausland et al. (2020) tested (and upheld) the postulates of signal detection theory and random utility theory using experimental designs for which the O of both theories is minuscule.
Another reason why researchers might feel reluctant to engage with weaker coordinations is the fact that the equations used in traditional ANOVA-type decompositions (main effects and interactions) are not suited to handle the now predicted inequalities. In response, we raise three points. First, many of the psychological theories that researchers engage with make predictions at the ordinal level (i.e., they predict one or more inequalities). In fact, we would argue that the ANOVA-type language imposed by many standard statistical methods has hindered psychological scientists in the sense that it very often interferes with their ability to think clearly and speak plainly about theoretical predictions and their encounter with data (e.g., Hatz et al., 2020; Rouder et al., 2019). Second, the joint test of the order constraints postulated by a theory allows for powerful omnibus tests that would not be possible with a piecemeal approach in which multiple main-effect and interaction tests would be necessary (for an excellent example, see Iverson, 2006). Third, the resources needed to apply order-constrained inference methods are now readily available (see Heck & Davis-Stober, 2019; Kalish et al., 2016; Regenwetter & Cavagnaro, 2018).
Conclusion
According to Lakatos (1976), auxiliary assumptions A serve as a protective belt around a theory. In a research program that is theoretically and empirically progressive by continuously resolving previous anomalous findings and confirming novel predictions, these auxiliary assumptions are constantly being analyzed and refined. Meehl (1978) despaired of psychological scientists attempting to test theories against null-hypotheses that are trivially false. The consequence of such practice is the development of degenerative research programs in which the rise and fall of theories is more reflective of fads and fashion than meaningful theoretical development.
In the current work, we argued that some of the problematic practices criticized by Meehl are still present despite apparent progress. We attribute this persistence to a continued neglect toward the problem of coordination, which leads to spurious support for or opposition to certain theories. Using the history of thermometry as a reference, we showed that the linear coordinations typically used are often implausible and never justified, giving researchers a false notion of precision, falsifiability, and empirical support (Chang, 2004; Sherry, 2011). This neglect corrupts the Lakatosian process of scientific development because it prevents researchers from investigating the impact of their assumed coordinations by incorrectly attributing their failures to the theories. As a first step, researchers should restrict themselves to assuming only monotonic coordinations that do not necessarily generalize across procedures. Fortunately, there is a rich toolbox of methods that can operate under such minimal structural constraints. The use of such methods, along with more plausible coordination functions, will result in less falsifiable T̅. This in turn will render theories T appropriately falsifiable, removing an important impediment to cumulative theoretical progress in psychology.
