Abstract
The issue of causal comparability in the social sciences underlies matters of both generalization and extrapolation (or external validity). After critiquing two existing interpretations of comparability, due to Hitchcock and Hausman, I propose a distinction between ontological and epistemic comparability. While the former refers to whether two cases are actually comparable, the latter respects that in cases of incomplete information, we need to rely on whatever evidence we have of comparability. I argue, using a political science case study, that in those cases of imperfect information, an epistemic homogeneity criterion can be an adequate justification for generalization.
1. Introduction
A key question in causal inquiry in the social sciences is causal comparability, that is, when we can fruitfully compare the causal relationship between a putative cause
Demanding causal comparability in the social sciences faces two main problems. First, arguably all instances of social phenomena have at least some causally relevant differences (cf. Kuorikoski 2012; Little 1991), and thus a population of cases under study may not contain similar enough individual cases to draw general conclusions. One may therefore argue that general conclusions in the social sciences are often unwarranted. The second issue with demanding causal comparability is that one almost never has a complete causal picture of the social phenomena under study (cf. Steel 2008), especially given the issues of multiple causation and INUS factors (cf. Cartwright and Hardie 2012). Without knowing what kinds of factors may interfere with
Although we may take these two problems, the lack of true comparability and lack of evidence for comparability, as a reason to stick with single-case research and avoid generalization altogether, this solution is unsatisfactory for some purposes; causal comparability is required for any kind of generalization, and generalization in turn is important for both theorizing and policy purposes. A push to forego causal explanations at the population level altogether is incompatible with social scientific practice.
How, then, should we analyze causal comparability? The main focus in the debates on causal comparability centers on criteria for comparability themselves. Broadly speaking, there are two interpretations of what causal comparability amounts to. One derives from Christopher Hitchcock’s highly theoretical characterization of the relationship between singular and general causation, the “causal homogeneity condition,” and the other from a more pragmatic characterization of this relationship by Daniel Hausman, the “average effect condition.” However, as I will show in this paper, both interpretations are problematic in the social sciences. The former is too strong, because Hitchcock requires a strict similarity between different individual cases that is almost never met. The latter is too weak, because Hausman cannot adequately describe causal explanation in cases in which a population contains many causally dissimilar subpopulations.
This paper moves the causal comparability debate on by starting from our limited causal knowledge of social phenomena. I argue that Hitchcock and Hausman’s interpretations are more pertinent to different epistemic situations: in the former, we are aware of all the relevant probabilities for individual cases in our population of interest and so we have evidence for the stronger generalization, and in the latter, we are unaware of the specific probabilities and so we merely have evidence for a weaker generalization. I then take this analysis further. Following an older distinction by Wesley Salmon, I distinguish between ontological and epistemic comparability. The former has been the focus point of debates in causal comparability, internal validity, and external validity; ontological comparability refers to whether two cases are actually comparable. Epistemic comparability, on the other hand, has been relatively ignored.
While ontological comparability refers to whether two cases are actually comparable, epistemic comparability respects that in cases of incomplete information, we need to rely on whatever evidence we have of comparability. I argue, using a political science case study, that in those cases of imperfect information an epistemic homogeneity criterion can be an adequate justification for generalization. I carefully distinguish different types of epistemic comparability: first, depending on whether one is concerned with comparability relative to one causal factor or several (“relative” vs. “total” epistemic homogeneity), and second, depending on whether the assumption of homogeneity is due to practical constrains on time and effort, or due to a true lack of knowledge (“pragmatic” vs. “true” epistemic homogeneity). This paper thus provides a framework for describing how generalization and extrapolation in the social sciences actually happen. I illustrate my claims by critically analyzing Nicholas Sambanis’ theory on ethnic civil wars using my “epistemic homogeneity” account. A direct consequence of my new taxonomy is that one can fruitfully discuss the comparability of cases even in the scenario that individual cases in a sample are not entirely similar. Using the notion of epistemic homogeneity, I argue that we can generalize even in such scenarios.
2. Two Interpretations of Comparability
To introduce the first interpretation of comparability, the Hitchcock causal homogeneity condition, consider the general causal claim that “smoking causes cancer” for a set of individuals including both John and Mary (Hitchcock 1995). Hitchcock interprets this general causal claim as saying that the relation between smoking and cancer is sufficiently similar for John and Mary. In brief, he argues that when we know that every individual will respond in the same way to the causal variable being at a certain value, all other things being equal, then we can make a general causal claim. John and Mary, all other things being equal, are equally likely to develop lung cancer when they smoke a certain number of packs per day.
More formally speaking, for Hitchcock’s causal homogeneity condition, we need the following ingredients. Hitchcock constructs a “little probability space” {Ω
i
,
The Hitchcock causal homogeneity condition is a very strong demand for social science generalization. First, it is difficult to find evidence that its requirements are met in any actual cases of interest; it will be difficult to calculate the individual probability spaces {Ω
i
,
Hausman shows that a general claim such as “smoking causes lung cancer” will be technically false if we interpret it, as Hitchcock does, as saying that the probability of lung cancer increases for every individual who smokes, all other things being equal. As Hausman points out, there are individuals whose smoking leads to a fatal heart attack before they can even develop lung cancer; these individuals do not fit with the general causal claim as defined by Hitchcock. After all, their background context will be different from the background context of individuals who do not develop a fatal heart attack. Call the set of all individuals whose smoking leads to a fatal heart attack before they develop lung cancer
This seems counterintuitive if, despite some individuals dying of heart attacks first, we still want to say that smoking causes lung cancer. For that reason, Hausman argues for his alternative characterization of the relation between singular and general causal claims, the average causal effect. Using the same notation as before, under this characterization, in population
Note that this interpretation of generalization does not ask us about any individual conditional probabilities for the
Although the average effect condition can be useful in those practical contexts, there are problems with this characterization of comparability. If there are many different subpopulations of
3. Epistemic Homogeneity Defined
There exist, then, two different interpretations of comparability, each of which has its issues. One way of construing this difference is to say that the two interpretations are more pertinent to different epistemic situations: in the one, we are aware of all the relevant probabilities for individual cases in our population of interest and so we have evidence for the stronger generalization, and in the other, we are unaware of the specific probabilities and so we merely have evidence for a weaker generalization. 2
Synthesizing the two interpretations, in the rest of this paper, I will argue for distinguishing epistemic homogeneity and ontological homogeneity. The latter has been the focus point of debates in causal comparability, internal validity, and external validity; the former, on the other hand, has been relatively ignored. In the second half of this paper, I show how epistemic comparability respects that in some situations we may not know the full story and thus need to rely on whatever evidence we have of comparability, and I carefully distinguish different types of epistemic comparability. I illustrate my claims by critically analyzing Nicholas Sambanis’ theory on ethnic civil wars using my “epistemic homogeneity” account. This paper thus provides a framework for describing how generalization and extrapolation in the social sciences actually happen.
3.1. Wesley Salmon’s Reference Class Rule
The notion of epistemic homogeneity that I will defend here is based on an earlier argument by Wesley Salmon. Salmon first proposed the distinction between objectively and epistemically homogeneous reference classes in an attempt to solve the reference class problem for probabilistic explanation (Salmon 1971). After introducing Salmon’s argument, I will use it to give an account of causal inquiry.
The reference class problem is the issue that calculating a probability of an event requires us to specify a reference class for that event, while there is no straightforward “correct” reference class but rather a variety, each of which gives a different probability. For example, to calculate the probability that John will die of a heart attack, we must know what reference class to use: the class of all Caucasian men? All middle-class people? All professional tennis players?
Salmon solved the reference class problem by arguing for the use of the “broadest homogeneous reference class.” Let me first introduce the concept of a homogeneous reference class with an example. Assume we are investigating a class of events or phenomena,
Consider a few more examples. The class of all pregnancies is not homogeneous for the property of babies born with physical, developmental, and functional problems, because we can further divide the class according to the place selection that picks out mothers who drink more than two glasses of wine per day. The class of university graduates is not homogeneous for the property of income, because we can further divide the class according to the place selection of gender.
A first thing to note is that, at least in the university graduates example, we could also partition the class with other place selections, for example, in this case according to social class. It is not always clear which of these classes we should work with, which is why Salmon argues for investigating “the broadest homogeneous reference class to which the single event belongs” (Salmon 1971, 43), that is, the homogeneous class with the most members. For Noa, a female graduate from a working-class background, this homogeneous class will be the class of all female working-class graduates. For Clarence, a male graduate from an upper-class background, this homogeneous class will be the class of all male, upper-class graduates. Note that if it were to later turn out that the highest level of education attained by Noa and Clarence’s parents also matters to Noa and Clarence’s income, this property would have to be included in the delineation of their homogeneous class, and thus we would have to narrow the homogeneous classes further.
A second thing I wish to stress is that in Salmon’s framework, a class
Problematically, in the social sciences, researchers are often unaware of the full causal picture, that is, which properties are relevant for the cause and effect under discussion, and which are not. Salmon’s suggestion, in those cases of “incomplete information,” is as follows: “[w]hen we know or suspect that a reference class is not homogeneous, but we do not know how to make any statistically relevant partition, we may say that the reference class is epistemically homogeneous” (Salmon 1971, 44).
How does this solution to a problem for probabilistic explanation relate to the interpretations of generalization I have outlined in Section 2? As argued there, the two interpretations are meant to apply to different epistemic situations. In Hitchcock’s interpretation, we are aware of all the relevant probabilities for individual cases in our population of interest, and in Hausman’s interpretation, we are either unaware of, or choose to ignore, 3 the specific probabilities. If we take Hausman’s interpretation as a suggestion to look for epistemic homogeneity rather than ontological homogeneity, that is, to assume homogeneity until proven otherwise, we will be able to make epistemic progress regarding a set of cases.
3.2. Epistemic Causal Homogeneity in a Formal Framework
Let me return to the framework from Section 2, in order to clarify what epistemic homogeneity would amount to there. A general causal claim, under Hitchcock’s interpretation, can be made over the set of those individuals for which the little probability space is the same, that is, for those sets
So far this framework cannot deal with cases of incomplete information. This is a problem because researchers are hardly ever in possession of all relevant information about the distribution functions
So, in general, if researchers have incomplete information about the causal structure of the area they are investigating, they might not be able to come up with a partitioning
4. Does Causal Modeling Trivialize the Epistemic–Ontological Distinction?
Before I illustrate the distinction between ontological and epistemic homogeneity with a case study, I wish to clarify one potential misunderstanding. So far, I have argued that we are hardly ever aware of all the relevant probabilities for individual cases in our population of interest, but may not wish to settle for a weaker generalization that merely averages out over the probabilities in these individual cases. One response might be that causal modeling approaches can help us find all relevant variables for a particular causal structure (cf. Pearl 2000; Spirtes, Glymour, and Scheines 1993; Woodward 2003), and that as such the distinction between ontological and epistemic homogeneity is relatively trivial. In this section, I will show why this is an oversimplification by discussing a limitation of causal modeling that is closely related to the issue of causal homogeneity, focusing on Spirtes, Glymour, and Scheines’ account (cf. Spirtes, Glymour, and Scheines 1993). I will also briefly rebut a closely related objection, namely, that econometric methods that uncover heterogeneity in populations trivialize the distinction between ontological and epistemic homogeneity.
Key within Spirtes, Glymour and Scheines’ causal modeling technique is the notion of a causal search. A causal search is, in simple terms, a computer algorithm that takes statistical data (including a list of all measured variables) as its input and produces an equivalence class of all the causal structures that fit these data as its output. Such structures can then, in turn, be used to answer predictive questions; for instance, we may ask what would happen to a particular variable in the causal structure under an intervention. Moreover, and worth noting given the discussions in this paper, a causal search can show whether a data set is compatible with the existence of one or more unspecified, latent common causes. As such it can, in some cases, tell us whether there is or is not enough information in the data to answer causal questions.
This, I believe, is where a potential misunderstanding could arise. A critic may argue that Spirtes, Glymour, and Scheines’ techniques of causal modeling trivialize the epistemic–ontological distinction, because they allow us to detect hidden variables. Thus, the techniques allow us to break down the epistemic boundary between the variables we know might have an effect, and the ones we do not know about, but actually do have an effect. However, this response is an oversimplification of the issues at hand. Without going too far beyond the scope of this paper, I wish to evidence my response briefly by describing particular data sets that a causal search has difficulty with, and which arguably illustrate why causal modeling does not (yet) trivialize the problems of comparability in this paper.
In their 2004 paper “Causal Inference of Ambiguous Manipulations” (Spirtes and Scheines 2004), Spirtes and Scheines draw attention to an issue that could arise when one uses “defined variables” in a causal search. In some of those cases, one’s causal search includes measurements that define a variable as homogeneous while in fact (without our knowledge) it consists of two subvariables which have heterogeneous effects under manipulation. Manipulating one of the subvariables will give a different result than manipulating the other will; therefore, we call such manipulations “ambiguous” (Spirtes and Scheines 2004, 834).
Spirtes and Scheines give the example of the variable “total cholesterol” (call this
Given these difficulties with defined variables, I believe it is an oversimplification to say that causal modeling techniques trivialize the question of comparability and homogeneity in this paper. Causal modeling will struggle with ambiguous defined variables, which abound in social science. Think, for one, of the concept “democratization.” There are many different kinds of democratization and so if, like in the LDL/HDL example, states undergoing certain types of democratization behave differently than states undergoing other types of democratization, it is highly likely that a manipulation of democratization as a total category will lead to many “can’t tells.” As such, we need the distinction between ontological and epistemic homogeneity to accurately describe the issues faced when building general theories in the social sciences. 4
In a similar vein, a critic may argue that recent developments in econometric methods trivialize the distinction between ontological and epistemic homogeneity. In econometric terms, this paper describes a taxonomy of how one may describe (our knowledge of) populations in light of unknown independent variables, and how one can taxonomize populations in the social sciences even in light of the inherent variability between subpopulations. A critic may argue that some econometric methods (such as regression analysis) do just that; specifically, some methods are specifically attuned to find out the potential effects of any unknown variables by systematizing these variables as error terms. These methods tell us the ceteris paribus effect of the systematized independent variables of interest on the dependent variables, exactly as Hitchcock requires for his causal homogeneity condition.
However, arguably such econometric methods do not trivialize the distinction between ontological and epistemic homogeneity but rather underline it. Information about a sample and its statistics, including the ceteris paribus effect of independent variables on dependent variables, should be described in terms of epistemic homogeneity. This information should not be confused with a full causal picture of the population and its parameters, which we can describe in terms of ontological homogeneity.
Cases of selection bias provide a relevant illustration of when a sample and its statistics on the one hand differ from a population and its parameters on the other. In these cases, we are concerned with the possibility that the differences within a sample at hand damage a study’s internal validity, that is, the degree to which a causal conclusion is warranted for the entire sample studied. Although one can limit the influence of selection bias using methods such as instrumental variables and Heckman corrections (cf. Antonakis et al. 2010), these methods rely on specific assumptions about the data under scrutiny and thus we cannot always implement them. For some econometric models suffering from selection bias, “it is difficult to anticipate whether the biased regression estimates overstate or understate the true causal effects” (Berk 1983, 390; see also Xie 2011). Here, we can fruitfully employ the distinction between ontological and epistemic homogeneity. If we can control for selection bias, we make epistemic progress by further narrowing down the reference classes under study. If we cannot, we are in a situation of epistemic, and not ontological homogeneity.
To rearticulate the rebuttals in this section, this paper does not aim to comment on econometric solutions, nor on causal modeling approaches, but rather aims to provide a more foundational analysis, that is, of how one can describe inherently variable populations using the concepts of ontological versus epistemic homogeneity. We can then use this analysis to describe any epistemic progress made by econometric methods or causal models.
5. Example: Civil War Studies’ Search for Epistemic Homogeneity
In this section, I wish to discuss an example of the search for causal homogeneity, namely, the move from a general theory on civil war onset to a more specific theory on ethnic civil war onset by Nicholas Sambanis. I will show that the class of states at civil war was an epistemically homogeneous class with respect to several properties, including economic and political factors, until Nicholas Sambanis figured out a way to make a statistically relevant partition in the class, namely, between ethnic and non-ethnic civil wars. 5 I will make clear what the results of the partitioning of this class were, and thereby illustrate the notions of epistemic homogeneity and epistemic progress discussed above.
5.1. Civil Wars as an Epistemically Homogeneous Class
In 2001, Nicholas Sambanis asked whether ethnic and non-ethnic civil wars start for the same reasons. The ordinary theories in the civil war literature at that time assumed so: the then prominent “economic” theory of Collier and Hoeffler (2004) aggregated civil wars since 1960 into one class 6 and suggested that such wars start mainly because of economic factors (such as financial incentives for the rebels) rather than political factors (such as the level of democracy of the country and of its neighboring countries).
In the terminology introduced in the previous section, one may say that Collier and Hoeffler treated the class of all civil war onsets as an epistemically homogeneous reference class with respect to the properties under investigation (i.e., with respect to the potential causes). Sambanis’ contribution to the civil war literature was suggesting that there might exist a relevant partition, “ethnicity,” which divides the class of civil war onsets into causally dissimilar subgroups. In this section, I will highlight the differences between ethnic and non-ethnic civil wars in Sambanis’ framework. I briefly outline the aggregate theory of civil wars as presented by Collier and Hoeffler, and then discuss the properties relative to which Sambanis believes there is a partitioning which shows the ethnic heterogeneity of the class of all civil wars.
Collier and Hoeffler and Sambanis define “civil war” in a similar manner, based on the definition used by the Correlates of War database, one of the most commonly used sources for data in quantitative studies of war. Collier and Hoeffler do not distinguish between different types of civil war, but instead consider the class of civil wars as causally homogeneous. They discuss this decision in later writing, noting that although some data sets define conflicts in terms of the underlying issues (as is the case when scholars make a distinction between ethnic and non-ethnic civil wars) they have decided not to do so because “the classification of conflicts according to their causes does not seem helpful . . . if we want to analyze the causes of civil war” (Collier and Hoeffler 2001, 5).
In terms of the interpretation outlined in the previous section, Collier and Hoeffler consider the class of all civil wars,
5.2. Partitioning Civil Wars into Ethnic and Non-Ethnic Wars Relative to Ethnic Heterogeneity
In contrast, Nicholas Sambanis (2001) argues for partitioning
Sambanis defines ethnicity following earlier work by Donald L. Horowitz (1985), in which an ethnic group can be defined in terms of anything from “color, appearance, language, religion, some other indicator of common origin, or some combination thereof” (Donald L. Horowitz 1985, 17-18) and covers other terms like “tribes, races, nationalities, and castes” (Donald L. Horowitz 1985, 53). A civil war is an ethnic civil war, Sambanis argues, if the core issues in the conflict are “integral to the concept of ethnicity” (Sambanis 2001, 261-62). Or, alternatively, an ethnic war is a “war among communities (ethnicities) that are in conflict over the power relationship that exists between those communities and the state” (Sambanis 2001, 261). He codes a conflict as an ethnic civil war if it is a civil war (for exact requirements, see Sambanis 2001, 262) and if it counts as an episode “of violent conflict between governments and national, ethnic, religious, or other communal minorities (ethnic challengers) in which the challengers seek major changes in their status” (Sambanis 2001, 262).
In light of the discussion of Salmon’s reference class rule, whether one can call “ethnicity” a proper partitioning for the class of civil wars in relation to the property of ethnic fragmentation is dubious. At first glance, one may suspect that taking “ethnicity” as a place selection breaks the rule that a partition of class
On the other hand, ethnic fragmentation refers to the number of different ethnic groups within a state. Thus, it is not the case that the partition into “ethnic” and “non-ethnic” civil wars refers to the property of “ethnic fragmentation.” As Sambanis puts it, “not all wars that involve ethnic groups as combatants should be classified as ethnic wars. The issues at the core of the conflict must be integral to the concept of ethnicity” (Sambanis 2001, 261-62). A country can in theory have a low degree of ethnic fragmentation and still be at ethnic civil war; two ethnicities is all it takes. And, vice versa, a country can have a high degree of ethnic fragmentation without an ethnic civil war being fought there—and if there is a civil war being fought in the country, it may have started for different reasons. 8 For those reasons, one might argue that “ethnicity” is a proper partitioning, despite first appearances to the contrary.
However, the correlation between ethnic heterogeneity and ethnic civil war is arguably not unexpected given that both are defined in terms of ethnicity; they might not be interchangeable but they are closely related. Unless we can show the conceptual independence between the two, it is not useful to make a distinction between ethnic and non-ethnic civil wars merely on the basis that they have a different causal connection to the property ethnic fragmentation. What Sambanis needs to show, I would argue, is that ethnic wars are different from non-ethnic wars in a way that goes beyond their causal history of ethnic fragmentation.
Though it is not immediately obvious that Sambanis’ distinction between ethnic and non-ethnic wars is a proper partitioning for the property of “ethnic fragmentation,” there are other properties that Sambanis investigates. He shows that ethnic and non-ethnic civil wars also differ in relation to those other properties, that is, the class of “civil wars” is heterogeneous relative to other properties besides the (dubious) ethnic heterogeneity. Sambanis considers several such properties, including political variables such as the polity score of the country, the polity score of its neighboring countries, and economic variables such as real per capita income. He shows that while the polity score of a country is statistically significant for ethnic war onset, 9 its polity is non-significant to non-ethnic war onset (Sambanis 2001, 276). Moreover, real per capita income is more significant to non-ethnic war onset than it is to ethnic war onset, leading Sambanis to speculate that economic variables are a more important causal factor in non-ethnic war onset than in ethnic war onset. And indeed, Sambanis concludes that “[ethnic] wars are predominantly caused by political grievance, and they are unlikely to occur in politically free (i.e. democratic) societies” (Sambanis 2001, 280).
5.3. Lessons Taken Forward
So, I have shown in this section that civil war onset was treated as a homogeneous class, despite indications to the contrary from other areas of the civil war literature. Sambanis showed that for the property of ethnic fragmentation, the class could be fruitfully partitioned into two causally dissimilar subgroups, that is, ethnic and non-ethnic civil wars. I have analyzed to what extent this partitioning is a proper partitioning. I have also shown that Sambanis’ partitioning into ethnic and non-ethnic civil wars was relevant for other properties which are more straightforwardly independent of ethnicity, that is, political and economic factors. If we accept Sambanis’ statistical results, we must also accept that general theories of the causes of civil war cannot have all civil wars as their scope; the scope conditions had to be limited to ethnic civil wars or non-ethnic civil wars (not both). 10
This case study of Sambanis can be taken to reiterate that we must check for the independence of the place selection and property, a requirement I introduced in my analysis of Salmon’s reference class rule in Section 3.1. But the case study also highlights a relevant distinction in the notion of epistemic homogeneity: whether we consider epistemic homogeneity relative to one particular property, or total epistemic homogeneity. In the former case, researchers are merely interested in a particular cause, and in the latter, researchers aim to find a complete causal picture of a particular social phenomenon. I will discuss this distinction in more detail below.
6. Further Refinements of the Notion of Epistemic Homogeneity
6.1. Classes That Are Epistemically Homogeneous Relative to More Than One Variable
So far, I have argued that if we do not know what the right partitions in a class are to show causally relevant differences, then we may call this class epistemically homogeneous. We might say that things are equal until proven different. 11 Without such a ruling, any policy that requires some assumption about the probabilities of all cases under its scope will be incomplete. As already indicated when I discussed Hausman’s average effect condition, assuming a population is homogeneous will have more serious consequences if there turn out to be large causally relevant differences in subpopulations; if, for instance, territorial ethnic civil wars respond quite differently to international interventions than non-territorial ethnic civil wars do, then the average effect of international interventions for the class of all ethnic civil wars will be a misleading source of information for anyone deciding whether or not to intervene in an individual conflict.
In the case study above, I found that there are cases when a researcher does not simply investigate the homogeneity of a population relative to one particular variable. There are instances when researchers are not simply interested in the relationship between one property and the class under consideration (as when we try to investigate whether the polity score of a country is a [contributing] cause of civil war). Instead, researchers may wish to link a whole list of properties to the class under consideration (as when we are interested in a complete causal picture of civil war). Sambanis, in his case study, looks at more than just ethnic fragmentation; he concludes that there are “statistically significant differences between the means of core variables (e.g. political rights, ethnic heterogeneity, and war duration) sorted by war type” (Sambanis 2001, 272). This leads him to argue for researching the differences between ethnic and non-ethnic wars in more detail. We may call the sort of homogeneity we are looking for when it comes to a whole list of properties “total causal homogeneity,” as opposed to the “relative causal homogeneity” researchers are looking for when they only consider one (potentially causal) property.
As anticipated in the introduction, a critic might respond that there are no “social kinds,” that is, there are reasons to believe that all social science cases have causally relevant differences (cf. Little 1991). This means that total ontological causal homogeneity is far too strong a requirement for social science. Nevertheless, I would argue, one can fruitfully employ a notion of total epistemic causal homogeneity. In other words, though after comparative research it will always turn out that all homogeneous classes have size one, that is, contain only one element, researchers may work with what they believe to be a completely causally homogeneous class until they find a relevant partition. 12
Consider the following formalization of this search for total causal homogeneity. Looking for total ontological causal homogeneity would involve the following. Call
Total epistemic causal homogeneity, then, is the weaker notion of generalization that we adopt when we do not have sufficient information to formulate such partitions
Table 1 summarizes the distinction between total and relative homogeneity on the one hand, and epistemic and ontological homogeneity on the other.
Distinguishing total and relative, and epistemic and ontological homogeneity.
6.2. Pragmatic Epistemic Homogeneity
The criticism that there always exist causally relevant differences between social phenomena now brings me to another way in which we can taxonomize epistemic causal homogeneity. Although one may agree that there are always going to be causally relevant differences between cases, this has so far not stopped social scientists from attempting to generalize, whether relative to one particular property (in the case of epistemic causal homogeneity relative to a property) or relative to all properties (in the case of complete epistemic causal homogeneity). In this section, I will highlight what reasons researchers can have to generalize despite these relevant differences, and distinguish between “true epistemic causal homogeneity” (when researchers truly cannot think of a relevant partitioning of a class) and “pragmatic epistemic causal homogeneity.” Researchers are working with the latter when they know particular partitionings of the class exist, but treat the class as homogeneous anyway for pragmatic reasons (including, for instance, not having the time, energy, or resources to adjust policy to any subclasses). 13 In the latter case, the causal claims one makes about a class are not tracking the true causal structure of the world, or even what we believe may be the true causal structure given epistemic limitations, but such claims are nonetheless useful simplifications.
To highlight these two constraints on homogeneity considerations, consider the following example. If we wish to know whether there is a type-level causal relation between students’ ethnicity and their educational attainment, we must measure the relation between the two (properly systematized) variables for a particular population: for example, we may calculate the correlation between ethnicity and educational attainment for all Key Stage 4 pupils in England.
Now, crucially, assuming that each ethnic group in English Key Stage 4 pupils is causally homogeneous with respect to educational attainment will rely on two considerations. First, as I have argued in the first part of this paper, such an assumption will depend on whether researchers have any evidence to the contrary (whether there exists a place selection that would show there exist causally heterogeneous subgroups within at least one ethnic group of English Key Stage 4 pupils with relation to the property of educational attainment; think of, for example, gender).
Second, which group researchers treat as causally homogeneous will depend on the purpose for which they seek the causally homogeneous group. To name but a few aim-related considerations, which group researchers treat as causally homogeneous will depend on whether they can afford (the time, energy, finances) or risk to generalize over causally heterogeneous subgroups. Do they have the time, energy, or finances, say, to research educational attainment for each ethnic subgroup? What would they risk if they did not, and what might the consequences be? Moreover, even if they have good reasons to believe there exist causally relevant differences between some subgroups, do the researchers have the time, energy, or finances to treat each subgroup differently?
The distinction between true and pragmatic causal homogeneity is also related to the following distinction. Generalization always amounts to describing properties of some population. On some occasions, all one cares about are facts about the population at large; in that case, averaging out over subpopulations we know to be heterogeneous is acceptable, that is, we can work under pragmatic epistemic homogeneity. On other occasions, one cares about painting as accurate a picture as possible for each individual in the population, and not just for the population at large. In that case, averaging out is unacceptable; we will not work under pragmatic epistemic homogeneity.
The distinction between true and pragmatic causal homogeneity completes my taxonomy of the different kinds of homogeneity. We can thus expand Table 1 to include this distinction, into Table 2.
Distinguishing pragmatic and true homogeneity.
The distinction between true and pragmatic epistemic homogeneity raises a few questions. First, we may ask whether true epistemic causal homogeneity can still be called epistemic (as opposed to ontological). I would argue that it can; it is a description of the state of our knowledge about the causal structure of a particular class of units, and not a description of the state of the world. As Table 2 illustrates, a class can be ontologically heterogeneous yet epistemically homogeneous. Second, we may ask whether there is any reason to use the term true epistemic causal homogeneity given that, as stated above, among others Daniel Little (1991) has convincingly argued there are no social kinds: there are always causally relevant differences between units researchers aim to generalize over. I would argue there are reasons to use the term: it describes the state of our current knowledge. To make a distinction between true epistemic causal homogeneity and pragmatic epistemic causal homogeneity allows us to highlight the difference between not knowing a way to subdivide a group into idiosyncratic individuals, and knowing such a subdivision exists, yet not acting on it.
7. Conclusion
In the above, I have argued that ontological causal homogeneity is too strong a requirement for comparability, and thus for building general theories. It requires one to have knowledge about all the events or phenomena in one’s class of interest, which does not reflect scientific practice. I defended one particular way of weakening the notion of homogeneity so as to better understand the process of generalizations in social science, that is, making a distinction between ontological and epistemic homogeneity. I showed how the notion of epistemic homogeneity can bring us closer to accounting for what constitutes an adequate justification for the scope conditions of a general theory, and illustrated this with the development of Nicholas Sambanis’ theory on ethnic civil wars. I then refined the notion of epistemic homogeneity by showing one can delineate a class of events or phenomena according to one property (is this class of pupils homogeneous with reference to their educational attainment?) or according to all properties (is there class of civil wars homogeneous with reference to all causally relevant variables?), that is, I distinguished between relative epistemic homogeneity and total epistemic homogeneity. Moreover, I discussed and illustrated a further distinction between pragmatic epistemic homogeneity, which takes into account the practical constraints on how much time and effort one can put into distinguishing a group into subgroups, and true epistemic causal homogeneity, which merely refers to the case when a researcher does not know a relevant partitioning.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
