Abstract
Small and Calarco have done the field a great service; we must go further and arm readers with better understandings of when authors have in fact fulfilled Small and Calarco’s strictures.
“The men of corruption are witty and slanderous; they practise types of murder that no longer need daggers or assault — they know that that which is said well is believed.” – Friedrich Nietzsche
Introduction
This really is a wonderful book, and it treats of a crucial issue that has really been overlooked—not how qualitative researchers can be better, but how others can know this. There are few books that I think absolutely everyone in the social sciences should read, and this is one. But the (deliberately short) book does not complete the task. An expanded, second, edition will be necessary that further begins from the perspective of the reader, and walks through, in greater detail, precisely how one can assess whether the research in question satisfies the strictures Small and Calarco lay out. In some cases, they pay explicit attention to these issues, but in others, they move rather quickly, assuming that if the researcher is, for example, empathetic, the reader will be able to tell.
I want to walk through the main points of the book and assess where there needs to be further attention to the process of assessment. The question is how a reader can detect when a researcher claims to be empathic, self-aware, and so on but is not. Here I am not interested in cases of deliberate falsification (with one marginal exception)—just as we do not teach practical statistics by focusing on cases of flagrant fakery, so there is no payoff to considering similar malfeasance in qualitative research. The question is how to tell the difference between sloppy and/or false applications of the criteria that Small and Calarco (henceforward SC) provide. I assume that others have described the overall approach and arguments of the book, and I can jump right in to considering their key criteria in detail, focusing exclusively on issues related to the Nietzsche dictum which I used as an epigraph. 1
Cognitive Empathy
I think that the most problematic criterion they give is the first, cognitive empathy, “the degree to which the researcher understands how those interviewed or observed view the world and themselves—from their perspective” (23). But how do we know when our qualitative researcher (QR) is giving us the perspective of the research subjects? Without, so far as I could tell, explicitly saying this (for they do not concentrate on what interpretations their example researcher makes), in this chapter SC emphasize getting more data. The implicit rule is that readers should beware of leaps of reasoning, in which an interpretation is given which is compatible with the data—but other interpretations might also be compatible. (This will intersect with their treatment of palpability discussed below.) And this seems to me true and significant. But a great problem lies therein.
Let us examine their example of a seemingly reassuring description from the hypothetical study of Maria (page 39). Their extract runs as follows: “In a predominantly black, high-poverty neighborhood in Philadelphia, a woman who appeared to be in her late teens stepped off a bus, wearing a blue-and-white school uniform, a school bag slung over her shoulder. She had dark brown skin and a large Afro. Heading north, she tiptoed around barbeque leftovers on the sidewalk—foil, charcoal, a few paper plates. It was a crisp September afternoon; a plane roared overhead. As she crossed the street, she saw two men in their early twenties spraying graffiti on the brick wall of a corner townhouse. Both wore fitted black pants, high-top sneakers, and oversized sweaters. By the end of the block, near another bus stop, she stared at a dark yellow stain on the pavement—the block smelled of urine. ”
Consider, in contrast, this possible extract: “In a transitional neighborhood near the University of Pennsylvania, a woman who appeared to be in her early twenties stepped off a bus, wearing a blue-and-white uniform of some type, a knapsack slung over her shoulder. She had dark brown skin, large tortoise-shell glasses, and somewhat bulky shoes. Heading north, she walked a bit uncertainly, winding her way through tree roots poking through the pavement or litter scattered about. It was a warm September afternoon; laughter could be heard coming from around the corner. As she crossed the street, she saw a child, perhaps 6 years old, staring at her from an upper window. The child seemed like she was pointing at the woman, perhaps saying something to someone else in the room. The woman looked down uncomfortably and continued. By the end of the block, near another bus stop, she stared at a dark spot on the sidewalk, trying to determine if it was slippery.”
We can imagine that these two extracts describe the exact same events, and illustrate another point emphasized by SC, that different observers may see different things. The key is that the first account, saturated by the observer's conviction that this-is-a-case-of (TIACO) a high schooler disgusted with the environment, not only focuses on certain things, but makes certain potentially implausible linkages (knowing what she is looking at and staring at); the second account, saturated by the observer's conviction that TIACO someone with mobility problems, assumes very different things. And in both cases, I would be somewhat concerned that the results were unreliable.
The take-away—and this will be the main theme of my engagement with SC—is that what is more important than displays of cognitive empathy is the cognitive plausibility of the account. The central problem is that past a certain point, the sort of rich detail and mutually supporting bits of evidence that give a qualitative report a literary feel and make it more convincing (the Nietzsche principle) actually should be signs for concern. Reading the SC extract, I wonder, how did the QR take down all these observations? Was he following the focal subject and talking into a recorder? Was he situated somewhere with a good angle of view and scribbling notes? Why would he write down that there was a plane overhead? Why that there were foil, charcoal and paper plates on the ground? Did the QR already know that TIACO environmental disgust? Was that the QR's main interest? Is, perhaps, the only robust finding that the QR is disgusted?
The red flag that such an account should raise is that it may be a reconstruction from memory, and what we know about such reconstructions is that (1) they are what we can call “downwardly permeable”—our higher order TIACO sense (e.g., “this was a car crash”) influences what particularities and details we notice, and leads us to resolve ambiguities in a way consisted with our TIACO decision (e.g., “the stuff on the ground was broken glass”); (2) there is no relation between the strength of our conviction that certain elements were actually present and whether they were in fact there. Many people believe that they have photographic memories, and that they can bring up the scene again in their minds (for example, they are sure that they can “see” the Lincoln Memorial in their mind's eye), but if asked to use the scene to answer a straightforward question that one could with a picture (for example, to count the columns), they cannot. They do not have an image in memory, they have a memory of an image—which is one often put together from beliefs and associations that lead the memory to approach a more conventional schema held by the person in question. Thus what may seem to signal the reliability of the account—its richness—may in fact signal the opposite, that it is (unknown to the author) a partial fabrication. One principle, then, is that readers should always know exactly how observational notes were taken, and any reconstructions should be explicitly labeled as such. The reader should be judging whether it is plausible that the QR was able to ascertain all the information he claims he did, and whether his inferences are justified.
But the reader must also judge whether the tasks that QRs set respondents are plausible. The first of only two places where I am not convinced by SC is where they suggest that cognitive empathy can be firmed up by asking hypotheticals. Of course, there is no question-type that can never yield important information, but think about the cognitive task that we are asking of a respondent with such a hypothetical: to construct an imaginary world, perhaps one often fantasized about (and hence related to a host of issues we probably do not understand), but more often one never seriously entertained before, transport one's self-understanding to this world, and then determine the most likely course of actions (and, in many cases, though not SC's [33] version, what others would have done)—within the 3 s of normal latency in a conversation.
In sum, a good rule of thumb to determine when we should trust text that seems to display cognitive empathy is to determine whether the cognitive processes that produced the data, both those of the QR and those of the subject, can plausibly be considered to belong to normal human beings. If not, richness of detail is a warning sign.
Heterogeneity
Before contemporary astronomy, people all over the world knew the astral configurations well, as these were necessary for agricultural and navigational purposes, as well as being aesthetically and experientially compelling in their own right. But why don’t all societies see the same constellations? The answer is that you get to pick and choose which stars you put in which constellation. And this brings us to SC's second point, the importance of the QR revealing heterogeneity. Readers can indeed often assess this heterogeneity, and so I would like instead to focus on the relevance of this principle for the other principles. The principle of heterogeneity implies the possibility of “cherry picking”—selecting details that are tendentious, that lean towards one interpretation, and the principle of downward permeability implies the regularity of such tendentious interpretations, not because QRs are deliberately cherry picking, but because that's the way that human minds work. There are, I think, two implications that fit the arguments made by SC, but are not explicitly emphasized by them. 2
The first is to be suspicious of extraneous detail. When interviewers do not merely describe the interviewee, but give a thumbnail description of the living room they are in (“the room had an ostentatious feel, with expensive but somewhat mismatched furniture, and impersonal decorations”) we should guess that they have formed a TIACO of the respondent, and that this TIACO has led them to notice (if they have actually taken notes on the surroundings at the time) or embellish (if they have not) certain details consonant with this TIACO. This is grounds for concern: perhaps the QR has been encouraged to make the piece more literary, but perhaps the actual data collected (in this case, interviews) are not sufficient to support the claims being made.
It is not that QRs must put on blinders and ignore contextual information, but that if they do include such details, they must conduct what Mitchell Duneier (2011) has called “inconvenience sampling”—figuring out what their TIACO is, what sorts of details would be inconvenient for such an interpretation and looking for them. If the arrangement of the room seems to contain valuable data, it should be asked about (SC's principle of Follow Up). “I notice you have several Thomas Kinkade paintings up…is he someone you like?” “Well, it depends what you mean; he's my uncle, and he painted these for us. I can’t say I really like them, but it was so sweet of him to give them to us…we do hide them on occasion…..”
It is for this reason that I would strongly insist that the fabrication of detail should be taken as totally unacceptable. Strangely enough, for some time, this was considered a fine thing for QRs to do. The notion was that to prevent identification of respondents, one should not describe them too accurately, but to make for a plausible picture, one should “scramble” some of the details. The principle of heterogeneity should show why this is so dangerous—the scrambling is always done to increase the weight of the evidence for the QR's TIACO. If the respondent is an immigrant Lebanese Muslim, and the QR decides that what is most important about him is that he is an immigrant, the counterpart might be a Laotian immigrant. If the TIACO is that he is Muslim, it might be a Singaporean Muslim. And so on. Details that could support an alternate TIACO are suppressed.
This brings us to the second of two places where I am hesitant to endorse SC's arguments, and this pertains to heterogeneity of motives. SC (56) write that “many actions are driven by multiple and even contradictory motives,” and a good QR will reveal this. I am not quite sure whether “contradictory” means ambivalence (I do want to go to the gym and swim this morning but at the same time I don’t) or whether the motivations push for the same action but evoke contradictory ontologies (I want to tell Brad to shut up because only then will he respect me and we will be real friends, and I want to tell Brad to shut up because I never want to see him again). But in either case, although the statement seems unobjectionable, I am not actually aware of any evidence that supports it. Motivation is a tricky beast to observe; like the Loch Ness, we have many unconfirmed sightings. For those who think we should be able to study motivation in causal terms, laboratory experiments usually can only evoke a single one at a time, and the application of multiple treatments would not in itself determine whether they combined or acted separately. The rest of us rely on talk about motivations, also known as justifications, explanations, and accounts. The fact that people come up with more than one of these (“It was like that when I got there, it was made poorly anyway, and I wasn’t even there”) is hardly evidence of the heterogeneity of motivation. I think that the Everett Hughes-Howard Becker-C.W. Mills tradition that tends to eschew asking about motivation produces more palpable data. And this is what SC tell us we should be aiming for.
Palpability
Thus we see that it is not simply palpability, but plausible and systematic palpability, that we must look for. Implausible palpability arises when QRs seem to remember far more than we would expect of even a trained observer, and when respondents remember certain convenient information with more consequential detail than consonant with the results from psychological studies of memory. Unsystematic palpability arises when the key bits of concrete detail seem ad-hoc—we do not know that the researcher was looking for them, or how they were observed and recorded.
That does not mean that QRs should be held to a rigid standard in which it is impermissible to introduce any ideas that were not deliberately used to structure the research design. As SC emphasize, one of the advantages of qualitative research is the iterative process of learning and refining one's conceptual orientation. It is often the case that early interviews or observations are of some use, but not as much use as are later ones. That does not mean that they should be excluded. But good researchers will make a distinction between (a) a pilot project, (b) this initial orientation phase, a bit akin to a “burn in” in Markov Chain analysis, and (c) the main research, at least when it becomes relevant for determining comparability. (For example, the QR may give text from an interview, and then note that at this time, the main theme had not emerged, and so unfortunately, this respondent was not asked about this or that.)
SC do a masterful job at communicating the importance of palpability, emphasizing that abstractions aren’t data (81), but they are less worried than I about the use of detail to paint a picture. Irrelevant detail that helps make a scene vivid for a reader may make for a better book, or they might lead us to confuse literary and scientific merit. 3 But, most important, they may not simply be irrelevant (as SC discuss on 84); they may illegitimately smuggle in non-systematic support for a TIACO, influencing the reader to accept an interpretation for which the more systematic data are lacking. I think that we have to wrestle with the question of whether the use of detail can become a “questionable writing practice”—a way of giving a claim more weight than it deserves”—and that palpability can have a side that should raise, not allay, concerns.
In the early aughts, HBO produced a television series The Wire, which inspired much praise from people I knew. All assured me that though they knew I hated television drama, this one was different, because it was so incredibly realistic and accurate. What bemused me was that those telling me had no first hand understanding of policing in late twentieth century Baltimore, nor had most done any research. So how could they tell that it was realistic? I think what they meant was that it seemed, it had the appearance of, reality—it had verisimilitude, or what early modern aestheticians would have called probability. The problem with evaluating qualitative research may be similar—that realistic seeming fictions are declared superior to unlikely truths. For this reason, it is essential that readers know how to focus on particularities that can reveal the presence of questionable writing practices. Most important is undisciplined selectivity—assembling ad-hoc collections of data here and there to paint a picture, as opposed to describing where the weight of the evidence lies. We distinguish ourselves from investigative journalism by our commitment to systematicity, and we must beware of allowing the journalistic aspects of our writing, fine in themselves, to be used to convince readers where the data do not.
Follow-up
SC's wonderful discussion of the follow-up principle is different, as most readers should be able to tell if a QR has indeed followed up…unless, of course, the aping of hypothetico-deductive logic has led them to deliberately obscure the fact that they have followed up. This does sometimes happen, as some reviewers seem to have a discomfort with any author who show signs of having learned anything in the field. This is a shame, because following-up often is far more truly hypothetico-deductive than what passes for such in most quantitative research. One of the masters of following-up was Becker—in Boys in White he (Becker et al. 1977 [1961]) describes how his team would develop a potential interpretation, a hypothesis, and then seek to test it by making a deduction (“that implies that in situations of type X we should see action of type Y…”) and going off to test it via targeted observations.
It might, indeed, be possible to amplify SC's discussion here, and propose that we should judge a qualitative report superior when it allows us to reproduce the iterative nature of the processes that the QR used to produce the data being analyzed, and inferior when it does not—at the scale of the project as a whole, at that of an interview, and even at that of an extract of dialogue. SC (66) have a wonderful discussion of how the same question, embedded in a different conversational context, can provoke opposite answers. But what they don’t emphasize is that, for this very reason, QRs should not give only the text of respondents’ statements, but their own words as well. QRs give the respondent a task, and just like we cannot understand whether “47” is a right or wrong answer if we do not know the problem, so we do not know what “I made it because I have a network that supports me” means if this is taken out of context.
Self-awareness
Self-awareness—“the extent to which the researcher understands the impact of who they are on those interviewed and observed” (119)—like empathy for others, is almost certainly an important contributor to good research, but it is entirely opaque to the reader. Do we mean the appearance of self-awareness? Perhaps we do. SC show that what they are looking for are plausible hypotheses as to how the nature of the QR might have affected access and likely conclusions, which is indeed an excellent practice, and they emphasize that this is not a simple issue of demographic matching (129). I am not sure, however, that their focus on “disclosure” (131) quite makes sense. This seems to contradict an expansive reading of SC's heterogeneity principle by suggesting that there is, for any study, only one real truth that all QR's should, if all goes well, retrieve, and that this truth is hidden from the casual observer but may be revealed to the trusted. It seems to me more likely that there are different interpretive glosses that actors give at different times and to different types of others; that sometimes the trusted insiders are at a disadvantage (because their interpretations may be more consequential for the fates of others, these others work harder to push the observer's interpretation in a certain direction); and that our great difficulties are in noticing, recording, and carrying out pattern recognition on the many small facts that do appear, and not in getting behind the appearance to the truth via disclosure.
If so, then it may be the case that one of the most important ways that a QR's presence affects the findings by come because he, perhaps unwittingly, encourages certain registers of discussion and not others, just as things about him might end up driving some potential informants away and not others (something SC do discuss [125, 127f]). The problem is that few QRs will write, “my obvious obsession with issues of cultural sophistication increasingly led my respondents to volunteer information that was designed to demonstrate to me that they were neither snobs nor boors, at least, that is true except for those who took an instant dislike to me, and who tried to deliberately push my buttons.”
Part of the problem, then, is that the things we most want QRs to be self-aware about are precisely those parts of their study and interaction that they can’t be self-aware of, because if they were, they wouldn’t have done it in the first place! It seems a paradox—how do we judge whether someone else is truly self-aware? I actually think there are things we can say to help readers follow SC's ideas here. First, notice what SC do not say that QRs should do: spend a great deal of time revealing things about themselves to the reader where these do not seem relevant to the issues of access and observations, boasting about the nobility of their goals, and grandstanding what we in academia call “politics,” or making unfounded, narcissistic claims (“as a member of X I was fundamentally attuned to the concerns of my subjects and understood them in a deeper way than others; my obvious commitment to the betterment of the entire planet also meant that I got much better results than anyone else, and when people occasionally threw rocks at me, I realized that this was because of their false consciousness….”). Such self-promotional, and self-indulgent, digressions have, in the past, been confused with self-awareness, which would suggest that somewhere, the very least self-aware people in the universe had gained control over what constituted self-awareness. Instead, it is more likely that the presence of discipline—a focus on what the reader needs to know for scientific purposes—in self-description, as opposed to narcissism, is one sign of self-awareness.
But even more useful from the point of view of readers is when the author substitutes the awareness that subjects had of her for her own—letting the reader see the most useful and valid data on how the QR was seen or did affect others. A wonderful example is in Duneier's (1999) Sidewalk, where a tape recorder left on captured a discussion that his informants had about him behind his back. Totally non-useful are “certificates of authenticity,” such as “you’re really one of us! You’re so different from all those phony academics we meet! You’re real.” Even if this was a common expression from the subjects, and even if it is how they felt, a truly self-aware QR would never print this—it's just too embarrassing. So if we see these sorts of certificates, we have good reason to suspect that we are reading something from a non-self-aware QR.
Conclusion
In addition to applauding the careful thought involved, the innovative nature of the work as a whole, and the transformative effect this book may have on the field (the emphasis on exposure over sample size as the key issue in evaluating proposals may really help turn things around), make this an unusually important contribution, certainly worthy of detailed attention from methodologists. Only after Small and Calarco had written this book does it become apparent just how badly a book like this was needed. I hope that they, and the field, can go farther in working out the principles that will allow readers to tell good qualitative research from bad. I think we can supplement their principles by focusing on the cognitive plausibility of the processes that would have been necessary to generate the data.
Footnotes
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
