Abstract
Appellate courts sometimes provide relief in cases where prosecutors engage in certain actions, either free from scrutiny during investigation (backstage) or under judicial oversight during litigation (front-stage), that go beyond their authority and the law. Yet little is known about how the nature and types of prosecutorial misconduct recognized by appellate courts systematically affect their decisions to provide relief. Using data from the Center for Prosecutor Integrity, we analyze 150 appellate court cases between 2010 and 2015 in which prosecutorial misconduct is substantiated by the courts. We find that higher courts are more likely to correct for cases involving multiple types of misconduct and for cases in which the misconduct occurs “backstage,” outside of judicial oversight, rather than during litigation.
In an age of skepticism over the ability of institutions to regulate the behavior of their members, increasing attention has been directed at mechanisms of professional oversight. For example, in the context of criminal justice, civilian complaint review boards have emerged in response to a growing disquiet over the ability of law enforcement to oversee the misconduct of individual police officers and systemic problems with police culture (Walker & Bumphus, 1992). The consequential role that an external body plays in the regulation, supervision, and oversight, of this state actor, however, is not parallel to how other agents of the state are policed (Gershman, 1985). As lawyers, prosecutors are occasionally subject to internal controls in the form of state bar associations or other disciplinary boards, but the major source of oversight for prosecutors, specifically, is the court community itself, in the form of self-regulation and judicial review. Said differently, for American prosecutors—who are amongst the most powerful criminal justice actor in the legal system given their enormous discretion in charging, plea bargaining, trial preparation and strategy, and sentencing recommendations (Davis, 2007; Gershman, 1985; Sklansky, 2018)—the central check on behavior comes in the form of trial judges’ rulings in cases that advance to trial, and appellate judges’ responses to claims of prosecutorial misconduct in the post-conviction realm. 1
In effect, court communities reflect an insular institutional check on the power of the prosecutor. These communities not only dictate punishment regimes and policies according to local conditions and organizational influences (Eisenstein et al., 1988; Johnson, 2018; Ulmer, 2005), but they also effectively police themselves. The degree to which prosecutors shape case processing sequences, in the absence of genuine oversight especially in the earliest stages of criminal case processing, provides opportunities for negligence or outright misconduct, even if these actions are ostensibly “noble” and based on an “ends justify the means” pursuit of justice (Gershman, 1985; Schoenfeld, 2005). While legal scholarship has focused on identifying various forms of prosecutorial misconduct and relating each to the likelihood that it contributes to miscarriages of justice (Medwed, 2012), little empirical research examines how higher courts handle such error.
In the United States, most jurisdictions have a judicial system characterized by a three-tier court hierarchy. Defendants who are convicted at the trial court level may appeal that outcome “as of right” directly to an intermediate appellate court, in which appellate judges evaluate any defense claims of error that occurred below. After the intermediate appellate court makes its decision, the losing party, whether the defendant or the state, then may seek “leave” to appeal that decision to the state’s court of last resort in which (often) a panel of appellate judges decides which appeals warrant full review (Halberstam, 2016). 2 If the state’s court of last resort grants review, it typically either affirms or reverses the intermediate court’s decision. Although appellate courts generally afford deferential standards of review in weighing the facts established by trial courts, they may decide, based on their own evaluation of the law, whether an erroneous legal ruling by the trial court merits reversal when misconduct was present (Halberstam, 2016).
Prosecutorial misconduct usually takes the form of due process or constitutional error, which may or may not affect whether the conviction will stand. Convictions may be upheld, even when a trial error took place, if the misstep is deemed “harmless” by appellate courts (Medwed, 2012). For constitutional errors, the conviction will stand only if it is deemed “harmless beyond a reasonable doubt” (Chapman v. California, 1967). For non-constitutional errors, jurisdictions have fashioned a range of different harmless error tests. And, when multiple errors occur, courts may consider them together in evaluating whether reversal is warranted, known as “cumulative harmless error” (Abrams & Garrett, 2017). Even if multiple errors exist in a criminal case, as the harmless error rule has evolved over time, appeals courts tend to focus on the quantity of evidence presented at trial and ultimately ignore error when the trial outcome would have been the same had the error not occurred. In essence, appellate courts may overlook prosecutorial misconduct if the rest of the evidence supports conviction (Lawless, 2008; Schoenfeld, 2005).
What, then, leads appellate courts to deem the conduct of the prosecutor so improper that they must correct for the lower court’s failure to sufficiently remedy it? Below, we first consider the literature on prosecutorial misconduct and the harmless error doctrine and posit how courts as communities may affect appellate relief of prosecutorial error. Next, using data from the Center for Prosecutor Integrity (CPI), we engage in an empirical study to explore how the nature and type of prosecutorial misconduct acknowledged in criminal trials affect appellate court determinations of relief in 150 state (n = 141) and federal (n = 9) appellate decisions that occurred between 2010 and 2015. 3 We end with a discussion of the theoretical and policy implications of our findings.
Prosecutorial Misconduct
Berger v. United States (1935) held that: “The United States Attorney . . . may prosecute with earnestness and vigor – indeed, he should do so. But, while he may strike hard blows, he is not at liberty to strike foul ones. It is as much his duty to refrain from improper methods calculated to produce a wrongful conviction as it is to use every legitimate means to bring about a just one.” Improper conduct involves various prosecutorial actions that occur when preparing for trial (the investigation stage) or when the trial is ongoing (the litigation stage). Goffman’s (1959) dramaturgical perspective offers an apt analogy to the phases during which prosecutors may commit errors or engage in misconduct while handling a criminal case: front-stage and backstage. “Front-stage” behaviors are effectively “performances” in front of others wherein actors are aware that the audience will perceive their actions in specific ways and have certain expectations for their behavior. In contrast, “backstage” activities can be more uninhibited. For prosecutors, backstage errors can occur with a greater degree of impunity, as they are more difficult to detect than those that surface in the presence of judges and juries (Medwed, 2012).
Backstage errors occur predominantly within the pretrial stage of case processing during investigation, as this is when prosecutors gather evidence and prepare for trial. Courts have consistently ruled the following types of pretrial conduct as out of bounds on the part of the prosecution: failing to divulge exculpatory evidence to the defense (also known as a Brady violation; Brady v. Maryland, 1963); charging errors, including overcharging; trial preparation errors, like falsification, tampering, or subornation of perjury; and plea bargaining errors, such as penalizing the defendant’s assertion of a right or privilege by charging more severely than warranted (Acker & Redlich, 2011; Hetherington, 2002; Lawless, 2008).
Litigation represents the front-stage of prosecutorial behavior such that it happens in front of other legal actors during the trial. In this context, prosecutors are prohibited from a range of activities, among them: impugning the defendant, the defense counsel, or witnesses for the defense; shifting the burden of proof or misstating the law or facts; vouching for government witnesses; presenting false or inadmissible testimony; mischaracterizing evidence; or implying that jurors are the conscience of the community and/or making improper or otherwise inflammatory arguments (Acker & Redlich, 2011; Hetherington, 2002; Lawless, 2008).
Existing studies on prosecutorial misconduct suggest that, despite appellate courts acknowledging that prosecutorial errors were overlooked or not properly corrected at trial, only a limited number of cases result in a determination of non-harmless error. For example, a study of cases of prosecutorial error in the state of California from 1997 to 2009 shows that only 159 (22.5%) of the 707 cases found to have involved prosecutorial misconduct were overturned because courts deemed the error not harmless (Ridolfi & Possley, 2010). Similarly, West’s (2010) review of court findings of prosecutorial misconduct claims among the first 255 DNA exoneration cases indicates that, of the 65 cases alleging prosecutorial misconduct through appeals or civil suits, error was acknowledged by the court in 31 (48%) cases with only 12 (18%) cases resulting in a determination of non-harmless error. Finally, the Center for Public Integrity finds that of 11,452 appeals involving prosecutorial misconduct between 1970 and 2002, only 17.6% (2,012 appeals) were reversed or remanded, suggesting the presence of non-harmless errors (West, 2010). Notably, no statistical analyses of how prosecutorial misconduct affected case outcomes were conducted in these studies. 4 Previous research demonstrates, however, that even when appellate courts concede that trial courts did not properly recognize or address prosecutorial misconduct, that conduct is typically viewed as harmless.
Determining Harm
The Supreme Court has endorsed several different harmless error tests. Under one of them, an error is found not harmless if it had some impact on the jury; this is known as the contribution-to-conviction test or Chapman test (Chapman v. California, 1967). Chapman specified a procedure whereby a defendant would first have to demonstrate the existence of a constitutional error and, upon doing so, the state would then have to prove that the error was harmless beyond a reasonable doubt. This is a high standard, meaning that there is no “reasonable possibility that the evidence complained of might have contributed to the conviction” (Carter, 2001, p. 231). Another test treats error as harmless if the record provides clear and convincing evidence of the defendant’s guilt absent the error; this is known as the overwhelming-evidence test or Harrington test (Harrington v. California, 1969). Finally, the Van Arsdall test fuses the first two basing the outcome on “either the severity of the error or the weight of the untainted evidence, depending upon what the court chooses to emphasize” (Mitchell, 1994, p. 1339; see also Delaware v. Van Arsdall, 1986). Of these tests, the Harrington test “has become standard practice for many appellate panels” (Edwards, 1995, pp. 1186–1187).
Harmless error standards vary across states and it is not mandatory that courts use any one of these tests uniformly. 5 Additionally, courts may make “cumulative harmless error” determinations by examining the aggregate impact of various errors present within a trial. Even if any single error alone in a particular case does not amount to reversible error, the presence of multiple errors could together deny the defendant a fair trial and justify reversal (Taylor v. Kentucky, 1978). Yet, there is no cumulative harmless error “doctrine” per se, and courts tend to apply the concept in a somewhat random and inconsistent fashion (Blume & Seeds, 2005).
What is common, however, is that appellate courts evaluate the quantity of evidence of guilt adduced at trial and then measure the alleged error(s) against that inculpatory evidence. Not only does this result in harmless error analyses that are biased in favor of the prosecution, as it enables courts to “find almost all error harmless” (Carter, 2001, p. 243), but it also may produce a range of problematic issues by allowing appellate courts to engage in “unguided speculation about guilt” (Mitchell, 1994, p. 1340). Requiring appellate courts to assess guilt complicates their function as “reviewer[s] of questions of law” (Carter, 2001, p. 244), undermines the function of juries, and essentially places appellate judges in the role of the “thirteenth” juror by having them focus more on guilt considerations than on the error itself (Little v. State of Mississippi, 2017; Mitchell, 1994). Also, in determining guilt absent the tainted evidence or argument, appellate courts typically limit their examination to the amount of evidence the prosecution has marshaled against the defendant, rather than weighing direct or rebuttal evidence of the defense against the prosecution’s case (Carter, 2001). This may distort the value of the untainted evidence. Finally, in evaluating the impact on the jury, appellate judges are in the untenable position of assuming how jurors, who were privy to all courtroom arguments, theatrics, and witness testimony, were affected by the error. Appellate courts thus typically defer to the trial court’s evaluation of the effect of prosecutorial error on the case outcome (Hetherington, 2002).
The mere positioning of appellate justices in the adversarial system thus highlights aspects of their decision-making that are reflective of a larger court community. This community would be defined by the adversarial court structure that envelopes the appeals process and establishes the going rate for how higher courts identify and respond to the misconduct or mistakes of their criminal justice brethren (Eisenstein et al., 1988; Ulmer & Johnson, 2004). Specifically, how appellate courts view and operate within this system can create pervasive adversarial norms in the legal community that facilitate shared expectations and enable greater certainty about the boundaries or limits regarding prosecutorial misconduct that will be deemed tolerable (Hester, 2017). 6
The vertical stratification of courts may also lend itself to granting considerable deference to other, even lower court actors. In federal habeas hearings, for example, the bar for reversing a state court’s ruling on a factual issue is high. Indeed, “as a matter of federalism and comity between state and federal courts, federal courts presume the adequacy of state proceedings and decisions” (Wolf, 2009, p. 237). It is only in cases where the federal court can identify “clear and convincing evidence” of a mistake in the state court’s handling of the issue that relief may be granted and, even then, the state court’s ruling must be deemed not just erroneous but “unreasonable.” Clearly, both the informal and the formal organization of the courts coalesce into a larger legal community reluctant to provide serious oversight of its members.
Research confirms these suspicions. For instance, in their analysis of 963 federal appellate criminal cases that occurred between 1996 and 1998, Landes and Posner (2001) report that flagrant errors increase the probability of the determination of non-harmless error. Intentional or flagrant errors, including when prosecutors withhold exculpatory evidence or imply a defendant’s guilt during closing argument by commenting on their right not to testify, can serve as a signal to the appellate court that the prosecution had an especially weak case. These authors do show, however, that reversal is somewhat dependent on the actor who errs, with judges holding other judges to a higher standard. That is, Landes and Posner (2001) find that prosecutorial error is more likely to be tolerated than error committed by a trial judge. There is an expectation for judges to be correct and impartial, while it is recognized that prosecutors may “go overboard to get a conviction” in the highly adversarial arena of a criminal trial (Landes & Posner, 2001, p. 188).
The expectation that representatives of the state may get carried away in securing a conviction suggests that appellate judges may be more willing to sanction misconduct when prosecutors are not bound by the confines of judicial oversight (i.e., when prosecutors have abundant opportunities and the errors are least likely to be identified). Errors that occur backstage thus may encourage appellate judges to feel compelled, as neutral arbiters of justice, to protect constitutional freedoms and the purity of the adversarial system.
One type of misconduct that occurs overwhelmingly in the absence of judicial oversight strikes at the heart of fairness and highlights the overly broad discretion granted to prosecutors: the withholding of exculpatory evidence prior to trial, or a Brady violation (Brady v. Maryland, 1963). These violations are highly problematic due to their hidden nature and the difficulties associated with exposing them (Landes & Posner, 2001; Sklansky, 2018), especially as they may occur at any point in the criminal trial process due to the ongoing obligation of the prosecution to add to discovery as it becomes available (Deal, 2007). Prosecutors must decide whether evidence collected during the investigation is Brady information (Medwed, 2012). Their adversarial zeal may foster a distorted perception of such evidence, leading to a determination that it is immaterial to the defendant (Ridolfi & Possley, 2010). Because the defense, to some extent, is reliant on this investigative material, Brady violations reduce public trust that the criminal justice system is convicting the guilty and weeding out the innocent.
Due to their egregious nature, Brady violations are “different” from many other trial-level errors made by prosecutors in that they are not subject to an independent harmless error analysis per se (Kim, 2017). Rather, Brady has a built-in analog through the “materiality” prong of its test. A court will only recognize a Brady violation if the evidence withheld by the prosecution is both favorable to the defense and material to guilt or punishment. Materiality exists if there is a “reasonable probability” that the undisclosed evidence would have affected the outcome. A finding of materiality is akin to a finding that an error is “not harmless”—that the misstep was of consequence to the result at trial (Blume & Seeds, 2005; United States v. Bagley, 1985).
In contrast to backstage errors, appellate courts may be less likely to provide corrective action for litigation or front-stage prosecutorial misconduct for the following reasons. When prosecutors engage in trial forms of misconduct during the litigation phase the likelihood that this error will be observed by trial courts and sufficiently remedied increases. At trial, the judiciary can correct for executive overreach by providing a curative instruction to the jury or sustaining an objection on the part of the defense, thereby curbing prosecutorial power and the abuse of adversarial justice (Medwed, 2012). Inferior court judges may also pay close attention to prosecutorial actions that violate the sanctity of due process in the hopes of avoiding potential reversal. Indeed, trial court judges may rule in ways that adhere to precedent and mirror higher court decisions to avoid the stigma that is associated with reversal (Caminker, 1994a, 1994b) and the full realization of the limited power they have under a hierarchical system of review (Caminker, 1994b, p. 78). Reversal could embarrassingly signal to other colleagues, practitioners, and scholars who comprise the professional audience of judges that their legal judgment or abilities are suspect (Caminker, 1994a, p. 827). For instance, a high reversal rate would add to a judge’s backlog and highlight overly hasty decision-making (Posner, 2005). Further, a high reversal rate could limit opportunities for promotion or appointment to other judicial commissions, thus adversely affecting professional recognition and advancement (Caminker, 1994b, p. 77; Posner, 2005). As such, appellate intervention may become less likely when front-stage misconduct occurs because lower court judges not only have a duty, but also a motivation to avoid the psychological and professional costs associated with reversal (Caminker, 1994b, p. 78), thereby encouraging them to correct for prosecutorial misconduct that comes to their attention. In cases where that remedy is a mistrial, appellate courts will only very rarely hear or opine on the prosecutorial misconduct that took place in the original trial.
The profound effect that prosecutorial misconduct could have on trial outcomes has led to recent efforts by nonprofit, legal, and academic institutions to assess prosecutorial behaviors and performance. For instance, Measures for Justice (2021) is an organization that attempts to increase criminal justice transparency, which, in relation to prosecution, concern assessing the number of cases not prosecuted, and the number of full-time and part-time prosecutors. In addition, the Deason Criminal Justice Reform Center at the Dedman School of Law has implemented the Prosecutorial Charging Practices Project to critically examine prosecutorial practices in relation to screening, charging, police interaction, and evidence evaluation (SMU, 2021). And, a recent collaboration of scholars at Florida International University and Loyola University Chicago has developed Prosecutorial Performance Indicators, a technological tool to assist prosecutor offices in being more data-informed about a wide range of practices to ensure that prosecutors are effective, efficient, and fair. Specifically, prosecutor offices are assessed on procedural and ethics violations, dedication to conviction integrity, commitment to law enforcement accountability, charging integrity, and discovery compliance (FIU, 2021; PPI, 2021). These initiatives highlight what prosecutors do in relation to the prevention of misconduct.
Furthermore, after a period of relative silence on the prosecutorial role, prosecutors’ behavior, and their considerable discretion in case processing, an emerging body of scholarship has begun to focus on: cultural shifts toward progressive prosecution; deflection, diversion, and declination decisions; plea processes; contributors to electoral success; participation in post-conviction review; and involvement in sentencing determinations (Bazelon, 2019; Davis, 2007; Johnson, 2018; Levine & Wright, 2017; Lynch et al., 2021; Sklansky, 2018; Webster, 2019). Little is known, however, about how system actors, such as appellate court judges, react and respond to prosecutorial misconduct that undermines the adversary process.
Current Study
While it is our hope that findings from our research can be used to lessen prosecutorial overreach, this study is not focused on ways to prevent misconduct. Instead, we explore how other criminal justice actors view and police misconduct that is recognized to have occurred. Namely, we examine the conditions under which ameliorative actions are taken by appellate courts to address the (mis)handling of prosecutorial misconduct at trial by lower courts. First, we investigate how the number of unique types of prosecutorial misconduct in cases affect appellate decisions to grant relief. One or two types of error in a single case may be easier to excuse than multiple forms. The presence of only one type may signal that the prosecutor inadvertently stepped over the line, which, if recognized, should be easily remedied by the trial judge’s curative instruction, thereby reducing the likelihood that appellate courts would overturn a trial decision (Landes & Posner, 2001). However, an accumulation of various types of errors makes it more difficult to argue, or for appellate courts to rationalize, that the misconduct was unintentional or benign. We thus hypothesize that:
Hypothesis 1: Appeals courts will be more likely to provide corrective action in cases in which prosecutors are found to have engaged in multiple types of misconduct in comparison to one or few types of misconduct.
Second, we contend that the stage of criminal case processing in which prosecutorial misconduct occurs likely influences whether relief is granted. This is in part a consequence of the enormous discretionary power of the prosecutor in the criminal justice system, particularly during the investigation phase of case processing. Judicial oversight over charging decisions is almost nonexistent; prosecutors choose which cases to pursue and what charges to bring, and they can make decisions to order arrests, pursue an indictment against the accused, or decline to prosecute (Davis, 2007; Gershman, 1985; Sklansky, 2018). We argue that this unchecked discretion may encourage appellate judges to be especially invested in attempting to correct for backstage prosecutorial power for the following reasons: (a) Brady violations usually occur in the investigation phase when prosecutors gather evidence, determine its significance, and then choose to withhold information that is, in fact, favorable and material; (b) other types of prosecutorial misconduct that occur during the investigation stage are equally non-transparent; (c) appellate justices are more likely to assume that lower court judges can discover and correct for misconduct that emerges during the litigation phase; and (d) in the potentially rare instances in which backstage misconduct is discovered, it may be especially problematic or egregious. We therefore predict that:
Hypothesis 2: Appeals courts will be more likely to provide corrective action when prosecutorial misconduct occurs during the investigation rather than litigation phase of case processing.
In testing these hypotheses, we examine how one set of actors in the criminal legal system, appellate courts, pass judgment on the decisions and actions of lower court officials, in this case prosecutors and, by extension, the effectiveness of trial judges in controlling prosecutorial misconduct (Joy, 2004).
Data and Methods
We use data from the Center for Prosecutor Integrity (CPI) to examine when appellate courts grant relief for acknowledged prosecutorial misconduct. The CPI collects information on appellate cases wherein prosecutorial misconduct was substantiated by the trial court or higher courts. 7 For a case to be entered into the database, the CPI follows a series of steps which include: (1) identifying potential cases from individual submissions, media reports, and LexisNexis legal database searches using the following terms: Brady violation, ethical, ex parte communication, failure to disclose, false evidence, improper argument, inadmissible evidence, inflammatory statement, lack of candor, mischaracterizing evidence, perjury, professional conduct, professional responsibility, prosecutor improp!; and prosecutor! misconduct; (2) confirming case eligibility by locating a court (trial, appellate, or supreme) or bar disciplinary committee decision; (3) case analysis by a legal analyst; (4) case review by the Registry Director; and (5) record enhancement, which entails searching for additional information from supplemental sources like bar disciplinary records, defense counsel reports, and media accounts.
While the CPI database extends backward in time, we restrict our analysis to the finding years of 2010 to 2015 for a host of reasons. First, this period coincides with a growing concern over prosecutorial misconduct. In 2010, the Department of Justice presented new initiatives directed at the training of federal prosecutors due to the attention over discovery failures in the prosecution of U.S. Senator Ted Stevens (Green, 2010). In addition, in 2013, for the first time ever, a prosecutor was sentenced to jail for wrongfully convicting an innocent man (Godsey, 2013). Second, at this time a growing number of courts expressed open dismay at the “epidemic” of prosecutorial misconduct—in particular, Brady violations—and their role in impeding and remedying it (United States v. Olsen, 2013). Likewise, in 2011, a key Supreme Court decision (Connick v. Thompson, 2011) effectively barred “one of the few remaining avenues for holding prosecutors civilly liable for official misconduct” (Keenan et al., 2011, p. 204); in turn, this shined a light on how ineffective state disciplinary associations are in holding prosecutors accountable for errors and misconduct that have grave consequences for defendants. Finally, the CPI data are presently available only through 2015 due to funding limitations.
We scrutinized each of these appellate court opinions to address coding errors and conducted additional research to supplement missing data. We also incorporated a series of measures, including whether the court explicitly recognized that prosecutorial error led the defendant to be deserving of relief in the text of its decision, the judicial panel size, and the number of types of misconduct in cases. As our focus is on whether appellate courts provide relief based on prosecutorial misconduct, we excluded cases that consisted of trial court decisions 8 and disciplinary hearings, which resulted in a total of 159 cases. We also excluded those cases that were disposed via plea (n = 4) and those in which judicial panel size was unknown (n = 5). Our final sample consists of a total of 150 trial cases between 2010 and 2015 in which appellate courts substantiated at least one instance of prosecutorial misconduct from a total of 13 states (Alaska, Arizona, California, Delaware, Florida, Hawaii, Illinois, Louisiana, New York, Rhode Island, Tennessee, Texas, and Wisconsin), Washington D.C., and the U.S. Virgin Islands. 9 We do not presume that findings from our analysis can generalize to all prosecutorial misconduct that occurs in practice or that is alleged by defendants. Instead, the scope of our research is limited to how appellate courts handle the misconduct that they recognize has occurred at the trial-level. 10
Due to the dichotomous nature of our dependent variable, we estimate a logistic regression model with robust standard errors to account for the clustering of cases by state. Note that there is insufficient variation in level-2 units (e.g., states) to warrant employing multilevel models in this case. Our analytic strategy allows us to test how the stage in which prosecutorial misconduct occurs (front-stage vs. backstage) and the number of unique types of misconduct affect appellate courts’ willingness to correct for prosecutorial error, controlling for a series of additional case characteristics that could bear on the decision.
Dependent Variable
To determine whether courts found that prosecutorial misconduct was deserving of relief, we created a dichotomous measure coded “1” when the court acknowledged that the alleged prosecutorial misconduct was sufficient to undermine the integrity of the case and the defendant was therefore deserving of relief, and “0” otherwise. For example, the former condition was satisfied when the court referred to prosecutorial misconduct as a basis for reversing the conviction, remanding the case for a new trial, or providing a partial reversal of the conviction or sentence. Cases were coded “0” when the court found that the prosecutorial error existed, but that it had no effect on the trial outcome; this includes 12 cases in which relief was granted on other grounds. 11 In all, prosecutorial misconduct was identified as deserving of relief in 32.0% (n = 48) of the 150 cases (Table 1). Nevertheless, our data show that more than two-thirds of appellate court decisions (n = 102; 68.0%) continue to treat prosecutorial error as effectively harmless even when it is acknowledged to have occurred in pursuit of the conviction.
Descriptive Statistics (n = 150).
Refence category.
Independent Variables
Our first independent variable is the number of distinct forms of prosecutorial misconduct recognized by the court in a particular case. To distinguish cases in which prosecutors engage in multiple types of misconduct, we created a variety score representing: (a) Brady violations; (b) charging errors; (c) evidence errors, including falsification, witness tampering, subornation of perjury, improper elicitation of evidence, and failure to enforce perjury; (d) plea bargaining errors; (e) making inflammatory statements and/or engaging in witness harassment; (f) introducing inadmissible or false evidence, such as arguing facts not in evidence or showing lack of candor; (g) impugning the defense or others; (h) misstating the law and/or shifting the burden of proof; (i) mischaracterizing the evidence to the jury; (j) vouching for a prosecution witness; and (k) other types of prosecutorial misconduct (CPI, 2016). This variety score distinguishes different types of misconduct in a case but not the number of instances of misconduct. A prosecutor may, for example, make more than one inflammatory statement but all instances would count as one type. 12 On average, appellate courts recognized 1.61 (SD = 1.02) types of prosecutorial misconduct across cases, ranging from a single type to a maximum of six.
Our second independent variable examines whether prosecutorial misconduct occurs backstage in the absence of judicial oversight, or front-stage in the presence of other court actors. This categorical variable distinguishes between: (1) cases involving any backstage misconduct (n = 23; 15.3%) that occurred during prosecutorial investigation and case preparation, including six cases that involved both investigation and litigation stage misconduct on the part of the prosecution; and (2) cases involving only front-stage misconduct that occurred during litigation (n = 127; 84.7%) (reference category) (Table 1).
Control Variables
We control for crime type (violent or non-violent), size of the appellate panel, type of case (federal or state), and the state in which the case was originally disposed. First, crime type serves as a control because the seriousness of the offense may lead police and prosecutors to pursue violent crime cases more vigorously (Worrall et al., 2006). Controlling for crime type serves an additional purpose; while appellate judges may be somewhat more impervious to concerns over public safety and security than their trial counterparts, they are not expected to be immune from misgivings over granting relief or impunity, particularly in egregious cases (Brace & Boyea, 2008). This variable is coded such that “1” indicates violent crime, including sex crimes, assault, attempted murder, murder, and other acts of violence; and “0” indicates non-violent crime, including burglary, drug crimes, and other acts of non-violence. Nearly three quarters (n = 110; 73.3%) of the cases involve violent crimes rather than non-violent crimes (n = 40; 26.7%, reference category). Second, we control for panel size since, in comparison to smaller judicial panels, larger panels are more likely to reverse lower court decisions (Halberstam, 2016). 13 On average, four appellate judges (SD = 1.34) decided cases with panels ranging in size from 3 to 9.
Finally, court resources and quality may also play a role in affecting appellate relief (Halberstam, 2016), thus implicating the importance of jurisdiction. Most of the appellate decisions in our data were made in state courts (n = 141; 94.0%, reference category) rather than federal courts (n = 9; 6.0%). Two states were also overrepresented in these data: New York (n = 89; 59.3%) and Texas (n = 36; 24.0%). We thus created a categorical variable for state, with other states (n = 25; 16.7%) as the reference category. Distinguishing New York and Texas, respectively, from other states controls for some state-level differences in appellate review of prosecutorial error and for the possibility that a certain state is driving the analysis. 14
Results
The results of our logistic regression predicting whether appellate courts granted relief in cases where prosecutorial misconduct was acknowledged are shown in Table 2 and provide support for our hypotheses. In support of Hypothesis 1, we find that for each additional type of prosecutorial misconduct acknowledged by appellate courts, the odds are 3.9 times higher that courts will find the error deserving of relief (OR = 3.871 to p ≤ .001). Our second hypothesis predicted that the stage in which the misconduct occurred would affect appellate relief. We find that in comparison to front-stage forms of misconduct subject to judicial oversight, the odds of appellate courts granting relief for prosecutorial misconduct that occurs backstage, involving investigative prosecutorial functions, are nearly five times higher (OR = 4.845; p ≤ .01). Last, none of the remaining case attributes predict whether relief was granted, save for the state in which the case is heard. New York and Texas appellate courts respectively had an 83% (OR = 0.169; p ≤ .05) and 86% (OR = 0.144; p ≤ .05) lower odds of deciding to remedy the prosecutorial error they found in cases, compared to other states, controlling for a variety of features of the cases under review.
Logistic Regression predicting Relief on the Basis of Acknowledged Prosecutorial Misconduct (n = 150).
Reference category is litigation stage.
Reference category is non-violent crime.
Reference category is state case.
Reference category is other state.
p ≤ .05. **p ≤ 01. ***p ≤ .001.
Supplemental Analysis
It is not possible to estimate models predicting relief by specific types of prosecutorial misconduct due to limited events per variable; in some cases, the type of misconduct perfectly predicts the outcome. To account for these small cell sizes and to illustrate whether and how each type of prosecutorial misconduct is associated with appellate relief, we conduct a series of Fisher’s Exact tests (two-sided). Doing so illustrates which specific types of misconduct are significantly associated with ameliorative action in, effectively, a series of 2 × 2 contingency tables. Table 3 indicates that the courts provided relief in nearly all the cases in which a Brady violation is recognized to have taken place (n = 11 of 12; 91.7%; p ≤ .001). At the litigation stage, of the various types of misconduct that can occur, only mischaracterizing (n = 12 of 20; 60.0%; p ≤ .01) and vouching (n = 13 of 22; 59.1%; p ≤ .01) were independently associated with appellate relief. Moreover, the direction of effects is generally consistent with expectations for other non-significant types of investigation and litigation misconduct; fewer than half of cases involving each of the litigation types of prosecutorial misconduct resulted in relief, whereas all but two types of investigative stage prosecutorial misconduct were always or almost always associated with relief. Small numbers limit the power necessary to adequately test these relationships, but the direction of these relationships conform to our expectations.
Frequencies of Types of Misconduct and Fisher’s Exact Tests for Each Type With Relief Granted (n = 150).
Cases may include more than one type of prosecutorial misconduct.
**p ≤ 01. ***p ≤ .001.
Discussion
As members of a shared legal community embedded within a highly adversarial system, appeals courts are tasked with one direct aim, correcting for procedurally unjust outcomes, and one indirect aim, providing oversight of other community members. Outside of the literature on court communities, which tends to focus on collaboration and cooperation between community members to reduce uncertainty, criminal justice theorizing has been directed at understanding “variation in the responses of criminal justice to crime as the outcome” (Duffee & Allan, 2007, p. 16). Very little theorizing about criminal justice illuminates how systems of accountability internal to those communities operate in practice.
Focusing on when appellate courts decide prosecutorial misconduct is worthy of relief, we highlight how adversarial norms structure when higher courts police other criminal justice actors. We thus extend the literature on courts as communities by identifying the existence of an appellate court culture that may operate at the state and national level (Eisenstein et al., 1988). Like Hester (2017), who finds that state legal culture can minimize variation in judicial outcomes, we indicate that adversarial norms establish how appellate courts police other quasi-judicial court actors who violate the standards of their position by engaging in misconduct (Gilliéron, 2013). Here, the going rate (Ulmer & Johnson, 2004) for cases in which prosecutorial misconduct is substantiated by higher courts mainly results in tolerance of that error, thereby reflecting an emphasis on the sanctity of jury verdicts and the finality and conservation of trial resources when there is overwhelming evidence of a defendant’s guilt (Gershman, 1985; Medwed, 2012; Schoenfeld, 2005). Despite this tolerance, our findings demonstrate that decisions to intercede are patterned by cumulative error and whether the misconduct occurs under judicial oversight.
While courts may be forgiving of some prosecutorial overreach, they become increasingly unwilling to overlook these errors as the types of misconduct that occur in a single case increase. Said differently, the more ways that prosecutors err, the more flagrant and intolerable the misconduct is perceived to be. And the more flagrant and intolerable the misconduct, the more likely appellate courts are to intercede. Thus, courts utilizing the “cumulative harmless error” doctrine, whether explicitly or implicitly, will be more likely to grant relief than when evaluating specific types of prosecutorial error in isolation. This may be because: (a) appellate judges interpret multiple forms of misconduct as indicative of a weak prosecution case or as simply too problematic for any curative instruction of the trial judge to overcome in the minds of jurors; (b) finding numerous types of misconduct signals to the judiciary that prosecutors are flouting not only the ethical requirements of their role and the power of the judiciary, but also the credibility of the adversary system itself; (c) the presence of multiple types of error causes appellate judges to doubt the accuracy of the verdict and become concerned that an innocent person is languishing behind bars, or at a minimum that a serious violation of due process occurred; and (d) if prosecutors are regularly violating their special responsibilities, appellate courts may believe that the trial judge failed to recognize or properly handle the misconduct, thereby ineffectively serving as a neutral arbiter.
The presence of multiple types of misconduct remains significant even when accounting for the stage in which the prosecutorial misconduct occurs, which suggests that appellate courts evaluate prosecutorial misconduct on the basis of both quantity and quality. Indeed, this study’s second, and perhaps more important finding, is that appeals courts appear to be mindful of the stage of case processing in which prosecutorial misconduct occurred. When defendants are hampered in their defense by a prosecutor who violates the rules and ethics of the role, at a stage in which the judiciary has little or no oversight capacity, appellate court judges are more likely to view those errors as not harmless to justice compared to errors that occur in the open during litigation. This likely reflects both the power of the prosecutor’s role and the expectations other judicial actors have of prosecutorial duties (Davis, 2007; Gershman, 1988; Medwed, 2012).
Appellate courts thus appear to be less inclined to second-guess the decisions or actions of their colleagues in lower courts when transparency and trial judge oversight are greater. Out of professional courtesy or deference, appellate judges may assume that front-stage prosecutorial misconduct has already been adequately reviewed and handled by trial judges, and thus it must not have been worrisome enough to force a mistrial. Little scholarship speaks directly to this possibility, although O’Hear (2010) finds that appellate courts tend to be reluctant to revisit the sentencing decisions of trial judges, even if doing so would decrease sentencing disparities. This is because they defer to the optimal positioning of trial court judges in evaluating the nuances and unique circumstances of the various trials that lower court judges oversee (O’Hear, 2010). This deference could thus extend to litigation-stage prosecutorial misconduct considerations. Moreover, as members of the professional audience of trial judges, appellate justices should be aware of the psychological and professional motivations inferior court judges have to monitor this misconduct to avoid the likelihood of reversal (Caminker, 1994a, 1994b; Posner, 2005).
In contrast, appellate judges may see backstage misconduct as more egregious than front-stage misconduct for several reasons. First, once an abuse of backstage prosecutorial discretion is discovered, appellate judges may be more apt to act upon this misconduct because their trial brethren could not observe, first-hand, the flagrancy of the misconduct nor could they provide sufficient curative instruction to the jury. Second, appellate courts may hold prosecutors more accountable for this type of overreach precisely because it could represent just the tip of the iceberg. Without the diligent efforts of defense attorneys or investigators, it is possible that very few forms of backstage prosecutorial misconduct see the light of day, given the opacity of backstage decision-making and the incentives to cover serious errors in due process. 15 Last, appellate courts may be more likely to find backstage error not harmless due largely to Brady violations, which are believed to be one of the most pervasive forms of prosecutorial misconduct (Gershman, 2012; Medwed, 2012) and were the predominant type of misconduct present during the investigation stage in our data. If prosecutors abuse their investigatory advantages by withholding exculpatory evidence, these actions not only increase the likelihood of wrongful convictions, but they also violate the separation of powers by effectively transforming agents of the executive into members of the judiciary (Davis, 2007).
Appellate judges may also view the oversight of prosecutors as necessary when their misconduct significantly favors the position of the prosecution (Medwed, 2012). We thus interpret our findings as emerging within a specific historical context and as a potential reaction to changing power dynamics between criminal justice actors. Judicial discretion narrowed considerably over the past few decades as politicians, intent on enacting a “tough on crime” agenda, took aim at the powers of the judiciary, primarily in the realm of sentencing (Simon, 2007). With the introduction of legislative efforts to reign in the power of the judiciary, discretion in case processing fell increasingly into the hands of prosecutors, especially at the earliest stages (e.g., preparing for trial). Restoring some regulatory power over prosecutors who abuse their discretion, particularly in the backstage arena of investigation and trial preparation, may be in response to these recent changes.
Further investigation is needed, however, to assess this assertion and expand our findings, especially considering the limitations of our exploratory research. First, our sample is not nationally representative. Caution should be taken in generalizing these findings, particularly to states not represented here, and considering that our data are largely (although not wholly) drawn from published opinions. 16 Second, we cannot assess whether prosecutors who engage in multiple instances of the same type of misconduct within a case are viewed differently by the courts than those who engage in a particular type of misconduct only once. While the court opinions comprising our sample rarely identified multiple different instances of the same type of misconduct, it is an empirical question that is worthy of consideration in future work. Third, research is needed to investigate how cases involving both backstage and front-stage forms of misconduct compare to those cases that involve misconduct associated with a distinct stage. Here, our measure of stage of misconduct compared cases with any backstage misconduct, which could have also involved front-stage errors, to that of front-stage misconduct.
Our sample also represents a selective set of appellate decisions—ones in which the fact of misconduct is substantiated—that embody a mere microcosm of the much broader group of cases in which defendants allege prosecutorial misconduct to have taken place. This is important because it is possible that appellate justices are more willing to acknowledge error when they can do so without having to overturn the case, thereby upholding a correct verdict and the finality of the system. Egregious misconduct could thus go unaddressed in cases where the weight of the evidence against the defendant is weak. According to this logic, appellate courts systematically pick and choose which instances of prosecutorial misconduct to acknowledge, and thus offer symbolic admonishment against the behavior, but do so in a way that curbs any need to take concrete ameliorative action to remedy it. While we acknowledge this possibility, such that we effectively censor cases where appellate judges fail to affirm the defendant’s accusation of prosecutorial misconduct at appeal, internally consistent selection of this sort would seem to be quite an organized and, dare we say, malevolent feat for the courts.
Additional measures or approaches could augment our research on determinations of appellate relief in the future. For instance, a study examining all cases in which prosecutorial misconduct is alleged by defendants in their appeals could help distinguish whether differences exist in the nature and type of misconduct between those cases in which the appellate court substantiated it from those in which the appellate court found no basis for the allegation of misconduct. In a similar vein, bar disciplinary proceedings could also be examined, as these agencies do not typically discipline prosecutors unless the courts fail to address prosecutorial misconduct (Joy, 2004). It would be worthy to note how the conduct of the prosecutor in these proceedings differs from that which is corrected at the appellate level.
It would also be worthwhile to examine how quantity of evidence moderates the relationship between prosecutorial misconduct and corrective action. Given that trial records, and specifically the entire evidentiary package, are rarely disclosed in the narratives of judicial opinions, it is not possible to construct an accurate measure of evidentiary weight for each of these cases. Attention to the nature and type of curative instruction offered by trial judges may also shed light on patterns of professional deference. Finally, appellate judges themselves, in in-depth interviews, may best be able to articulate why they behave the way they do.
Despite these limitations, we bring the study of appellate oversight on prosecutorial misconduct, which has been predominantly pursued by legal experts, into the domain of social science by empirically assessing the conditions under which defendants receive relief for overzealous prosecution. Moreover, we identify an additional strand of research associated with the theoretical and empirical literature on courts as communities by alluding to the larger adversarial culture that structures interactions between appellate justices, trial judges, and prosecutors (Eisenstein et al., 1988; Ulmer & Johnson, 2004). Our research broadens criminal justice theorizing to consider when one major internal mechanism of oversight within the legal profession works to quell (or, more commonly, excuse) the bad acts of its members. That is, even when higher courts acknowledge clear error or misconduct on the part of prosecutors, they tend to legitimize, rather than condemn, prosecutors’ actions by failing to hold them accountable. We highlight how the decision-making of appellate courts around substantiated prosecutorial misconduct is driven not by whether prosecutors engage in misconduct at all, but by when they engage in specific types of misconduct and in how many ways.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the James Madison University Program of Grants for Faculty Assistance.
