Abstract
This article contributes to the practice of coding in meta-analyses by offering direction and advice for experienced and novice meta-analysts on the “how” of coding. The coding process, the invisible architecture of any meta-analysis, has received comparably little attention in methodological resources, leaving the research community with insufficient guidance on “how” it should be rigorously planned (i.e., cohere with the research objective), conducted (i.e., make reliable and valid coding decisions), and reported (i.e., in a sufficiently transparent manner for readers to comprehend the authors’ decision-making). A lack of rigor in these areas can lead to erroneous results, which is problematic for entire research communities who build their future knowledge upon meta-analyses. Along four steps, the guidelines presented here elucidate “how” the coding process can be performed in a coherent, efficient, and credible manner that enables connectivity with future research, thereby enhancing the reliability and validity of meta-analytic findings. Our recommendations also support editors and reviewers in advising authors on how to improve the rigor of their coding and ultimately establish higher quality standards in meta-analytic research.
Keywords
Systematic reviews that apply meta-analytic research methods have gained popularity in organizational and management research (cf. Aguinis et al., 2011; Kunisch et al., 2018). Moreover, meta-analyses are frequently cited (cf. Carlson and Ji, 2011; DeSimone et al., 2019; Kunisch et al., 2018), emphasizing their importance as a means of synthesizing existing knowledge and directing future research. Toward enhancing the rigor of meta-analyses, methodological resources, including handbooks (e.g., The Handbook of Research Synthesis and Meta-analysis, Cooper et al., 2019; Meta-analysis in Social Science, Glass et al., 1981; Cochrane Handbook for Systematic Reviews of Interventions, Higgins and Green, 2008) and reporting standards have emerged providing guidelines for conducting and reporting meta-analytic research (e.g., MARS, 2008, 2018; MOOSE, 2000; PRISMA, 2009; QUOROM, 1999).
As DeSimone et al. (2020) points out, “although these guidelines are comprehensive, they are often also quite general, which prevents the inclusion of detailed summaries or guidelines concerning specific aspects [such as] coding” (p. 2). In fact, among the different issues explained in existing guidelines for conducting meta-analyses, the coding process, i.e., how meta-analysts plan, conduct, and report their coding, has received less scholarly attention from methodologists. This marks a critical situation, as coding is the invisible architecture behind any meta-analysis; it spans the entire meta-analytic research project and connects research questions with theory, hypotheses, and meta-analytic findings. Perhaps due to the complexity and specificity involved, coding is rarely covered in methodological resources, leaving the research community with a lack of guidance on how to approach this sensitive task when conducting meta-analyses.
Insufficient guidance on the coding process can be detrimental to meta-analytic practice for several reasons. First, awareness is lacking that coding is not a purely technical task, such as extracting effect sizes from primary studies. Consequently, it cannot be treated independently or separately from the meta-analysis research procedure. The lack of attention on coding in existing methodological resources obscures its integral role in the development of the meta-analysis. For instance, the research questions posed in a meta-analysis determine the design of the coding process. A misfit between the research questions and the coding process design can, however, result in the researchers’ inability to adequately answer such questions. Second, considering the highly subjective nature of coding, meta-analysts face numerous decision points with far-reaching consequences for the reliability and validity of their results. Making coding decisions without reference to best-practices can lead to false meta-analytic results, which are then erroneously accepted and built upon by the research community. Third, the lack of guidance on the coding process has led to a void in documentation on published meta-analyses, as we observed in our content analysis of meta-analyses in top-tier journal publications (see Appendix A). This lack constrains the readers of meta-analyses from assessing the authors’ data extraction process and coding decisions and thus how their findings should be interpreted in the light of readers’ own future research. As such, insufficient reporting on the coding process results in missed opportunities for scientists to build upon existing meta-analyses in an informed and coherent manner.
The purpose of this article is to contribute to the practice of coding in meta-analyses by offering direction and advice on the process. To substantiate and illustrate our guidelines, we present findings from our content analysis of 124 top-tier meta-analytic reviews published over the last three decades and surveys of their authors. We also use data from an unpublished meta-analysis to present an empirical demonstration of the error-proneness of coding decisions.
Our assessment of such extant methodological resources revealed that advice on coding focuses primarily on “who” should perform the coding (the “who” of coding) and “what” information should be coded (the “what” of coding). However, recommendations on “how” the coding process should be planned, conducted, and reported (the “how” of coding) were largely absent. To supplement prior seminal meta-analytic handbooks and reporting standards that have contributed to raising quality standards for the “who” and “what” of coding in meta-analytic studies (e.g., Cooper et al., 2019; Glass et al., 1981; Higgins and Green, 2008; MARS, 2018; PRISMA, 2009), we here recommend best-practices for the “how” of coding.
We first define coding in the context of meta-analytic research, followed by a review and state-of-the-art assessment of existing advice in extant methodological resources on the task of coding. Next, we present our guidelines for each of four steps in the coding process: Coding approach (Step 1), Data extraction (Step 2), Categorization (Step 3), and Reporting (Step 4). We conclude with a brief discussion to sum up the recommendations provided.
Coding in Meta-Analyses
Defining Coding
When coding, researchers examine each study selected (Hunt, 1997), extract relevant information from it (Siddaway et al., 2019), and ultimately aggregate diverse data from numerous studies into a single database, making the data suitable for meta-analytic estimations (Cooper, 2017; Ellis, 2010). Hence, coding is the process of translating independent primary studies (i.e., the individual studies comprising the meta-analysis sample) into a common language to enable meaningful comparisons and syntheses of empirical findings (Glass et al., 1981; Lipsey & Wilson, 2001).
Analogous to survey research, coding in meta-analysis can be seen as the “interviewing” of a study rather than a person (Cooper, 2017; Lipsey & Wilson, 2001) and is one of the most technically demanding (Cooper, 2009, 2017; Hunt, 1997; Stanley et al., 2013) and time-intensive steps in a meta-analytic research project (Steinkamp, 1998; Wilson, 2009). Coders, therefore, must not only have a thorough understanding of the coding scheme but also be equipped with context-specific experience to properly understand the content of the studies to be coded (Lipsey & Wilson, 2001). In coding primary studies, the ultimate pursuit of meta-analysts is converting a complex and messy reality to a matrix of numbers (Orwin & Vevea, 2009). However, this endeavor is arguably so difficult to attain that some scholars consider it impossible to being perfect (Cooper, 2017), alluding to the vulnerability of the coding task to biases attributable to transcription errors, miscomprehension of coding instructions, or lack of construct validity in primary studies (cf. Murray, 2011; Wanous et al., 1989).
To arrive at that matrix of numbers ready for subsequent analyses—the final output of the coding process—both word and number coding is required (cf. Lipsey and Wilson, 2001). Word coding is the process of extracting information relevant to the meta-analyst's research questions such as labels, definitions, and measurement items of variables from primary studies. The number coding process entails both extracting numerical data such as correlation coefficients, Cronbach's alphas, and sample sizes from primary studies and converting the finalized word coding into numbers, i.e., attributing example codes to coding categories. While word coding is subject to varying degrees of subjectivity and coder bias, such as grouping variables into categories (Card, 2015; Cooper, 2017); number coding, extracting correlation coefficients from primary studies’ correlation matrices, requires less or even no judgment. The word coding process, however, is highly subjective, and consequently where error is likely to originate and where researchers are most likely in need of guidance. As such, our guidelines focus on word coding (hereinafter “coding” refers to “word coding”).
As existing methodological resources for conducting meta-analyses are quite general in nature (cf. DeSimone et al., 2020) and suggestions for coding best practices are highly fragmented, we performed a state-of-the-art assessment to synthesize existing advice on coding, which we present in the next section.
Existing Advice on Coding: The “Who” and “What”
Our analysis of methodological resources on conducting meta-analyses (e.g., Cooper et al., 2019; Glass et al., 1981; Higgins and Green, 2008; MARS, 2008, 2018; MOOSE, 2000; PRISMA, 2009; QUOROM, 1999) revealed that advice for coding orients researchers in terms of “who” should perform the coding and “what” information should be coded. This corresponds to the result of our analysis of top-tier meta-analytic publications (see Appendix A); those authors seemed to follow presently available methodological advice and primarily report on the “who” and “what” of the coding they perform.
The “Who” of Coding
The “who” of coding refers to the human element of coding, to the coders who perform the coding of the studies comprising the meta-analysis sample (cf. Cooper, 2017; Rothstein and McDaniel, 1989). Researchers are advised to (a) include multiple coders, at least two, who independently code the studies and to report the number of coders in their review (cf. Glass et al., 1981; MARS, 2008; Steel et al., 2021); (b) provide those coders with sufficient training before actual coding to familiarize them with the coding scheme and clarify potential ambiguities in the coding manual that need refinement and also to transparently document the details of the training provided (cf. Higgins and Deeks, 2008; Wilson, 2009); (c) select coders who are at least at the doctoral student level and who have sufficient experience to master the technically demanding aspects of coding and to explicitly report the coders’ identities (e.g., students or authors as coders; cf. Cooper, 2009, 2017; Lipsey and Wilson, 2001; MOOSE, 2000); (d) choose coders from complementary disciplines (e.g., a methodologist and a topic area specialist) and report on the coders’ level of knowledge and expertise in respective research fields related to the meta-analysis research topic (cf. Higgins and Deeks, 2008; Orwin and Vevea, 2009); and (e) to always perform a second coding (double coding) of the sample and report on the extent to which such second coding was performed (e.g., full vs. partial second coding; cf. MARS, 2018; PRISMA, 2009; see S1, supplemental material, for an extensive list of methodological references).
The “What” of Coding
The “what” of coding pertains to the data aspect of coding, the type of information coders extract from studies for estimation of hypothesized relationships (cf. Borenstein et al., 2009; Glass et al., 1981). Researchers are advised to: (a) specify the type of information that coders must extract when coding a study (e.g., what study characteristics, variables, effect sizes need to be coded; cf. Cooper, 2009, 2017; PRISMA, 2009; Wanous et al., 1989); (b) determine and report the inclusion criteria a study must meet for coders to consider it pertinent for coding (cf. Littell et al., 2008; Siddaway et al., 2019); (c) use at least two databases to identify relevant studies for coding and to report the databases used (cf. Card, 2015; Hunter et al., 1982); (d) decide which keywords best circumscribe the focal construct or phenomenon of the meta-analysis and display the depth of data extraction by indicating whether the title, abstract, or full text of the primary study was searched (cf. MARS, 2008, 2018; MOOSE, 2000; PRISMA, 2009; Steel et al., 2021); and finally (e) to reduce the risk of publication bias, always consider unpublished studies (grey literature) for coding and indicate the percentage of unpublished studies represented in the meta-analysis sample (cf. Glass et al., 1981; Rothstein and Hopewell, 2009).
Building upon these cardinal methodological resources, we now intend to bridge the gap between the “who” and “what” of coding by offering insight into “how” to proceed with coding in a meta-analytic research project. Our guidelines may empower scholars to rigorously plan (i.e., cohere with the research objective), conduct (i.e., make reliable and valid coding decisions), and report (i.e., in a sufficiently transparent manner for readers to comprehend the authors’ decision-making) the coding process. We hope to offer scholars consulting our guidelines the opportunity to better understand what meaningful coding practices entail, which pitfalls exist, and how coding decisions can impact the reliability and validity of meta-analytic results. Moreover, our guidelines may better equip researchers to effectively persuade journal editors and reviewers of the quality of their coding and increase the chance that other researchers will connect with and build upon their meta-analyses. In the following, we present our guidelines for the coding process.
Guidelines for the Coding Process: The “How” of Coding
The guidelines presented here establish best practices for the planning, conducting, and reporting of the coding process. Recognizing that knowledge on coding is fragmented and only rarely documented, we draw on our own extensive coding experience, on the coding experience of other meta-analysts, and on existing knowledge on the coding process, as reported in published meta-analyses (see Appendix A). We follow the common logic of the scientific process (cf. Cooper et al., 2019) to enrich and complete recommendations on steps to take when performing the coding process. These guidelines span integrating the research objective of the meta-analysis, the set-up of the coding process, the data preparation in terms of categorization, and, ultimately, transparently documenting it. As such, our guidelines acknowledge that the coding process is not an uncoupled but a highly integrated element of meta-analysis.
To substantiate our recommendations and illustrate the application of our guidelines, we refer to numerous supplementary works. These entail 1) our content analysis of 124 meta-analytic reviews published in seven top-tier journals (which assess the rigor and transparency meta-analysts apply when reporting on coding; see Appendix A and S2, supplemental material); 2) an unpublished meta-analysis on the functional diversity-performance relationship (to demonstrate how different coding approaches can cause variances in meta-analytic results; see S3, supplemental material); and 3) a survey with the authors of the corresponding meta-analyses in the content-analysis sample (to understand to what extent editors and reviewers ascertain the quality of coding during the review process; see S4, supplemental material).
Our guidelines offer support on how to perform the four major steps of the coding process (Coding approach, Data extraction, Categorization, and Reporting) coherently, efficiently, and credibly, thereby enhancing the reliability and validity of meta-analytic findings and promoting connectivity with future research (see Figure 1). Reliability and validity are pivotal to the accuracy of coded data and in turn to the findings of meta-analyses. Coherence, efficiency, credibility, and connectivity may not directly influence the findings of meta-analyses but may add to the quality and transparency of reporting and make the work more citable and accessible. We elaborate on these goals for each step below.

Guidelines for the coding process (the “how” of coding).
Step 1: Coding Approach
Step 1 involves identifying the research objective and conveying to coders which variables need to be coded and whether deductive or inductive coding should be applied. Research questions and hypotheses are important to consider when selecting data sources and establishing inclusion criteria (Creswell & Plano Clark, 2007). It is imperative that researchers realize that the research objective directs the coding process (i.e., the “how” of coding), and coherence and efficiency will be achieved through thorough reflection on the research objective. Accordingly, the planning phase of the coding process should precede the actual coding to ensure that the information collected is aligned with the research question and that iterative loops are minimized (cf. Booth et al., 2016).
Identify the Research Objective
Identification of the research objective of a meta-analysis is important for the coding process. In this context, we distinguish between “synthesis of a construct” and “synthesis of a phenomenon” (see Figure 1), the former being more specific and predictable and the latter broader and more iterative.
A construct refers to a mental abstraction that has “been deliberately and consciously invented or adopted for a special or scientific purpose” (Kerlinger & Lee, 2000, p. 40) and that can, as such, be measured. As ensuring construct validity and reliability in developing construct operationalizations can be complex (Churchill, 1979), researchers typically use the same established set of items to measure the same mental abstraction. Indeed, in meta-analysis, synthesis of a construct relates to the origin of meta-analytic endeavors, as first introduced by Glass (1976) and advanced by Hunter and Schmidt (1990). Here, the objective is to synthesize the effect sizes of studies that have investigated the same constructs (Geyskens et al., 2009). This typically entails integration of the state-of-the-art of a given construct (e.g., entrepreneurial orientation; see Rosenbusch et al. (2013) and Schweiger et al. (2019)), occurring in multiple primary studies using the same or similar measurements. The variables involved will have been researched for several years and pertain to established constructs. Thus, researchers seek to understand where differences in findings among individual studies and their relationship(s) to other variables originate. Specifically, researchers aim to describe boundary conditions of a given relationship by identifying moderator variables that vary the effect size of the relationship. Such contributions to the research community report on the state of extant knowledge on a construct in aggregate and suggest directions for future research.
By contrast, a phenomenon, in the context of meta-analysis, is a broader research subject around a general underlying and recurring idea. With synthesis of a phenomenon, researchers aim to make sense of a broad topic, frequently one with fuzzy boundaries, sometimes by integrating constructs from different research streams. The meta-analyst does not subsume similarly measured constructs but rather conceptually comparable variables associated with the same overarching phenomenon. For instance, meta-analytic syntheses of a phenomenon are conducted on applicant attraction outcomes and predictors of such by Chapman et al. (2005), on leader behaviors and organizational justice by Karam et al. (2019), and on exploratory and exploitative innovation by Mueller et al. (2013). Other meta-analysts synthesize a theory, which entails subsuming primary studies under an umbrella phenomenon that constitutes such theory, as in Van den Broeck et al. (2016) on self-determination theory, Geyskens et al. (2006) on transaction cost theory, and Heugens and Lander (2009) on institutional theory. With these syntheses of a phenomenon, meta-analysts contribute to the literature by abstracting from the constructs under study to arrive at higher levels of aggregation.
Identifying the objective of research is pivotal to the entire coding process as it determines where on the spectrum between synthesis of a construct and synthesis of a phenomenon the meta-analysis will fall. The differing degrees of homogeneity or heterogeneity of the pieces of research to be integrated are the determining factors for all subsequent steps of the coding process. As illustrated in Figure 1, this in turn leads to the next decision, whether to extract the information through inductive (i.e., data-driven) or deductive (i.e., theory-driven) means.
Decide on the Coding Approach—Inductive or Deductive
The research objectives and, accordingly, the type of syntheses selected determine the degree to which inductive or deductive coding approaches or combinations thereof apply (Creswell & Plano Clark, 2007). Deductive reasoning is concerned with testing existing theory by confronting it with observations in given circumstances (Eriksson & Kovalainen, 2015). Deductive coding is required when variables to be coded are explicit from the outset (e.g., meta-analysis on entrepreneurial orientation, Rosenbusch et al., 2013) and have been measured similarly in primary studies. Such a synthesis implies predefined and straightforward coding. The coder consistently searches every primary study for information (e.g., variable labels, definitions, measurement items, perceived vs. objective measurements, cross-sectional vs. lagged measurements, and the like) that relates to their focal constructs (e.g., entrepreneurial orientation).
Inductive reasoning involves developing theory by observing patterns in specific observations and arriving at broad generalizations (Eriksson & Kovalainen, 2015). Coders conducting inductive coding subsume different latent variables from primary studies among a broader phenomenon (e.g., meta-analysis on application attraction outcomes and its predictors, Chapman et al., 2005). Synthesis of a phenomenon requires more coder judgment when selecting information to include than synthesis of a construct as phenomena must be investigated from multiple perspectives. The observed information in primary studies is approached in a more explorative manner, requiring the coder to constantly allocate, reconsider, aggregate, and refine identified data from study to study. Thereby, meaningful categories of variables emerge from the original data. We refer to coding categories as distinct classes by which the variables of the primary studies are classified by the coder during the coding process (also, see Step 3). Such classification can lead either to a new umbrella category for the synthesized variables or an umbrella category based on a theory the researcher imposes on the original data. As such, the inductive (compared to the deductive) approach requires more openness and iterative coding behavior from coders, as they must oscillate between the collected data and the conceptual literature. Eventually, when categories have emerged, coders would then recode all studies in the coding scheme to validate the developed categories. Our content analysis of 124 meta-analytic reviews (see Appendix A) revealed that some authors already explicitly describe the extent to which they adopted deductive or inductive coding approaches, or both (e.g., Bhaskar-Shrinivas et al., 2005; Frazier et al., 2017; Kooij et al., 2011).
Combining deductive and inductive coding approaches within a single meta-analytic research project are common. For instance, in the unpublished meta-analysis on the functional diversity-performance relationship (see S3, supplemental material), functional diversity is the focal construct. Levels of functional diversity, frequently studied in team literature, are predominantly measured by the number of functional roles present in a team (cf. Ancona and Caldwell, 1992; Bunderson and Sutcliffe, 2002). Deductive coding has been adequate for meta-analyses of literature on functional diversity because the researchers knew which information (i.e., variable labels, definition, and measurement items) was of interest at the outset. By contrast, team performance is a less frequently systematically addressed construct of team research (Ilgen, 1999; Mathieu et al., 2008). Accordingly, the researchers who conducted the unpublished meta-analysis on the topic decided to code performance inductively thereby allowing performance dimensions, more specifically, categories of performance outcomes such as team efficiency, team innovativeness, new product financial success, new product novelty, etc.) to emerge from the data.
Step 2: Data Extraction
Step 2 involves developing the coding scheme and subsequent pre-testing, during which coding is practiced and improved upon before the complete set of primary studies is approached. In this step, the meta-analyst aims to achieve the reliability of the coding scheme, which serves as a standardized means of how information is collected. More precise coding schemes are more reliable and allow coders to conduct congruent coding. Further, this step aims to facilitate an efficient coding process by avoiding pitfalls that can result in the need to recode certain studies or even the entire sample.
Develop Initial Coding Scheme
The coding scheme (i.e., codebook, coding protocol, coding manual) documents the content coders extract from the original studies (Booth et al., 2016; Glass, 1976; Lipsey & Wilson, 2001; Wilson, 2009). Typically, coding schemes for meta-analyses of organizational research literature encompass 1) information on independent and dependent variables, moderators and mediators, and control variables (e.g., variable labels, definitions, and measurement items used in an original study, and if applicable, reliability estimates) and other variables of interest; 2) data (e.g., data type, perceived vs. objective, firm size, and firm age); 3) quality (e.g., journal ranking and study setting such as geographic location); and 4) quantitative information (e.g., sample size, effect sizes such as correlations between variables of interest). The research objective and synthesis type drive the content sought through the coding scheme, and consequently, determine whether that information should be coded deductively or inductively.
The coding scheme structures a uniform and standardized process of entering the information in a spreadsheet (Gaur & Kumar, 2018). A well-designed coding scheme directs coders in data reduction (i.e., when simplifying and abstracting original data) and data organization (i.e., when collocating the extracted data to make it interpretable). It enables coders to document qualitative (e.g., theoretical perspective) and quantitative (e.g., sample size) information relevant to the given meta-analysis (Cooper, 2009; Wilson, 2009). It presents the studies to be included in the meta-analysis in rows, and column labels guide coders on variables to be extracted from primary studies and on how such extraction should be conducted (e.g., type of measurement scale used for the entrepreneurial orientation-construct: 1 = Covin and Slevin (1989); 2 = Lumpkin and Dess (1996); 3 = other). For syntheses of a construct (e.g., on entrepreneurial orientation), very specific variable labels and well-established measurement scales are indicated to coders, and for syntheses of a phenomenon, broader column labels such as “antecedents” and “outcomes” are specified in the coding scheme along with broad descriptions of information that could potentially be subsumed by the category.
For information to be coded deductively, the coding scheme for how data should be extracted is quite predefined as variables of interest are sufficiently specific at the outset. That is, coders are guided by well-established definitions and measurements for the variables to be coded, and these are subject only to minor changes during the coding process. In contrast, an inductive coding approach entails allowing categories to emerge from the data presented in the primary studies, and so the initial coding scheme cannot be as straightforwardly predefined. The coding scheme is iteratively adapted as the coding of primary studies proceeds and meta-analysts and coders gain more insight into the data. In this way, they create meaningful coding categories for the coding scheme.
We recommend constructing the most fine-grained, comprehensive coding scheme possible (i.e., coding all information of potential interest) and as such ensuring that all coded data can be traced back to the original studies (Machi & McEvoy, 2016), thereby avoiding the need for multiple recoding. Comprehensive and fine-grained coding can be conducted by splitting multiple data entries. This requires simultaneous data entries with varying degrees of broadness, especially for information to be coded inductively. The first data entries in the coding scheme should be originally taken from a given study in a 1:1 manner—either as narrative text (e.g., original text passages of the study informing authors’ definition of a variable) or original word coding (e.g., extracting items used by the original authors for measuring a variable). While the first codes may be broad narrative entries, as coding proceeds, subsequent codes for the same data lead the coder to the progressively narrower first-, second-, and third-order categories that emerge.
For example, the authors of the unpublished meta-analysis coded “performance” inductively (see S3, supplemental material), and accordingly, the first coding entries for performance were original text passages that coders extracted verbatim from primary studies. The subsequent narrower codes for performance included only excerpts, such as the definition and key measurement items of the performance variable, from the first-order data. In this manner, the authors created multiple data entries for “performance” with varying degrees of specificity. In doing so, they identified seven performance outcomes at two different levels (team- and product-level) of analysis. In this way, the authors of the unpublished meta-analysis were able to regress functional diversity on specific performance outcomes (e.g., team innovativeness or new product novelty) rather than solely on a single aggregated performance outcome. Consequently, they were able to reach more fine-grained conclusions than authors of other meta-analyses on this topic (e.g., Hülsheger et al., 2009; Sivasubramaniam et al., 2012), as they could delineate the extent to which functional diversity is beneficial to the performance of a team, respectively for a new product (see Table 1 in S3, supplemental material).
As such, if provided in the majority of primary studies in the sample, we recommend extracting information that further aids in disaggregating variables (e.g., level of analysis, rater level, data source dependence, and the like), as this could make more in-depth analyses possible. A fine-grained coding approach and a detailed coding scheme can provide opportunities for conducting more in-depth analyses, which can lead to novel meta-analytic insights.
When the dataset becomes very large, maintaining an overview of coded information becomes increasingly difficult. We recommend combining word coding (i.e., the extraction of original text from primary studies) with assigning preliminary number codes to similar content (e.g., the country where the study was conducted: Switzerland = 1, Germany = 2, Austria = 3) wherever possible from the beginning of coding onwards; such numerical filtering makes databases considerably easier to handle. This is more easily accomplished in deductive coding because categories are often predefined. By contrast, when coding inductively, categories emerge as coding proceeds; numbers then cannot typically be assigned until the last studies have been coded. This entails that coders include as much detail as possible when they perform word coding to facilitate categorization and subsequent number coding based on those entries.
Conduct a Pre-Test
For quality enhancement and minimization of iterative loops, we recommend pre-testing the coding scheme on a subset of primary studies that, ideally, represents the full sample in terms of journal quality characteristics and information to be coded (Cooper, 2009; Higgins & Deeks, 2008; Littell et al., 2008; Orwin & Vevea, 2009; PRISMA, 2009; Weber, 1990). Pre-testing enables meta-analysts to evaluate the coding scheme along with associated decision rules and data extraction procedures, leading to greater precision and, in turn, greater reliability. In addition, the results of a pre-test can further refine the study sample. For instance, during the pre-test, coders may encounter studies that are unsuitable for the intended constructs or phenomena to be synthesized, leading to adjustments of exclusion and inclusion criteria. Further, a pre-test can uncover missing codes, such as those that potentially constitute heterogeneity between studies and those that need further specification (Cooper et al., 2009). Despite pre-test benefits, our content analysis of meta-analytic reviews (see Appendix A) revealed that only a few authors reported to have pre-tested their coding schemes (e.g., Blume et al., 2010; Joshi and Roh, 2009; or Sleesman et al., 2012).
When information is to be coded inductively, boundary conditions are not definable from the outset, questioning the feasibility of a pre-test. Although it is not as straightforward as for deductive coding, we still recommend running a pre-test as an initial means of engaging with the coding scheme, gaining understanding of the literature, and activating discussion among coders. In contrast to deductively coded information, it is likely that the definitions of inductively obtained information will need to be amended after the pre-test due to the high degree of judgment and interpretation involved. When working with information to be coded inductively, the meta-analyst should employ pre-testing in an exploratory fashion without aspiring to determine final definitions and boundary conditions.
Step 3: Categorization
Step 3 entails developing meaningful categories—a highly sensitive task that influences the explanatory power and quality of the subsequent meta-analytic findings. Building categories means grouping information from primary studies into classes, and thereby abstracting from the data, leading to data reduction and simplification of comparisons between studies (Brewerton & Millward, 2001; O'Reilly et al., 2012). As this step is significant for the reliability and validity of the coding and the meaningfulness of the meta-analysis, we regard the following recommendations as pivotal. Specifically, coders need to consider the meaning of a primary studies’ variables and categorize based on the indicated measurement items. In particular, high-inference variables that require inductive coding pose a challenge to coders in developing mutually exclusive categories.
Categorize Based on Measurement
Variable labels and definitions often do not or only partially correspond to the measurements of the variables in primary studies (Churchill & Peter, 1984; Lipsey & Wilson, 2001; Rousseau et al., 2008). Indeed, this lack of construct validity, one of the challenges most frequently encountered during the coding process, threatens the internal validity of the meta-analysis itself (Card, 2015; Chapman et al., 2005; Damanpour, 1991; Fainshmidt et al., 2016; Steel et al., 2021).
Building on prior researchers’ calls (cf. Aguinis et al., 2018; Rothstein and McDaniel, 1989; Wanous et al., 1989), we examine whether and how authors’ coding decisions and subjective judgments can cause variances in meta-analytic results and consequently lead to erroneous conclusions. Using an unpublished meta-analysis for empirical demonstration (see S3, supplemental material), we found that different coding approaches (i.e., coding variables based on their measurement items vs. coding variables based only on the label provided by the original authors) can lead to different meta-analytic results. In fact, comparing the effect sizes generated from each of the two coding approaches reveals that the significance and magnitude of meta-analytic effect sizes and even the directions of relationships change substantially depending on the coding approach adopted (see Tables 1 and 2 in S3, supplemental material). If meta-analysts code variables based only on the primary study labels, considered the lowest level of specificity for coding (Christian et al., 2010), and ignore how the variables were actually measured in those studies (i.e., coding based on measurement items), they may attribute variables to the wrong coding categories, ultimately distorting the basis for meta-analytic calculations and the results of the meta-analysis.
Consequently, we advise coding measurement items, variable labels, and definitions as documented by the original authors of the given primary study (cf. Cooper, 2009; Larsen and Bong, 2016; Steel et al., 2021). This direct 1:1 approach has several advantages. First, coding the original information allows authors to reassure themselves at any point in the coding process, analyze what the authors of the original studies measured, and determine whether a coded variable has been attributed to an adequate coding category, and thus provides the basis for subsequent meta-analytic analysis.
Second, coding measurement items enables meta-analysts to categorize according to what the original authors actually measured rather than according to what they intended to measure, resulting in higher degrees of internal validity for the category and the meta-analysis. Our content analysis (Appendix A) revealed that only a few meta-analysts explicitly refer to applying a coding approach based on measurement items (e.g., Christian et al., 2010; Damanpour, 1991; Klier et al., 2017; Sihag and Rijsdijk, 2019; Zhao et al., 2007). A frequent challenge for meta-analysts when coding measurement items is that authors of the primary studies comprising the sample report only minimal examples of items they used for operationalizing their variables instead of documenting the entire measurement scale (full list of measured items). If not provided in the (web-)Appendix to the present study, meta-analysts are encouraged to contact the respective authors to receive full information for adequately categorizing variables of interest.
Third, coding measurement items as reported in the original studies allows meta-analysts to identify measurement differences within the same underlying constructs (cf. DeSimone et al., 2020; Steel et al., 2021). For instance, although the authors of the unpublished meta-analysis found that the construct “functional diversity” is predominantly measured by the number of functional roles in a team, they also found that alternative conceptualizations of functional diversity (cf. Bunderson and Sutcliffe, 2002) are being increasingly applied. When meta-analysts notice that different measures are being used for the same underlying construct, they can use that in-depth information to run moderation analyses (comparing different measurement scales for the same construct). Such a moderation analyses can provide new explanations for the equivocality of findings for a specific relationship (e.g., depending on how functional diversity has been measured, some scholars find it has a positive effect on performance, while others find a negative or non-significant impact). Considering that one purpose of a meta-analysis is to generate explanations for controversial findings in a research field (cf. DeSimone et al., 2019), a coding approach based on measurement items offers valuable information for creating and shaping the research contribution of a given meta-analysis.
Fourth, coding by the original variable labels (i.e., how authors name or label the variables in the study they conduct) is also important, because it allows the meta-analyst to determine degrees of uniformity or variety in the labeling of particular variables (Lipsey & Wilson, 2001). As Steel et al. (2021) pointed out, “there can be dozens of terms and scores of measures for the same construct (i.e., jingle) and different constructs can go by the same name (i.e., jangle)” (p. 5). For instance, in the unpublished meta-analysis, the authors realized that various labels were being used to refer to the construct “functional diversity” (e.g., cross-functional diversity, multi-knowledgeable teams, expertise diversity, and the like; cf. Lovelace et al., 2001; Park et al., 2009; Van Der Vegt and Bunderson, 2005). The degree of variety (or uniformity) of variable labels can constitute important insights for the research field. Scientists can use the different labels to enhance the quality of their literature searches, identify other scholars from related fields investigating the same construct or phenomenon, and ultimately, begin working towards greater construct consistency.
Develop Mutually Exclusive Categories
To varying degrees, categories are either theoretically imposed and predefined (deductive approach) or allowed to emerge from the data (inductive approach). The conceptual literature on the respective research topic, which will inform meta-analysts and coders about adequate coding categories, is relevant to deductive coding approaches. Some scholars suggest relying on previously developed procedures for classifying coded variables or using categories from other researchers’ coding schemes as an orientation (e.g., Christian et al., 2011; Gaur and Kumar, 2018; Judge and Ilies, 2002; Martin et al., 2016; Ng and Feldman, 2008).
For deductive and inductive coding alike, we recommend structured categorization approaches that lead to reproducible, consistent, valid categories. Specifically, we suggest defining category names and definitions, key measurement items, and boundary conditions as early as possible in the coding process, even when an inductive approach to coding is required. Further, we recommend a continuously feed of category descriptions during the coding process with example codes to further specify the meanings of categories (e.g., Loignon and Woehr, 2018; see Table 5 in their online appendix). As categories are to be altered as coding proceeds, this approach helps coders categorize distinctively between categories. Moreover, especially relevant to inductive coding approaches, when relations between categories are not predefined but rather emerge from the data, a hierarchical category system such as a tree diagram is recommended (e.g., Thomas, 2006). A continuously refined framework of categories is helpful for visualizing linkages between such categories.
Valid categories must be mutually exclusive; they must each have distinct theoretical scope and boundary conditions and not contain overlapping meaning. Thus, when categories are stable, indicating the end of the coding process, each coded variable will be clearly assignable to a single category. If coders can theoretically attribute a variable to multiple categories, this indicates that the categories are not mutually exclusive. While there are meta-analyses conducted with variables in multiple categories, such procedures distort the results and conclusions derived. Using data multiple times for estimating meta-analytic relationships can bias the results in the same way as using duplicate samples (cf. Wood, 2008). Hence, we recommend that each information feed into a single category only. As noted, visualizing linkages between categories can help meta-analysts develop boundary conditions and detect possible overlaps. If a variable or study cannot be clearly subsumed within a category, we advise revisiting category definitions and example measurement codes and determining the adaptations required to arrive at mutual exclusivity. Ultimately, it is also possible to remove such variables or studies from the database (Chamberlin et al., 2017).
Two varying variables being aggregated within one construct is another frequent challenge coders face. For instance, in the unpublished meta-analysis, the coders found that scholars sometimes aggregated different performance outcomes, such as team efficiency and team effectiveness, into a single construct, typically called “performance” (e.g., Ancona and Caldwell, 1992). As a decision rule on whether a variable can be coded to a category, we suggest adopting a two-thirds-cutoff approach (cf. Chamberlin et al., 2017). Specifically, when two thirds of the measurement items for a variable match the meta-analysis coding category, it should be coded as such. An alternative is to classify such variables into a mixed category to enhance the validity of the meta-analysis. In the case of the unpublished meta-analysis relied on here, the authors allocated variables that presented a lack of mutual exclusivity into a mixed category (e.g., mixed performance, see Table 1 in S3, supplemental material).
Step 4: Reporting
Step 4 involves the thoughtful—rather than extensive—documentation of the coding process, which is pivotal to the credibility of meta-analytic results (cf. Aguinis et al., 2018, 2020; Aytug et al., 2012; DeSimone et al., 2020; Gibbert and Ruigrok, 2010). To further facilitate connectivity with future research, and because coding decisions or judgments may vary substantially among teams of meta-analysts investigating the same topic and may lead to different meta-analytic findings (Cooper, 2009), such decisions or judgments should be explicitly recorded and reported.
Step 4 aims more at connectivity with future research than reproducibility. As such, we do not claim that authors must disclose coded numeric data such as effect sizes taken from the primary studies in their meta-analytic calculations. Instead, we argue that authors should disclose their coding schemes so that others can fully comprehend how the categories were built and which primary study variables were coded into which categories. Scholars require this information to assess the category of a given meta-analysis and to tie in with and build upon its findings in their own future research.
Ensure Transparency
While we found that the “who” and “what” of coding has been sufficiently reported on, our content-analysis of 124 meta-analyses published over the last three decades in top-tier journals (see Appendix A) suggests that the “how” of coding is rarely documented. Building on prior research investigating potential gatekeepers of the review process (e.g., Green et al., 2016), we surveyed the first authors of the 124 meta-analyses on what information editors and reviewers had requested from them during their review process to ascertain the quality of their coding (see S4, supplemental material). Our survey findings suggest that editors and reviewers primarily addressed coding features belonging to the “who” and “what” of coding (e.g., number of coders, choice of inclusion criteria, whether grey literature was included in the sample coded), leaving the coding process (i.e., “how” authors proceeded with their coding) relatively unchallenged. However, our analysis also reveals initial efforts in recent journal publications to disclose information on the coding process (e.g., Special Issue on Contemporary Meta-Analyses, Journal of Management Studies; Combs et al., 2019), indicating that journal editors and the research community are increasingly aware of the sensitivity of the coding process in meta-analyses.
Addressing emerging needs for increased consistency and coherence in reporting on coding while acknowledging the limited space available in academic journals, we recommend reporting the most crucial aspects that we have elaborated on in the previous steps of our guidelines. In the following, we refer to the minimum reporting we consider necessary in the interest of balancing full transparency with the space constraints of academic journals.
For Step 1, we recommend reporting which variables of the meta-analysis required inductive coding, as these are the primary sources of coder bias. For Step 2, data extraction, we suggest reporting whether a pre-test was conducted and resulting main insights that led to coding scheme amendments. Further, we suggest including the final coding scheme (online) in an appendix to the meta-analysis that reports the variables coded from each study (e.g., Knight et al., 2017, see File 2 in their online appendix; Montano et al., 2017, see File 1 in their online appendix). Similar to primary survey research, where information on conceptualization and measurement is common practice, a disclosed coding scheme should serve as an important means of transparency (e.g., disclosing example codes for coding categories and extracted information from original studies such as definitions and measurements of variables that are subsumed under a single category). For Step 3, categorization, meta-analysts should report which primary study variables were assigned to a category, each category's boundary conditions, and how mixed variables were treated (e.g., two-thirds cutoff criterion, mixed category, etc.). As good coding practice, we recommend explicitly stating in the methods section of the report on the meta-analysis that categories were developed based on measurement items.
We also suggest providing visual aids to help readers more easily gain a sense of the data collected. For instance, an advantage of including a frequency table (e.g., Sivasubramaniam et al., 2012, p. 820) that represents the coding of the entire sample is increased transparency regarding how authors grouped and categorized variables in their meta-analyses (e.g., Jiang et al., 2012, pp. 1289–1294; Karam et al., 2019; see Table 2 in their online appendix). Moreover, frequency tables allow for identification of highly researched variables at first glance as well as recognition of white spots in the research field that could become crucial avenues for future research (also, see our frequency table in S2, supplemental material). Additionally, it is reasonable to include a table displaying the construct definitions and key measurement items derived from selected references (e.g., Heugens and Lander, 2009, p. 68; Loignon and Woehr, 2018, see Table 5 & 7 in their online appendix); these function as example codes and guidance for coders during the entire coding process (cf. MARS, 2018). Such an operationalization table not only provides guidance to coders but also explains to the readership how the coding categories were built, allowing assessment of internal validity and the overall quality of the meta-analysis.
Describe Intercoder Reliability
Among the different measures of intercoder reliability (cf. Le Breton and Senter, 2008), the agreement rate expressed as a percentage (agreed-on codes/total number of codes) is the most popular, followed by Cohen's Kappa. Instead of simply measuring the percent of agreement, Cohen's Kappa takes into account that agreement can occur by chance (Cooper et al., 2009). Typically, intercoder reliability measures are used to justify coding decisions or claim that coding is valid. However, these statistics only report the degree of coding agreement and cannot reveal whether the coding is “correct” or biased.
Referring to our content analysis, we found that in the 124 meta-analyses (see Appendix A), the most frequently reported feature concerning the “how” of coding (i.e., the coding process) is the documentation of an aggregated intercoder reliability. That is, the authors report one-sentence statements such as “We had a high degree of interrater agreement with a Cohen's Kappa of 0.98.” This common reporting practice also corresponds with our survey (see S4, supplemental material), where an aggregated intercoder reliability value was also a prime indicator of coding quality that editors and reviewers requested from authors during the review process. Other features, such as the type of coding approach adopted or the coding scheme with extracted data, which would inform how coding categories were built, were seldom requested by editors and reviewers.
We argue that representing the coding process with an aggregated intercoder reliability value belies the highly iterative nature of coding. Indeed, coding consistency between coders is not achieved in an ad hoc manner and cannot be expressed as a single number alone. Instead, coding consistency should be the natural product of close collaboration between coders and should be part of an iterative process. As such, the explanatory power of a single number, such as aggregated intercoder reliability, is limited.
Moreover, an aggregated intercoder reliability calculated across all coding categories of a meta-analysis regardless of their degree of inference obscures which specific categories have been subject to debate during the coding process (e.g., contextual categories such as firm age usually reach intercoder reliability values of 1, as they are easier to code than high-inference categories representing psychometric constructs such as happiness or group satisfaction). Thus, we recommend reporting intercoder reliability measures for single coding categories, especially for high-inference categories (cf. Cooper, 2009; Orwin and Vevea, 2009; Steel et al., 2021) as in Loignon and Woehr (2018), Mackey et al. (2017), and Thatcher and Patel (2012). Reporting isolated statistics is still not sufficient to demonstrate valid coding. Often, high-inference categories involve complex back-and-forth between data, and evidence and hence are exposed to coders’ subjective judgment. As such, researchers should regard intercoder reliability measures as accompanying information that by itself cannot indicate quality control.
In addition to reporting intercoder reliability measures for single high-inference categories, we highly recommend reporting insights that disclose how the reliability level has been achieved and which coding categories were subject to debate and why. Typically, meta-analysts report that discussion meetings took place to solve coding disagreements until consensus among coders was reached (see Appendix A). More detail on the content of these discussions would enable other scholars to comprehend important coding decisions that shape the meta-analysis findings and category-development of their own meta-analytic research projects (DeSimone et al., 2020; Higgins & Green, 2008). Specifically, we recommend reporting the most challenging coding categories, i.e., related difficulties faced by coders, reasons for codes not matching, important final decisions in the case of debates, and clearly stating which categories underwent heavy regrouping, and ultimately, how procedural decisions on coding could potentially affect the meta-analytic estimations and findings.
Due to the generally limited explanatory power of intercoder reliability measures, we recommend reporting the simplest statistic (i.e., agreement rate in percentage in terms of agreed-on codes/total number of codes), as it should serve only as an indicator. We suggest calculating reliability measures at the very beginning, for instance after the pre-test and as needed for coder meetings as well as at the end of the coding process. Documenting an improvement in consistency and explaining its origin adds to the credibility and traceability of the coding process.
Discussion
This article contributes to the practice of coding in meta-analyses by shedding light on questions of “how” to plan, conduct, and report on the coding process in a coherent, efficient, reliable, valid, and credible manner that promotes connectivity with future research.
Compared to survey research, in which data and data measurement gain high levels of attention from the research community, there seems to be little awareness that coding entails actual data collection and the full meta-analytic data preparation process as well. In fact, coding is the integral part of a meta-analysis that connects research questions and meta-analytic findings. As it is susceptible to high degrees of subjective decision-making, we argue that the same rigor called for during data collection for primary research is likewise appropriate for the coding process in meta-analyses. Therefore, as in primary research, the coding scheme (cf. questionnaire), the coders (cf. interviewers), and the coding process (cf. data preparation) deserve attention; coding the targeted primary studies is pivotal to the validity and reliability of the meta-analytic findings (Lipsey & Wilson, 2001).
In the present research, we reviewed currently available guidance on coding for meta-analysts. By analyzing extant methodological resources such as meta-analytic handbooks and reporting standards (e.g., The Handbook of Research Synthesis and Meta-analysis, Cooper et al., 2019; Meta-analysis in Social Science, Glass et al., 1981; Cochrane Handbook for Systematic Reviews of Interventions, Higgins and Green, 2008; MARS, 2008, 2018; MOOSE, 2000; PRISMA, 2009; QUOROM, 1999), we found sufficient advice on the task of coding in terms of “who” should perform the coding and “what” information should be extracted from the sample studies. However, existing guidelines provide little information on the coding process itself, that is, “how” meta-analysts should plan, conduct, and report the coding of their studies. Indeed, our content analysis of 124 meta-analytic reviews (see Appendix A) and our survey of corresponding first authors (see S4, supplemental material) reveals that the coding process (i.e., the “how” of coding) is only partially addressed in meta-analytic publications and rarely subject to discussion during the review process in top-tier journals.
This lack of guidance for meta-analysts reflects lack of awareness regarding the crucial role of the coding process in meta-analytic practice. As our empirical demonstration based an unpublished meta-analysis suggests (see S3, supplemental material), how variables are coded directly impacts the results of a meta-analysis (see Step 3 in the present guidelines). Thus, following certain standards for the coding process determines the quality of the meta-analytic conclusions reached, and, consequently, advancement in the research field.
In four steps, our guidelines on the “how” of coding (see Figure 1) address key aspects that should be considered by meta-analysts before and during the coding process, as well as how to document this in their review articles. Step 1 (Coding approach) and step 2 (Data extraction) guide the meta-analyst in setting up the coding process efficiently and reliably. Step 3 (Categorization) determines how the data is framed and the validity of the meta-analytic findings. Step 4 (Reporting) describes the most important information a meta-analytic report should document to yield credibility and connectivity with future research. Greater transparency in reporting on the coding process will assist other researchers in making inferences about the coding decisions that back a meta-analytic review and in building on its findings in an informed and coherent manner.
Our guidelines should not be used as a rigid checklist. Rather, they should serve as an orientational and conceptual template for the planning, conducting, and reporting of the coding process. The relevance of each step and the extent to which iterations are necessary will vary depending on the research question driving the given meta-analysis, the research discipline, and the experience of the coders. Further, our guidelines cannot guarantee meaningful meta-analytic results per se, as there are numerous other aspects that influence the quality of a meta-analysis (cf. Gonzalez-Mulé and Aguinis, 2018; Wanous et al., 1989). For instance, besides the “how” of coding, the selection of primary studies in terms of the quality of the information they contain (i.e., the “what” of coding, cf. Hiebl, 2021), and coders’ level of experience in coding (i.e., the “who” of coding), are also important factors in the overall quality of a given meta-analysis.
When referring to our guidelines, meta-analysts and readers of meta-analyses can better evaluate what meaningful coding practices entail, the main pitfalls, and the extent to which the coding process should be reported. Our article aspires not only to empower experienced and novice meta-analysts but also to inform editors and reviewers on how to support authors in improving the quality of their coding and reporting, ultimately, to elevate the validity of meta-analytic reviews. Therefore, our guidelines on the “how” of coding seek to contribute to higher quality standards for meta-analyses like prior methodological resources offered on the “who” and “what” of coding (e.g., MARS, 2008, 2018; MOOSE, 2000; PRISMA, 2009; QUOROM, 1999; see also S5, supplemental material)—a standard for the reporting of coding that yields increased consistency among meta-analyses and, promotes interdisciplinary dialogue, empowering researchers to build their future scientific projects upon the knowledge that meta-analyses create.
Supplemental Material
sj-pdf-2-orm-10.1177_10944281211046312 - Supplemental material for Making the Invisible Visible: Guidelines for the Coding Process in Meta-Analyses
Supplemental material, sj-pdf-2-orm-10.1177_10944281211046312 for Making the Invisible Visible: Guidelines for the Coding Process in Meta-Analyses by Jessica Villiger, Simone A. Schweiger and Artur Baldauf in Organizational Research Methods
Footnotes
We express our gratitude to Markus Menz, the three anonymous reviewers, Paul Bliese, Sven Kunisch, Jean Bartunek, Laura Cardinal, and David Denyer for the valuable guidance they provided us throughout the review process and for giving us the opportunity to share our research with ORM's readership. We are also thankful to our friendly reviewers Katja Rost and Andreas Rauch, and the 35 distinguished scholars who participated in our survey. A special thank you goes to Isabel Stuber, Seraina Schönenberger and Lars Allgäuer for their ongoing support during this research project's journey.
Supplemental Materials
Online Supplement S1: Suggested Coding Practices for the “Who” and “What” of Coding in Methodological Resources
Online Supplement S2: Content Analysis of 124 Meta-Analytic Reviews: Identification and Coding of Meta-Analytic Reviews and Frequency of Reported Coding Features in the 124 Meta-Analyses
Online Supplement S3: Unpublished Meta-Analysis: How Different Coding Approaches Can Cause Variances in Meta-Analytic Results and Background Information on the Unpublished Meta-Analysis
Online Supplement S4: Survey: How Editors and Reviewers Ascertain the Quality of Coding During the Review Process
Online Supplement S5: Publication of Meta-Analytic Reporting Standards and Their Impact on Frequency of Reported Coding Features
References to the meta-analyses included in the content analysis (N = 124) are provided in the supplemental materials
Author Biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
