Abstract
Sound evaluation planning requires numerous decisions about how constructs in a program theory will be translated into measures and instruments that produce evaluation data. This article, the first in a dialogue exchange, examines how decisions about measurement are (and should be) made, especially in the context of small-scale local program settings. Rigorous measurement strategies will increase the credibility of a study’s conclusions, but they usually entail various kinds of costs. In making measurement decisions, evaluators must establish standards for strength of evidence that a given measure produces, weigh alternative measurement options, and communicate carefully with clients and other stakeholders about the measurement requirements in a given evaluation.
Social programs are called upon to produce evidence of their impacts, to a wide variety of interested parties that may include funders, agency partners, target audience members, legislators, internal administrators, and many others. When a program evaluation results in claims that desired impacts have been achieved, stakeholders’ acceptance of those claims will depend, in part, on their being confident that the evaluation study was well-planned, well-reasoned, well-executed, and focused on essential program outcome variables. 1 Experienced evaluators can anticipate these eventual stakeholder judgments early on, during the evaluation planning process, and address them through features of the evaluation design. In this article, I examine one aspect of that planning process: the measurement-related decisions that will have a bearing on the strength of the conclusions about program effectiveness and success.
The evaluations that I am concerned with here are, for the most part, local evaluations of programs that are delivered on a relatively limited scale. I work in the Cooperative Extension system, and there are dozens, if not hundreds, of programs fitting this description delivered in most states in any given year. Many of the evaluations are not high-visibility in the sense of reaching audiences beyond the direct scope of the programs. Outside of their intended local audiences, most of these evaluations do not get wider dissemination through publication or other means. Nevertheless, they often influence the fate of the programs in question (Braverman, Engle, Arnold, & Rennekamp, 2008). Similar conditions probably hold for many program evaluations conducted across a range of nonprofit and for-profit organizations.
The local nature and under-the-radar quality of these evaluations can raise challenging questions about rigor and the quality of evidence. In the absence of requirements for journal publication and other peer review mechanisms, the planning decisions about evaluation rigor can be perplexing. At present, decisions about rigor are often made by the evaluator, sometimes in consultation with others, sometimes not. I argue that the decisions should be resolved in the context of the local program setting with the involvement of local stakeholders. My discussion will concentrate specifically on the area of measurement and instrumentation. Certainly, every other aspect of the evaluation can be subjected to this analysis as well.
An example may serve to illustrate the kinds of measurement-related decisions that are involved. Many impact evaluations of community-based educational programs use participants’ ratings of their own skills and knowledge as the strategy for measuring the learning outcomes of interest. These self-ratings, considered a form of indirect assessment (Banta, 2004; Suskie, 2009), are often used in a standard pretest–posttest design, in which they are collected before the program begins and again after it has ended. Self-ratings are also sometimes used in a retrospective pretest design, in which both sets of ratings are collected after the program’s conclusion and participants are asked to reflect on and report their prior, preprogram status on the variables. But self-assessments of learning are susceptible to numerous potential biases and sources of error that are difficult to appraise. “True” pretest ratings and retrospective ratings have been found to be prone to different kinds of bias (Aiken & West, 1990; Hill & Betz, 2005; Lam & Bengo, 2003; Nimon, Zigarmi, & Allen, 2011; Pelfrey & Pelfrey, 2009; Taylor, Russ-Eft, & Taylor, 2009). An alternative strategy would be to develop a direct assessment or test of relevant skills and/or knowledge, to be administered before and after the program (possibly involving parallel forms to avoid repetition). The two alternative strategies of direct assessment and participant self-ratings may both target the same construct of interest, for example, learning within a specified content domain, but they approach the task in distinctly different ways. Direct assessment will generally yield a far superior quality of evidence, but will also entail considerably more time, resources, and expertise. Thus, resource-strapped local program staff may ask whether a strategy of indirect assessment, by itself, can ever be considered adequate. In view of considerations regarding measurement rigor and implementation feasibility, arriving at an informed, defensible, and satisfactory answer to this question for a given program evaluation setting is complicated. As stated at the outset, the judgments of important stakeholder audiences—actual or anticipated—will have a significant influence on the answer.
I begin with an examination of the concept of methodological rigor and its relationship to properties of evidence and the larger concept of evaluation validity. Next, I discuss the development of evaluation measures and I consider the factors that may influence the adequacy of measures for achieving particular evaluation purposes in particular contexts. The final section explores challenges faced by the evaluation planning team and the ways that the evaluator can address and negotiate rigor-related measurement decisions. Some of the examples I consider are drawn from comprehensive, high-resource evaluation studies, but for the most part my focus is on local, small-scale evaluation settings.
Rigor, Evidence, and the Validity of Evaluation Studies
Building from a previous formulation (Braverman & Arnold, 2008), I view methodological rigor as a set of characteristics of an evaluation study, which support the study’s underlying logic and influence the confidence with which conclusions can be drawn. In a similar vein, Coryn (2007) states that rigor concerns “the means by which integrity and competence are confirmed” (p. 26). Johnson, Kirkhart, Madison, Noley, and Solano-Flores (2008) note, “Rigor addresses the extent to which an evaluation adheres to strict standards of research methodology” (p. 199).
A rigorously conducted evaluation will be convincing as a presentation of evidence in support of an evaluation’s conclusions, and will presumably be more successful in withstanding scrutiny from critics. Rigor is multifaceted and relates to multiple dimensions of the evaluation. For example, in the area of study design, it relates to the determination of how, when, and from whom data will be collected, and the structure of the critical comparisons that address the questions of interest. In the area of measurement, it relates to the internal logic of the linkages among constructs and outcomes, as well as the adequacy with which those outcomes have been translated into measurement strategies and instruments. Methodological rigor can be assessed from the evaluation plan and the quality of the evaluation’s implementation.
The concept of rigor is understood and interpreted within the larger context of validity, which “concerns the soundness or trustworthiness of the inferences that are made from the results of the information gathering process” (Joint Committee on Standards for Educational Evaluation, 1994, p. 145). The concept of validity has its origins in theories of measurement, but has been extended to apply also to research and evaluation studies. There is relatively broad consensus that validity is a property of an inference, knowledge claim, or intended use, rather than a property either of a research or evaluation study (Chen, Donaldson, & Mark, 2011; Shadish, Cook, & Campbell, 2002) or of a test or measure (American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, 1999; Kane, 2006; Lissitz, 2009).
The technical aspects of an evaluation study that are associated with methodological rigor are directly linked to the quality of evidence that the study is able to produce. Schwandt (2009) has described three properties of evidence that relate to its value for supporting inferences: (a) the relevance of the evidence to the inference being made, (b) its credibility or believability, and (c) its probative (inferential) force, that is, the strength with which it supports the inference. An evaluation’s measurement-related planning decisions and implementation activities, that is, its measurement-related rigor, will influence the quality of evidence with respect to each of these three dimensions.
The relevance, credibility, and strength of evidence form the basis for the evaluation study’s subsequent interpretations and inferences. Those interpretations may involve the meaning of the measurement scores, relationships between variables, integrity of treatment, and representativeness of samples. The validity of the inferences will be assessed based on judgments about the evidential support for those inferences. Thus, validity does not come into play until the study is completed and one or more inferences, such as whether an intended program impact has been achieved, are drawn from the study’s findings.
In recent years, a great deal of attention has been given to the concepts of rigor and standards of evidence, but those discussions have been focused almost exclusively on the relative merits of experimental and quasi-experimental designs in comparison to alternative approaches for generating evidence in support of causal claims (e.g., Chatterji, 2007; Donaldson, Christie, & Mark, 2009; Julnes & Rog, 2007; Mosteller & Boruch, 2002). Much less attention among evaluators has been directed to how standards might be developed for rigorous measurement in order to yield valid interpretations about desired impacts.
Moving From Constructs to Variables to Measures
Measurement as Simplification
The sociologist Hubert Blalock (1979, 1982), writing about the relationship between theory and measurement, noted that measurement in social science, by its nature, is a process of simplification and ordering of the world. Simplification serves the purpose of advancing the generalizability of concepts across units, so that claims may be made about groups rather than individuals. For example, a questionnaire measuring weekly exercise and physical activity may use a single metric, such as minutes per week of vigorous exercise, to classify individuals along a single dimension to establish comparability, even though their actual patterns of exercise will vary from each other in multiple ways that the metric ignores. This simplification process, however helpful, will necessarily entail assumptions and potential biases. Blalock noted that social science theories are, explicitly or implicitly, accompanied by “auxiliary measurement theories” that underlie the use of whatever specific measures have been chosen. “In short,” he wrote, “we must become much more attentive to the need for stating explicit auxiliary measurement theories and for examining comparability of measurement, just as we must also be concerned about the generalizability of our substantive theories… If either process lags too far behind the other, we shall find ourselves stymied” (Blalock, 1982, p. 31).
Blalock’s view of measurement theory places importance on the ways that measures reflect their underlying constructs, and thus it has obvious parallels to construct validity, which many psychometricians consider the most fundamental form of validity in the measurement literature, incorporating aspects of both criterion and content validity (Kane, 2006). “Arguably in its most common current use, construct validity refers to the degree to which inferences can be made legitimately from the observed scores to the theoretical constructs about which these observations are supposed to contain information… The term ‘construct validity’ has therefore evolved to be shorthand for the expression ‘an articulated argument in support of the inferences made from scores.’” (Zumbo, 2009, pp. 68–69). In other words, construct validity is concerned with how well desired outcomes are represented by the scales, tests, survey items, observation forms, scoring rubrics, and other measures that evaluations employ.
The Specification Process
The measurement specification process generally begins with the broadly conceived target construct, which reflects, often in everyday language, the issue that the program is designed to address. Many of the stakeholders who come together to create social programs express their goals for the desired anticipated outcomes in broad terms. For example, they may want to reduce obesity, increase student interest and enrollment in science courses, or promote safe sexual practices and reduce incidence of sexually transmitted diseases. In stakeholders’ conversations about program goals, their conceptions might not be more precise than this. These general constructs need to be interpreted and operationalized before the program and the evaluation begin. This process presents a chain of specific decision points, each of which can present a wide array of options for the evaluation planning team.
Figure 1 presents a model and illustration of this process in four phases, mapping a chain of potential decision options for a program focused on nutrition education and healthy eating. The phases are marked by increasing specificity, and the expansion of options at each of the steps is readily apparent. The figure’s identification of those options is necessarily incomplete, as many more options can be suggested for each phase.

Illustration of the possible progression from construct to instrument (some paths left incomplete).
The Choice of Variables
A variable is a specific characteristic of a unit—person, family, and so on—that has differing values within a population of those units. The variables selected for an evaluation must be more precisely specified and measurable than the constructs from which they are drawn. One aspect of rigor in this step of the translation process will be the presence of a clear and logical link between the construct and the variable. Ideally there will be general agreement among the evaluation team and the primary stakeholders that the variable is a satisfactory representation of the guiding construct.
In small-scale, low-resource evaluations, one common scenario in which this link may be open to challenge can occur when a program aimed at producing behavioral change does not have the resources to follow participants for the long term, and must end data collection concurrently with the conclusion of the program. In such cases, the staff may feel constrained to measure a short-term outcome identified in a program’s logic model and claim it as a proxy for a long-term outcome further down the outcome chain (see, e.g., Funnell & Rogers, 2011; Weiss, 2000). For example, the behavioral construct of greatest interest, healthy eating, may be replaced by a presumed psychological mediator, dietary intentions, as the evaluation’s primary outcome variable. This substitution may or may not be justifiable in terms of the overall goal of the evaluation, and there may or may not be reasonable alternatives given the available resources. But program staff, evaluators, and other stakeholders need to be aware of the ways that the substitution weakens the ability to draw conclusions about the program. In this example, the program might produce impressive increases in behavioral intentions, but the original target construct—eating behavior—will remain unexamined.
The Choice of Approach
The options for a measurement approach will vary depending on the program context and the nature of the variable. For variables that reflect individual behaviors, the available approaches might include self-report by the participant, direct observation by the evaluation team, reports by others (e.g., parents or teachers on behalf of children), and physiological or biochemical measures, among others. For variables that reflect participants’ skills and aptitudes, the options might include direct tests, self-assessment by the individual, assessments by others, or existing indicators (e.g., standardized test scores). For variables that reflect internal psychological states (e.g., motivation, interest in a topic, satisfaction with a service), there may not be viable alternatives to self-report.
To measure variables that reflect presumably objective facts or status, such as body mass index, weight, or grade in math class, there might seem to be little choice of appropriate measures. But even in these cases, decisions must be made that will affect the reliability and accuracy of the results. In the case of weight, for example, evaluation planners will need to decide how carefully to control the time of day of the measurements, and whether the weight should be self-reported or recorded by a member of the evaluation team. In the case of grades, they may decide to ask the student or may seek more direct access to grading information.
A convergent validity study by Sussman and Stacy (1994) provides a valuable illustration of the divergences that can exist across different measurement strategies for the same conceptual variable. They compared alternative strategies to derive estimates of two institutional-level variables: school-level student cigarette use and school-level student alcohol use. For cigarettes, the five strategies were students’ self-reports of their own use, students’ estimates of school prevalence, staff estimates of school prevalence, structured direct observation of student behavior at locations near the school, and refuse evidence observed at selected sites such as parking lots and sports fields. For alcohol, they used the same strategies except direct observation. (The strategies of student self-reports, student prevalence estimates, and staff prevalence estimates entailed interviews with small samples of students and staff.) The researchers found the correlations between these measures to be strikingly divergent, ranging between .01 and .54 for cigarette use and between .09 and .63 for alcohol use. One is impressed not only by the lack of correlation in some of these pairings but also by the moderate limits on the high end: none of the correlations exceeded .63, indicating that the two most strongly related strategies (student and staff estimates of alcohol use prevalence) shared only 40% of their variance. Considering that these strategies were all intended to assess the same variables, this example illustrates the dramatic implications that the choice of measurement strategy can have on the apparent levels of an evaluation’s outcome variables.
The use of multiple approaches to supplement self-report
Because all measurement approaches have characteristic limitations, plans that combine strategies will provide a stronger evidence base. This is particularly true for self-report, which is almost ubiquitous in local evaluations. Self-report information is easy and fast to collect, and it can address states of mind and private behaviors that may be impossible to obtain in other ways, certainly not with the same degree of efficiency. However, self-report also has numerous well-known limitations and problems, most importantly the susceptibility to bias due to errors of memory, misperception, or intentional misleading (e.g., Donaldson & Grant-Vallone, 2002). Therefore, when possible, measurement strategies that buttress self-report with another method to assess the variable of interest, such as biological or physiological measurement (when relevant), will considerably increase validity.
The Selection or Development of Instruments
Finally, when the variable and measurement approach have been decided, the evaluation planning team must decide the structure, format, and (if appropriate) wording of the instrumentation. Recommendations about instrument design are beyond the scope of this article, but a few examples can illustrate the decisions to be considered:
The use of multiple measures and triangulation
Any given measure will have limitations, which might include the kind of populations it can be used with, its reliability (as measured by internal consistency, interrater consistency, consistency over repeated administrations, etc.), its dependence on human memory or other fallible sources, the specificity of information it provides, and so on. The use of more than one measure can help neutralize the limitations or weaknesses of any single measure.
Multiple-item versus single-item measures
Some variables, such as age, gender, and other demographics, can be addressed completely through the use of a single item. For more complex variables, including many kinds of behaviors, the necessary number of items within a measure will be less clear. It can be tempting to cover the information of interest with a single item (e.g., How much do you exercise each week?), but single items are susceptible to significant problems of reliability. The use of multiple items, if appropriate for the variable under consideration, generally increases reliability by increasing the sampling of both content and respondent performances (Cronbach, 1990). However, this also increases the time required for response, so if a questionnaire covers a great deal of ground, this is a trade-off that needs to be weighed.
The Adequacy of Measures
The evaluation planning team needs to consider an array of factors when deciding how the program’s constructs will be represented through specific measures. Two strategies can be immediately highlighted as paths to avoid. The first is simply to select the easiest measurement option—the one that minimizes planning time, testing time, program resources, and due diligence—and hope for the best when the study is done, in terms of its usefulness and influence. The second strategy, which sometimes occurs when local programs work jointly with an outside evaluator, involves the evaluator making the decisions on what approach is required—based on his or her methodological training, desire for a publishable study, or other reasons—and expecting program staff and participants to conform. The first strategy runs the risk of completing an irrelevant evaluation study, whereas the second risks implementation challenges and the breakdown of the plan.
Ideal and Adequate Rigor: Considerations of Cost for Alternative Measurement Strategies
In cases of well-funded, high-stakes evaluation studies, there is typically consensus that the evaluation plan needs to be as rigorous as possible in order to produce strong, credible evidence and a high level of confidence for the study’s interpretations, inferences, and conclusions. This may mean, for example, very close monitoring of the program’s operation, careful attention to sample selection and retention, multiple measurement points that can track the linkages specified in the program theory, and selection of measures that are the most valid and comprehensive representations of the program theory’s constructs.
However, rigor will involve costs of various kinds, and in smaller local evaluation contexts it may be productive to ask what may be an adequate level of rigor for measurement or other aspects of the evaluation study. Chatterji (2007) argued along similar lines in discussing the strength of evidence afforded by different kinds of research designs for purposes of inferring causality. She maintained that the quality of evidence should not be judged simplistically as either strong or weak, but rather in terms of “grades of evidence” that vary along certain dimensions.
A concern for identifying criteria for acceptable levels of rigor, as well as for the strength and credibility of evidence, can result in useful standards or benchmarks as the planning proceeds. These determinations will be situated within the local context of what the program is trying to do. The primary stakeholders, and the forms of information that they need and expect, must be clarified. Program stakeholders can be drawn into preliminary discussions in which they clarify their own intended uses for the information and the strength of evidence they require—questions they may not at first be readily able to answer. The importance of engaging primary stakeholders is widely acknowledged, but matters of evaluation methodology are not always considered in these discussions.
Inadequate rigor
An obvious point, which nevertheless deserves emphasis, is that throughout the process of negotiation and decision making, care must be taken that all parts of the study are at least minimally adequate to fulfill the evaluation’s purposes and to be consistent with established standards for effective evaluation practice (Yarbrough, Shulha, Hopson, & Caruthers, 2011). Options brought to the table that are judged to have inadequate rigor, such that their prospects for producing credible evidence are unacceptably low, must be rejected, independent of other considerations. Poor evaluation practice provides no benefits, and can easily do harm to programs and their stakeholders by creating inaccurate or misleading information.
The Rigor-Feasibility Dynamic
Braverman and Arnold (2008) enumerated several categories of costs associated with increases in evaluation rigor—money and other material resources, time, and participant burden. Time can constitute a cost in multiple respects, including preparation and planning time (which may translate to money), the time available with participants during data collection, and the time horizon available for tracking change in the target groups. Participant burden can take several forms. Filling out questionnaires takes time and effort, and in some evaluation studies the questionnaires can become quite long. There may be multiple measures, as well as items that ask for revealing or sensitive information. The data collection may need to be completed at specific points in time, with which participants need to comply. If the program’s participants decide that they do not wish to tolerate these infringements, they will decline, and the evaluation will suffer accordingly.
In sum, these increased costs of rigorous methodology may lead to the judgments that some measurement strategy options are more feasible for the evaluation than others. These considerations may persuade program staff, target audiences, and others involved in the planning to reject the more demanding approaches in favor of the simpler ones.
Here is a thought experiment: If we can conceive of both the rigor and feasibility concepts as single, linear dimensions, one may imagine the relationship between them as depicted in Figure 2. This is a highly simplified representation, because there is no reason to assume that either of these concepts must be either unidimensional or linear. Nevertheless, the figure is offered to illustrate a basic conceptual point. Rigor in evaluation is typically associated with increased demand on people and/or resources. To the extent that the feasibility of any approach implies certain resource requirements and acceptance by affected audiences, feasibility and rigor are probably negatively related, more often than not, so that high levels of one will be associated with low levels of the other. As Figure 2 illustrates, an overemphasis on one of these dimensions may result in the other being so low that the evaluation becomes unacceptable or unworkable. (The existence of a discrete point signifying the change from acceptability to unacceptability is another oversimplification for purposes of illustration.) Thus, if rigor is too low, the level of confidence will be so compromised as to render the evaluation not worth doing. On the other hand, if feasibility is too low, there will be uncertainty about whether the evaluation can be completed.

Hypothesized (and simplified) conceptual representation of the relationship between evaluation rigor and implementation feasibility.
This leaves a range of options between the two thresholds, in which both rigor and feasibility may be relatively higher or lower, in opposite relation to each other, but nevertheless acceptable. This intermediate range, however the concepts may be operationalized in terms of individual evaluation decision options, is the territory for negotiation and for attentive, thoughtful decision making—a “sweet spot” for planning, so to speak. In my experience, this deliberative decision-making process about methodology does not take place as often as it should during evaluation planning—at least not in a way that is explicit, conscious, and inclusive of diverse stakeholder influence.
A final consideration is how the thresholds for acceptability get set. In the local program context, this also can derive from discussion and negotiation, guided by the expertise and methodological sophistication of the evaluator in advancing arguments to help parties understand the bases for credibility of evidence. Of course, as the scale, complexity, stakes, visibility, and resources for an evaluation increase, questions about acceptable rigor quickly become considerably more complicated, and, perhaps, less negotiable.
When feasibility issues are taken into consideration, it may be determined that a moderate level of rigor—that is, something less than ideal—may be acceptable for the primary stakeholders and planned uses of the evaluation evidence. In the next section, I turn to the interpersonal aspects of the evaluator’s role, particularly the processes of guiding discussion and building consensus. In some cases, it may be appropriate for the evaluator to make the case that the rigor, and the ensuing evidence base, should be stronger.
Interpersonal Processes and the Evaluator’s Active Role
In arriving at an evaluation plan, the evaluator may need to engage in a range of interpersonal and group activities—negotiating, teaching, consensus building, advocating—to bring various stakeholder perspectives and priorities into focus. Ideally, the resulting evaluation plan will represent a consensus among the program’s stakeholder groups. But at a minimum, it will reflect the thoughtful weighing of concerns, resulting in an evaluation that should be maximally useful given the identified challenges, constraints, and opportunities. The process can entail several components.
Underlying this conception of the planning process is a perspective that evaluation is fundamentally a process of argument (e.g., Greene, 2011; House, 1980; Wallace, 2011). Ernest House cogently characterized this position: “Evaluation techniques are often presented as being nonargumentative, as, for example, being based on valid and reliable instruments, as employing sound statistical procedures, and so on. In fact, all statements made on the basis of an evaluation are subject to challenge and are arguable—if properly challenged. The more technical and quantitative the evaluation, the less a naïve audience will be able to challenge it, and the evaluation will appear to be more certain than it is” (House, 1980, p. 74). In this spirit, the recommendations I offer for evaluator activities can, hopefully, help to reduce stakeholders’ naiveté and increase the sophistication with which they approach questions about rigor.
Establishing the Standards for Evidence
I have participated in numerous evaluation planning meetings in which the program staff’s primary questions center on their concerns about the ease of implementing data collection. They wish to know how long various instruments may take to administer, whether the data collection might generate boredom or frustration, how much time the data collection will take away from program delivery, and so on. These are legitimate local concerns, which can justifiably be the basis for accepting some measurement options and rejecting others. However, the concerns can often obscure the longer term aims of the evaluation to produce strong and credible evidence that can support useful conclusions about the level of program success.
Thus, one of the evaluator’s tasks is to make sure that the challenges inherent in conducting the evaluation are balanced against the evaluation’s potential long-term utility and value. The planning group and the program stakeholders who are being consulted must take explicit account of the broader utilization context. The evaluator will also be able to remind the group about the parties not in the room who will nevertheless be important audiences for the study, which might include—depending on the individual case—program funders, target audience members, state- or county-level legislators, clearinghouse committees, journal editors, and so on. The evidence standards for these audiences will need to be identified or estimated, ideally through direct consultation but, at a minimum, through knowledgeable appraisal.
Patton (2008) has described procedures and exercises in which the evaluator can lead stakeholders through a futuring process, to anticipate potential findings and what the utilization response will be in each case. Issues about the credibility and strength of evidence need to be part of these discussions, the result of which will be guidance for determining how rigorous the evaluation needs to be. Thus, taking into account all of the reasons for which the evaluation is being done, the group will determine the criteria for the strength of the evidence base resulting from the evaluation, in order to support the conclusions they may wish to draw with a desired level of confidence.
This process suggests that the adequacy of evidence—in terms of its relevance, credibility, and strength (Schwandt, 2009)—lies in the eyes of the beholders, who in this case are the eventual audiences for the evaluation. This is consistent with the view that evaluation is, indeed, a process of argument. Consider the example of a smoking cessation program, for which the primary outcome of interest is postprogram smoking levels. Given an expected social desirability bias that may lead to underreporting of smoking, the planning team must decide whether participant self-report, by itself, will be a sufficiently rigorous approach to address the questions and concerns of the evaluation’s primary audiences. If those audiences consist of community members and organizations, the answer may be yes. However, if the primary audiences also include public health researchers, editors, or critics who have already expressed an interest in defunding the program, the answer may be no. In the latter case, the evaluation planning team will need to consider supplementing the self-report with biochemical validations such as saliva tests for cotinine or expired air samples for carbon monoxide (Hukkanen, Jacob, & Benowitz, 2005). The general point is that evidence cannot be judged as sufficiently or insufficiently credible in the absence of the context in which it is being used, the audiences who review it, and the purposes to which it is being put.
Weighing Alternatives
Based on the evaluator’s training, acquired knowledge, and experience, he or she will probably be in the best position to educate and advise about the strengths and shortcomings of various kinds of evidence. Consider the example presented earlier, in which the long-term outcome variable of primary interest is healthy eating behavior (e.g., 6 months postprogram), but a potential short-term proxy under consideration is dietary intentions, measured at the program’s immediate conclusion. The difference in costs between these alternatives will be considerable and will favor the shorter term option. A determination must be made with respect to the net benefit of the longer term (behavioral) option, whether it is worth those added costs, and how much credibility would be lost if only the intentions variable were used. In this case, the evaluator might need to describe the relationship that has been found previously between intentions and behaviors in various domains. In general, the relationship is positive and significant, but certainly limited (e.g., Fishbein, 2008; Webb & Sheeran, 2006). More specifically, the findings have been mixed for the relationship between intentions and behaviors in the specific area of eating behavior (e.g., Lohse, Wall, & Gromis, 2011; Shaikh, Yaroch, Nebeling, Yeh, & Resnicow, 2008; Zoellner, Estabrooks, Davy, Chen, & You, 2012). Depending on the audiences for the evaluation and the discussions about this research base, the use of intentions as an outcome variable may or may not be sufficiently convincing to enable the evaluation study to achieve its intended purposes.
The evaluator can also lead the planning team through the process of identifying and weighing uncertainties regarding the consequences of alternative decision options. For example, the benefits for data analysis inherent in using a more comprehensive and detailed set of measures may be offset by issues of sample representativeness if substantial numbers of participants refuse to participate in the evaluation. However, the effects of each of the alternatives on participation rates may not be easy to predict beforehand. If a high participant-demand measurement strategy results in lowering participation by, say, 20–30% compared to less demanding options—for example, smokers declining to participate in a study that includes saliva tests compared to only self-report—the study’s rigor, as a whole, will not be enhanced by the ostensibly more rigorous option. The variety of consequences associated with each option cannot be predicted with precision, thus introducing degrees of uncertainty into the mix. Examining the experiences of prior evaluation studies will be useful in reaching reasoned predictions, and the evaluator is probably in the best position to identify these precedents and assess their relevance to the context at hand.
Negotiating the Selection of Measures
Bringing the planning team to a set of final decisions about which measures to use in the evaluation will incorporate judgments about the most critical variables for capturing the program theory’s central constructs, the most appropriate measurement strategies, the required strength of evidence, and the anticipated feasibility of the approach. In Blalock’s (1982) terminology, these decisions will reflect the auxiliary measurement theory that supports the primary program theory driving the evaluation. If there is an initial consensus about the measures, the team will be fortunate. If not, the group will need to resolve its differences in some way that is reasonably acceptable to all. The evaluator will play a key role in this consensus-building process, but may or may not be the individual who leads it.
Consider again the scenario mentioned earlier in this section, in which the program staff’s immediate focus is the ease of implementing the data collection procedures. It may be that the evaluator believes that more rigorous measurement is needed, compared to what the staff and/or other members of the planning team are advocating. The evaluator can remind the group that if they are fortunate enough to wind up with an evaluation that shows program impact, the demonstrable rigor of the study’s plan and its implementation will be of tremendous political benefit for the program. It may be that the feasibility concerns still carry the day. In that case, the decision will at least have been well considered, and the long-term consequences, in terms of relative reductions in the validity or persuasiveness of the study’s conclusions, will presumably have been anticipated.
Explaining the Measurement Requests to Program Participants
When describing the measurement requirements of impact evaluations to program participants or, in the case of children, to their parents, it may be helpful to stress to them the value of the information that is being sought about the program and/or the social condition being addressed. To the extent that the participants whom it serves value the program, an argument that the program needs the opportunity to convincingly demonstrate its effectiveness may be persuasive.
A colleague of mine, Dr. Katherine Gunter, studies bone health in children and has developed physical activity programs that emphasize high-impact exercises such as jumping and running. These exercises, when performed with sufficient frequency, produce long-lasting benefits in bone density and mineral content in the hip, spine, and other skeletal areas (Gunter & Kasianchuk, 2011; Gunter et al., 2008). In evaluations of these programs, the measurement of the primary outcome variable of bone mass entails the use of dual-energy X-ray absorptiometry (DXA), a scan that requires an extremely low dose of radiation, approximately the amount that an individual would receive from background environmental radiation exposure in a normal day’s activity. In the course of program evaluations over several years, Dr. Gunter has spoken to many parents of the children who participate in the programs. She explains the procedure, in accord with her university’s Institutional Review Board (IRB) protocols, and the parents make an informed decision about whether to let their children undergo the DXA scans. The great majority of parents and children do provide their consent, but, in many cases, only after considerable discussion to address the parents’ concerns and questions.
This example illustrates a case in which the presentation of the evaluation’s measurement procedures to prospective participants requires significant knowledge and expertise. In the case of DXA scans, there are no alternative procedures that can address the primary outcome variable of bone mass. The evaluator, researcher, or other individuals charged with this task must communicate with honesty and objectivity so that participants can reach an informed decision.
In some program settings, the ethical issues involved in discussions with participants will be particularly challenging. Mitchell, Nakamanya, Kamali, and Whitworth (2002) report on a process evaluation they conducted to examine the local acceptability of a randomized trial, implemented in several dozen villages in rural Uganda, of an intervention aimed at reducing HIV/AIDS prevalence through education and the improved management of other sexually transmitted diseases. The outcome measures for the randomized controlled trial included, in addition to psychometric and behavioral variables assessed through questionnaires, the collection of blood samples to assess HIV prevalence. Mitchell et al. describe the intensive efforts that were needed to dispel rumors and misinformation about the serological survey. One widespread rumor was that the blood was being collected, not for its stated purpose, but for resale abroad. In addition, they note, “There was evidence that enthusiastic supporters of the survey within the community sometimes used the promise of an impending cure to cajole neighbours and peers into providing blood” (p. 1085). With continuing communication efforts as the study progressed, the rumors and misstatements subsided over time. Nevertheless, faced with these substantial problems, the question arises as to whether the blood collection was necessary or advisable for ensuring the validity of the study’s findings. The authors believe that it was: “One might argue that alternative methods of measuring HIV (such as urine or saliva) may have engendered less controversy. The drawback of these alternative methods is that they do not allow for so many tests for other STD’s” (which constituted an important component of their program theory). “Furthermore, earlier pilot studies conducted in the study population found neither urine nor saliva to be any more acceptable than blood” (p. 1088). This example dramatically illustrates the difficulties that can exist in explaining to participants the need for a particular choice of measures.
The utilization of a measurement protocol that entails high participant burden is frequently addressed by offering compensation to participants in the form of cash or other incentives. Compensation can increase participation in evaluation studies and communicate to participants that their time and good will are valued. Federal research regulations provide latitude for IRBs to set their own standards about compensation, but they direct IRBs to insure against “undue inducement” (Dunn & Gordon, 2005; Permuth-Wey & Borenstein, 2009; Wertheimer & Miller, 2008), a term that is generally interpreted to mean a level of compensation so disproportionate that prospective participants might be induced to act against their own best interests. Ultimately, for any individual evaluation study, the decision about whether to compensate participants, and in what amounts, will need to be resolved in consultation with the IRB that is overseeing the project.
Defense of the Measurement Decisions
A final function that may fall to the evaluator is defending the decisions that were made during the evaluation planning. When the study has been completed, there may be skeptics, critics, or simply interested stakeholders who ask why a more rigorous approach was not taken in some area of the study, such as the choice of variable, the timing of measurement, or the technical characteristics of instruments. If the choice can be defended in terms of the balance it provides across critical planning considerations, it can make an important difference in the credibility afforded to the study and the extent to which the study is eventually used. Criticisms regarding disputable levels of rigor can be countered by identifying the significant feasibility challenges that existed. A further strength in this debate would be the opportunity (if true) to report that the measurement choices reflect the input gathered from multiple stakeholders.
Summary
In this article, I have addressed two primary issues. First, the level of rigor associated with the plan for measuring an evaluation’s outcome variables needs to be determined during the planning process. Rigor can be explicitly negotiated during the evaluation planning phase between the evaluator, program staff, program funder, and other primary stakeholders. A range of levels of acceptable rigor is possible, assuming that due consideration has been given to different measurement options and their feasibility. Rigor-related decisions should be guided by the information needs of the most important audiences for the evaluation and the study’s planned uses.
Second, the evaluator should provide leadership for a variety of critical functions with regard to planning measurement strategies and choosing or developing particular measurement instruments. The evaluator must have the requisite expertise needed to communicate with a variety of audiences, build consensus, guide the decision-making processes, and defend the evaluation plan against potential critiques.
The measurement specification tasks that I describe in this article—starting with one or more constructs and translating them into one or more measures that form the basis of data collection—are done in virtually every impact evaluation and can scarcely be avoided. However, the evaluation will benefit if this process is accomplished consciously, thoughtfully, with identification of competing options, and with broad stakeholder input. In local, small-scale program settings, these topics might be the springboard for some valuable discussions about the specific ways in which concrete planning decisions can affect the ultimate usefulness of program evaluations.
Footnotes
Acknowledgments
I thank Norman Constantine, Michael Hendricks, Roger Rennekamp, Darlene Russ-Eft, Thomas Schwandt, and Jana Kay Slater for their valuable comments on earlier versions of this article. I am also grateful to Katherine Gunter, Lisa Leventhal, and Siew Sun Wong for sharing their expertise with regard to specific sections.
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
