Abstract
Racially and ethnically diverse populations from minoritized backgrounds are often exposed to research methodologies that amplify structural racism and negate their sociocultural reality. Although cross-cultural validation of measures is considered a requisite step to multigroup comparisons, researchers apply measures validated and standardized in the dominant White culture to under-researched populations (without assessing measurement equivalence first). Seeking to align with calls for more culturally sensitive measures, an argument is made to leverage youth participatory approaches in the cross-cultural validation of study instruments. Using an illustrative case study, this paper describes how 16 youth researchers (predominantly Hispanic) partnered with an academic team to examine the validity of the Prosocial Behavior Scale. A new tool, the Key Informant Validity Index, is introduced to determine if items maintain adequate levels of content validity when applied to populations that differ from the original norming and validation study samples. Youth researchers rated items on relevance, clarity and coverage, and use a guided protocol to conduct cognitive interviews with 32 youth participants. Recognizing the inherent challenges of reducing hierarchal power dynamics within youth-adult relationships and/or alleviating insider/outsider tensions across lines of cultural difference, an intentional focus is placed on naming key strategies that facilitated the collaboration process.
Keywords
Scientific racism, or the ways in which scientific discourse has been used to sustain systemic inequality and exacerbate racial disparities, continues to maintain a strong foothold in research today (Scheurich & Young, 1997; Thomas & Sillen, 1972; Winston, 2020). Referred to by Saini (2019) as a “toxic little seed at the heart of academia,” scientific racism has hidden behind the intellectual veneer of rigorous research practices, while reinforcing White-centered narratives (p. 205). Despite decades of critiques advocating for research scientists to work with racially diverse populations, the pursuit of race-neutral universal phenomena continues to negate our understanding of racialized experiences in developmental and psychological processes (Bell & Hertz, 1976; Betancourt & López, 1993; Kline et al., 2018; Nielsen et al., 2017; Parson, 2019). In fact, a recent query of more than 26,000 empirical articles published between 1974 and 2018 revealed that only 5% of studies highlighted race, which was defined in the paper as groups conceptualized as ancestrally, phenotypically, culturally, and/or socially distinct (e.g., African American, Asian, Biracial, Hispanic, Latinx, Native American, etc.). Although the authors acknowledge that not all study phenomena vary as a function of race, the almost non-existent attention to how racialized experiences shape how people think, develop, and behave is described as “a disservice to psychological science, especially in the face of increasing racial diversity, segregation, and inequality” (Roberts et al., 2020, p. 1299).
This problem is in part driven by the tendency of empirical studies, in both standard and cross-cultural research, to operate from the false assumption of equivalence, or the notion that using identical research methods will produce equivalent and externally valid data, even across disparate cultural contexts (Kline et al., 2018). In other words, using measures originally developed and normed on a different population than the target sample, without first establishing equivalence across groups, increases the likelihood of biased results. This can include inaccurate estimates of prevalence, exaggerated projections of risks, and/or false narratives around differences in mean scores (Hall et al., 2016).
To date, researchers endeavoring to better understand the generalizability of developmental and psychological constructs typically rely on indirect methods for testing validity. For example, using statistical procedures to examine the internal structure of a measure or investigating cross-group similarity in relations to theoretically-relevant constructs (e.g., Pina et al., 2009). These psychometric approaches are a necessary, but insufficient condition for establishing cross-cultural validity, or the extent to which measures originally generated in a single culture are applicable, meaningful, and thus equivalent in another culture (Matsumoto, 2003). Establishing cross-cultural validity requires accessing local knowledge and involving non-academics who have expertise as individuals who live the research issue (Israel et al., 2001; Teixeira, 2015).
The collaborative nature and process-oriented approach of participatory methods opens the door to directly engage with local stakeholders and generate knowledge in partnership. Paving a path between science and practice, interest in participatory methods has grown rapidly in recent years, spanning multiple disciplines, and settings-from promoting community advocacy needs to enhancing school program development to supporting healthcare prevention efforts. Moreover, increasing evidence points to the effectiveness of participatory approaches when engaging underserved and research-cautious communities (Wallerstein et al., 2017).
Participatory Methods: Historical Context, Philosophical Underpinnings, and Key Principles
Situated at the intersection of positive youth development and social justice youth development (Iwasaki, 2016), scholars acknowledge that participatory methods have largely been derived from multiple paradigms including critical theory, constructivism, and pragmatism (Israel et al., 2001; Park, 1993; Reason, 2006; Reza, 2007). On an ideological front, these approaches are aimed at breaking down power structures, alleviating social injustice and removing systemic barriers to equity. On a methodological front, there is broad consensus that participatory research serves as a blend of investigation, knowledge production and action, however the process is typically iterative, and the methods remain flexible (though stakeholder participation is a standard marker; Reza, 2007).
Rooted in the same ontological paradigm and guided by the same founding principles, multiple terms are often used interchangeably to reference the methodology including: action research, participatory action research, community-based participatory research, action science, action civics, user-centered design research, etc. Broadly, “participatory” places emphasis on increasing participant voices and shifting power in the research context, while “action” aims to improve conditions and practices that are perpetuating dissatisfaction, alienation, and/or the injustices of oppression and domination (Kemmis et al., 2014). To cast the widest net and acknowledge the heavier weight placed on youth engagement in the presented case study, the umbrella term of participatory research will be used herein. More importantly, although there is still ambiguity around a system of nomenclature, the key tenets of each approach include: (1) exploring local knowledge and perceptions (Israel et al., 2001); (2) empowering stakeholders by considering them agents who can investigate their own situations (Teixeira, 2015); (3) providing a forum that can bridge across cultural differences among participants (Israel et al., 2001); (4) employing a range of methodological approaches and techniques; and (5) committing to a shift in power from the researcher to the researched.
The concept of expanding the reach of participatory research to youth is not new, however, there has been an increased endorsement of youth participatory, empowerment-centered approaches more recently (Teixeira et al., 2021). This serves as a notable departure from traditional research which has almost exclusively engaged adolescents as subjects, respondents, and informants, with minimal effort to involve them in the discovery of new knowledge (Kirshner et al., 2005). Advocates of participatory approaches suggest empowering youth to engage as partners and co-researchers can improve the research quality by both generating more reliable data and strengthening data interpretation (Powers & Tiffany, 2006). In return, the partnership with trusted adults and peers meets a number of adolescent developmental needs including cultivating supportive relationships, exercising agency, developing leadership, supporting advocacy skills, and promoting self-efficacy as change agents (Sabo, 2003; Suleiman et al., 2019). Additionally, obtaining stakeholders’ perspectives strengthens the ties between theory and research, and promotes an inclusive and culturally sensitive approach to counter the longstanding use of a deficit or pathological lens (Bennett & Cohen, 2019; Ozer & Douglas, 2013).
The Case for Youth Participatory Research in Measurement Validation
The primary purpose of this article is to explore the utility of participatory approaches in advancing measurement validation procedures. Active participation of youth has been leveraged at various stages of the research process, including: recruitment (e.g., respondent-driven sampling; Bianchi et al., 2003); data collection/analysis (e.g., designing, conducting, and analyzing peer interviews; Lile & Richards, 2018); intervention mapping (e.g., developing objectives and ideas for planned interventions; Toraif et al., 2021); and program implementation/evaluation (e.g., delivering program content, examining pre-and post-survey data; Beatriz et al., 2018). Yet despite its lauded benefits, the use of participatory approaches to improve measurement validity is an area that remains largely unexplored (Gonzalez & Trickett, 2014). Recognizing the limited guidance currently available to extend the use of such approaches to measurement validation, an illustrative case study is presented to describe the process of engaging youth in that work and to draw attention to the many benefits gained on both sides of the partnership.
To begin, a closer look is taken at sources of measurement bias and assumptions of equivalence in studies employing the same measure across distinct racial and ethnic groups. This is followed by a description of a case study leveraging an academic-youth partnership to examine the cross-cultural content validity of the Prosocial Behavior Scale. To generate estimates of content validity, a new instrument is introduced called the Key Informant Validity Index, which adds to the toolbox researchers draw from when examining the culturally specificity of study measures and/or inviting stakeholders to transition from the passive role of informant to an active role of co-researcher. Finally, toward the goal of drawing more explicit links between abstract guidelines and concrete practices, the final section focuses on lessons learned throughout the partnership and names key strategies employed.
Conceptual and Methodological Challenges in Cross-Cultural Measurement Validity
Calls have grown louder for increased precision and sampling coverage in measurement practices, as well as greater consideration of the sociocultural influences across race and ethnicity (Lee, 2018; Putnick & Bornstein, 2016). This requires addressing the often ignored reality that the majority of measures have been conceptualized around a dominant cultural frame and subjected to a universalist bias. Accordingly, methodologists have increasingly stressed the importance of testing for measurement invariance to establish psychometric equivalence and to avoid drawing erroneous conclusions (Little, 2000; Vandenberg, 2002).
Moreover, mainstream approaches have largely ignored the perspectives of youth in the measurement development process, overlooking, and undervaluing the contribution of their knowledge to promoting content and face validity. The following two subsections describe ongoing efforts to address these issues, drawing attention to where they fall short and where there may be opportunities for participatory approaches to offer alternative pathways forward.
Addressing limitations of testing measurement invariance
Although the conceptual importance of testing measurement invariance (i.e., equivalence of a construct across groups or measurement) has been recognized for over 50 years (Struening & Cohen, 1963), statistical techniques (e.g., confirmatory factor analysis, item-response theory-based methods for assessing differential item functioning) for testing invariance have only become more accessible to, and expected from, the research community relatively recently (Putnick & Bornstein, 2016). Despite the increased adoption of testing measurement invariance in developmental and psychological research, notable limitations persist. First, achieving full invariance is rare and there is limited guidance as to how to proceed when violations of measurement invariance are found (Putnick & Bornstein, 2016). Second, researchers have questioned the ability of measurement invariance to detect real, meaningful differences in constructs across groups, particularly given the ambiguity around identifying what threshold needs to be reached for a true difference to exist (Millsap, 2005; Vandenberg, 2002). Finally, testing invariance of measures takes place prior to testing mean differences or differential relations across groups but after the items of a measure have already been selected (i.e., addressing the problem of cross-cultural validity retroactively, rather than proactively).
Participatory methods offer two possible inroads to address the aforementioned concerns. Local stakeholders can be invited into the measurement development process earlier to generate items with similar meaning across different groups, or once measurement noninvariance is found, the target population can be consulted to identify potential sources of bias. The first strategy represents an a priori approach to increase the relevance and representation of a target measure to the intended construct, while the second strategy focuses on soliciting a posteriori knowledge that can help unpack unanticipated results from tests of invariance. The case study presented in the current article demonstrates an example of the latter.
Amplifying youth voice in the measurement validation process
Several adolescent self-report measures, including the measure of focus in the case study, were originally designed with a different target population in mind (e.g., elementary school-aged children, college students, and/or adults), yet measurement equivalence across age groups is seldom addressed (National Research Council, 2011). Additionally, when adaptations are constructed to accommodate a sample in single-study use, limited information is typically offered as to how the necessary modifications (e.g., rewording of items, adding/eliminating items, etc.) were determined or applied (e.g., Guevara et al., 2015; Mestre et al., 2017). This is further complicated when considering the multiple terms, often used interchangeably to conceptualize developmental stages, including adolescence (e.g., “youth,” “young adults”; Sawyer et al., 2018). Even in the current paper, the reference to “youth” in the background and literature review is more broadly alluding to the time between childhood and adulthood. However, the case study that follows focuses more narrowly on a particular group of “youth researchers” aged 13 to 15. And in cases where a particular developmental stage (e.g., early vs. middle vs. late adolescence) or age range is specified, there are still several sources of maturational heterogeneity in physical, cognitive, and psychosocial development during adolescence that are not reflected by age (Sanders, 2013). This underscores the importance of confirming measurement invariance across age groups and highlights the potential utility of engaging youth in the measurement validation process to ensure measures operate the same way across age groups of interest.
An additional challenge in engaging youth as co-researchers is overcoming the longstanding perception of children and adolescents as objects of concern, rather than persons with voice (McLaughlin, 2015). Although this issue was raised almost 30 years ago, the literature on adolescent development has continued to be largely shaped by adult frameworks employing a deductive conceptual approach (Teixeira et al., 2021; Wong et al., 2010). Moreover, the bulk of findings are drawn from studies focused on mainstream adolescents, with far less consideration given to how these developmental changes may unfold differently in marginalized groups. Thus, many current measures are White normative and adult-centric, that is, largely constructed through a narrow White adult lens, with the perspectives and real-life experiences of diverse youth populations omitted (Bennett et al., 2003; Daiute & Fine, 2003).
To broaden the interpretive lens, participatory approaches create a structure that centers around validating the knowledge of the target population (in this case historically marginalized youth) and places existing instruments or those in development under the necessary scrutiny (Schilling et al., 2007). By inviting adolescent experiential experts into the measurement validation process, everyday conceptions of the target construct can be explored, balancing the influence of adult “expert definitions” with adolescent “lay usage” (e.g., Delle Fave et al., 2016). Furthermore, the shift in traditional roles within the research process also helps to replace risk-saturated and deficit-based discussions with more adequate explorations of the cultural assets and inherent strengths of non-White populations.
In sum, measurement instruments are required to be culturally sensitive, and reliability and validity should be established in the group for which the measure is intended (American Psychological Association, 2002). To accomplish this, measures can be psychometrically tested in the population of interest, however, this alone cannot determine its age and cultural equivalence across groups. With this in mind, the current paper explores the use of participatory approaches to facilitate the validation of measures to currently under-researched populations in hopes of producing more comprehensive, broadly applicable results in developmental research and practice.
Youth Participatory Approaches in Measurement Validation: Overview of a Case Study
The current case study employs a youth-focused and strengths-oriented approach to the engagement and empowerment of high-risk, marginalized adolescents. Sixteen seventh and eighth grade Hispanic students (predominantly Puerto Rican) partnered with a university research team to evaluate the cross-cultural validity of the Prosocial Behavior Scale. Prosocial behavior, broadly defined as overt actions intended to benefit others (Batson & Powell, 2003), is open to numerous definitions, conceptualizations, and methodological approaches. However, the underlying assumption remains that it is a socially contingent, culturally-anchored construct that changes over time, both in terms of individual life course changes as well as changes in sociocultural context. As such, a participatory approach was applied to take a closer look at one of the construct’s commonly used measures. The Prosocial Behavior Scale is comprised of nine statements rated on a 1 to 3 scale: 1 = never, 2 = sometimes, or 3 = often (e.g., “I try to make people happier when they are sad”; see Appendix A for full instrument). The validity and reliability of this scale has mostly been demonstrated in European samples (e.g., Caprara et al., 2005; Caprara & Pastorelli, 1993) and therefore the current study sought to determine whether the scale provides sufficiently equivalent measurement across individuals of different racial and ethnic backgrounds, namely Hispanic youth of Puerto Rican and Dominican origin.
Case Study Site
The study school district is located in New England and serves a high concentration of under-resourced students in a city that includes a large population of working-class immigrants. About 87% of families live in poverty and about a quarter of the district’s 5,500 students are homeless or in temporary housing. The city’s population is 18.8% college educated and demographics indicate inhabitants are 44.7% Hispanic, 41.9% White, and 3.27% Black or African American. The study school site, however, serves a predominantly Hispanic student body (89%), with approximately 67% of students identifying as Puerto Rican, 17% identifying as Dominican, and 3% identifying as bi-racial.
Historical Context of University-School Partnership
In 2015, the study district was placed under receivership as a result of increasing out-of-school suspensions (more than five times higher than the state), chronic absenteeism, and non-rising graduation rates. During 2017, in response to the district’s announcement of planned efforts to re-engage disconnected or at-risk youth, a partnership was established between a diverse team of five university-based researchers (identifying as Black, Hispanic, Middle Eastern, White, and bi-racial) and one of the middle schools in the district. Although the university-school partnership had collaborated on several research endeavors by the time the case study launched in 2018, there had been no previous attempts to actively involve students in the planning or execution of data collection, analysis, or interpretation. The idea for doing so largely emerged from ongoing conversations with school leadership and staff regarding the observable disinterest and irritability students displayed when asked to complete surveys (which was a far more frequent ask since the district had been placed under receivership).
Rationale for Youth-Adult Research Collaboration
The decision to provide middle school students with a more active and meaningful role behind the scenes of survey administration was primarily motivated by four factors. First and foremost, school leadership and staff were seeking to increase transparency around why data was collected and generate more trust and confidence that data could (and would) ultimately be used to bring about positive change (something students had indicated was notably absent in their past experiences with surveys and therefore repeatedly referred to them as “useless” or a “waste of time”). Second, the new partnership approach coincided with an equally pressing goal of generating unique leadership opportunities for students who teachers described as “having leadership potential with no outlet to exercise it.” Third, school leaders were interested in exploring whether the lack of expected findings in previous administrations of student surveys could be attributed (in part) to issues of survey validity and wondered if consulting directly with students would help to align “expert” and “lay” conceptualizations of the target constructs being measured. Finally, a previously conducted needs assessment involving school leadership, staff, and students had revealed longstanding frustration with the lack of cultural humility often practiced by outsiders involved in the school’s evaluation efforts and a deep mistrust in research more generally. Leveraging participatory methods with students was therefore intended to draw more strategically from local context and cultural values; to reject universalistic or deficit approaches, and to take an initial step in countering the preconceived hierarchy between researchers and those being researched.
On the flip side, by engaging students in the examination of surveys prior to their administration, the university researchers (also referred to as academic researchers) stood to gain valuable insight regarding potential sources of bias in the wording of measurement items. Thus, forming an academic-youth partnership was driven by pragmatic goals (improving quality of student data collection efforts), responsive to local needs (addressing frustrations at the student-, teacher-, and administrative-levels), and ethically responsible (engaging marginalized youth in the decision-making of a process that directly impacts them).
Engaging Youth Research Partners in Measurement Validation
At the time that the case study was initiated, the academic-youth partnership was considered the first of a series of steps to establish a long-term, systematic process for student voice to be amplified when selecting, adapting, and/or validating measures. As such, the scope of the first attempt to do so was limited to a single instrument-youth research partners (also referred to more simply as the “youth researchers”) were tasked with choosing one out of eight surveys administered bi-annually to students in grades six through eight. These surveys were chosen by school leaders in previous years to evaluate ongoing efforts to promote social and emotional learning, as well as assess student perceptions of belonging and positive school climate.
Once the youth researchers had selected the Prosocial Behavior Scale through majority vote, collaborative working sessions got underway, taking place once a week for 1.5 hours (see Appendix D for an overview of the partnership process and Appendix E for a detailed breakdown of the 12 weekly sessions, both documents are adapted for a youth audience). Youth researchers received training in four main areas: measurement validation, cognitive interviewing, qualitative analysis, and research dissemination. The training sessions included balancing larger group discussions with smaller group activities, leveraging collaborative brainstorming approaches (e.g., concept mapping); using culturally relevant everyday examples (e.g., referring to a recent headline or viral video); diversifying the format of content delivery (e.g., videos, podcasts, and gamified lessons); opening up multiple feedback channels (online polls, exit slips, and anonymous question box) and introducing each lesson with an unexpected “introductory hook” to spark engagement and remind the youth of the significant purpose behind their work (e.g., a surprising finding from a research study showing cultural generality, a personal anecdote demonstrating a “research fail” from one of the academic team members, etc.). This last strategy was particularly important given the youths’ initial hesitation to view themselves as difference-making researchers. Exposing science’s role in normalizing hierarchies between racial groups and reiterating that despite decades of work, the adults may not know all (or any) of the answers to the problems that they too were grappling with, reinforced the importance of having youth voices heard and their knowledge affirmed.
Youth Research Partners
A total of 16 seventh-and eighth-grade students aged 13 to 15 partnered with the university research team. The youth researchers were recruited through flyers, information sessions, and emailed invitations (see Appendix B for a sample recruitment flyer). Thirty-three students submitted forms of interest to participate and those selected for the current case study were intended to reflect a representative demographic breakdown of the Hispanic population within the school, including students who identified as Puerto Rican, Dominican, and biracial, all of whom were bilingual (Spanish and English). About 8 of the 16 youth researchers identified as female, 31% reported unstable housing situations and all students were eligible for free or reduced lunch (see Table 1 for additional demographic information on both the university and youth research team members).
Demographic Characteristics and Attendance of Youth Researcher Partners (N = 16) and Youth Research Participants (N = 32).
Youth Research Participants
An additional 32 seventh and eighth grade students (aged 13–15) participated in cognitive interviews conducted by the youth researchers. About 16 of the 32 participants identified as female, 75% identified as Puerto Rican, 42% reported unstable housing situations, and all were eligible for free or reduced lunch. Students were recruited primarily through word of mouth, although flyers and classroom announcements were also used (see Table 1 for additional demographic information).
Methods
Content validation is a multimethod process and has previously been estimated qualitatively, quantitatively, or using a combination of both methods to determine a degree of consensus among experts about an item or measure in question (Haynes et al., 1995). Here, quantitative estimates of relevance, clarity, and coverage were calculated first using a newly developed measure, the Key Informant Validity Index (see Appendix C). Then, the youth researchers were trained to conduct cognitive interviews with their peers and qualitative accounts obtained from the process were used to determine how well the survey items were understood and what factors led to variability in responses. This two-pronged approach allowed for potentially problematic items to be flagged during the quantitative strand, while follow-up data collected in the qualitative strand offered a clearer understanding of “why” the flagged items may not function as intended.
Pilot testing the Key Informant Validity Index
The Key Informant Validity Index (KIVI) builds on the previously established Content Validity Index (CVI; Polit & Beck, 2006) to provide a more comprehensive approach of examining validity during the instrument development or adaptation. Key aspects of the original tool are maintained (e.g., overall structure, use of Likert-type, and ordinal scale scoring procedures), however, the KIVI also employs a mixed method validation process which is a departure from the CVI’s use of a quantitative-only approach to examine a single item on relevance. Designed as a two-stage process, an index of inter-rater agreement is calculated first, followed by extracting themes from cognitive interviews. Of note, the items and interview questions included in the KIVI were developed based on a literature review conducted prior to the start of the case study.
To begin, the KIVI asks key informants (in this case study, 16 Hispanic adolescents), to rate each item on a 4-point scale. A separate CVI value is computed for relevance, clarity, and coverage (e.g., response options for relevance include 1 = not relevant, 2 = somewhat relevant, 3 = quite relevant, and 4 = very relevant). The item-level CVI (I-CVI) represents the total number of key informants that gave a rating of either 3 or 4, divided by the total number of key informants providing ratings—in other words, the I-CVI represents the proportion in agreement for each item. For example, if seven key informants provide ratings, and five of them rate an item as 3 or 4, that particular item’s I-CVI would be 0.71 (or 5/7).
Similarly, the scale-level CVI (S-CVI) determines the proportion of items on a scale that achieved a rating of 3 or 4 by all key informants (e.g., a score of 0.70 on a 10-item scale indicates that 7 of the 10 items received a rating of 3 or 4 by all of the key informants while 3 items did not reach universal agreement). If the item-CVI or scale-CVI for a particular dimension is at or above “acceptable” cut-off values, then the item/scale is considered to have achieved a satisfactory level of content validity (see cut-off values in Table 2; Lynn, 1986; Polit & Beck, 2006; Yusoff, 2019). The recommended number of informants to review an instrument varies from 2 to 20 individuals, but to have sufficient control over chance agreement, at least five people are suggested to review the instrument (Zamanzadeh et al., 2015). For five or fewer informants, the I-CVI must be 1.00—that is, all informants must agree that the item is content valid. However, with the addition of more informants, room is made for a modest amount of disagreement (e.g., when there are six informants, the I-CVI must be at least .83, reflecting one disagreement, when there are nine, the I-CVI must be at least .78).
Item Content Validity Index Scores Calculated from Youth Researcher Expert Ratings on the Key Informant Validity Index.
Note. I-CVI scores reflect the proportion of experts who agree on the relevance, clarity, and coverage of an item. Bolded statements reflect items flagged as problematic (I-CVI < 0.78).
Reverse coded item.
Cognitive interviewing
The KIVI also includes a cognitive interview protocol with follow-up probe questions that can be used to investigate problematic items (items that did not receive a 3 or 4 on relevance, clarity, and/or coverage). Cognitive interviewing is an evidence-based qualitative method specifically designed to determine whether a survey question satisfies its intended purpose by more closely examining how respondents comprehend, interpret, and respond to survey items (Willis, 2004).
During the collaborative work sections, the academic team employed reciprocal teaching strategies (“I do, we do, you do”), first modeling the semi-structured interview protocol, then role-playing the interviewee for the youth researchers and eventually creating space for the youth researchers to offer one another feedback as they honed their interviewing skills. All interviews were transcribed and responses to each item were chartered into a spreadsheet (items on the horizontal axis, participants on the vertical axis). A traffic light system was used to distinguish between comments that called attention to a problem with an item (highlighted in red), comments that named an item strength (highlighted in green), and comments that came across as neutral or ambiguous (highlighted in yellow). This served as a starting point in organizing items around broad categories. Next, thematic analysis was used to further assess relevance of the item (based on the extent to which the item was viewed as important to the measurement of prosocial behavior), clarity of the item (based on respondent’s comprehension of the question), and coverage of the measure (based on the extent to which the Prosocial Behavior Scale was perceived as comprehensive).
Results
The item CVI scores for relevance and clarity were calculated for each item of the Prosocial Behavior Scale using the KIVI rating scale. Four statements were flagged as problematic (I-CVI < 0.78; see Table 2) and placed under further scrutiny using the KIVI interview protocol.
Item 1: I Try to Make People Happier When They Are Sad
The phrase “cheer up” appeared to elicit mixed reactions. Concern was raised that the act of trying to “cheer up” someone may actually create a sense of emotional dismissal and perhaps unintentionally invalidate the recipient’s circumstances or exacerbate feelings of distress. In two interviews, participants shared personal stories of how attempts to uplift someone in a distressed state was a particularly negative experience when the recipient did not believe their circumstances were within their control. Although there was consensus that successfully making someone feel better would be perceived as prosocial, participants suggested the state of happiness may not be the right end goal. More specifically, participants noted that sometimes individuals experiencing hardship want and/or need to process the negative experience and are often seeking companionship or support while doing so, without the expectation or desire of a positive mood shift. Overall, there did not seem to be any confusion regarding the intent of this item, therefore the item was considered clear, however further examination of the item is needed given the word choice may not align with observed norms for providing emotional support.
Item 3: When I Have to do Things I Don’t Like, I Get Mad
The negatively worded and reverse-coded structure of this item created a lot of confusion. Participants re-read the item multiple times and still struggled to understand exactly what was being asked (often interpreting it as getting angry when asked to do things you don’t want to do). There also appeared to be agreement across interview participants that although displaying emotional regulation was valued in their interactions with peers, this item was not necessarily capturing that. The more relevant examples that came to mind were almost always situated in times of conflict. With the lack of consensus on what the statement was conveying and the difficulty drawing connections to the broader construct of prosocial behavior, this item was considered neither clear, nor relevant.
Item 8: I Like to Play With Others
While male participants found this item to be both clear and relevant (often referring to examples of “playing sports” or “playing video games”), female participants took issue with the word “play,” suggesting it was not developmentally appropriate for this age group (e.g., “I’m not a child. I don’t ‘play’ with my friends. I hang out or chill with them”). There was also concern expressed regarding the extent to which “playing” (i.e., “hanging out”) was an important criterion to consider in the measurement of prosocial behavior (particularly as compared to the other items which seemed to be referencing higher cost prosocial behaviors). With the consistently strong reaction to the word “play” and the lack of consensus regarding its importance, this item was not considered clear or relevant.
Item 9: I Trust Others
Second only to the reverse-coded item (#3), this statement produced the most confusion among participants, with the majority interpreting the statement as referring to how trustworthy they themselves were (as opposed to the extent to which they trust others). This led to additional commentary regarding whether or not the ability to trust others was important in measuring prosocial behavior, particularly in comparison to being trustworthy yourself. With most students interpreting this statement backwards and the difficulty reaching consensus regarding its importance, this item was considered not clear and not relevant.
In summary, the above findings suggested that four of the nine items on the Prosocial Behavior Scale may be vulnerable to varied interpretations among Hispanic youth and particular word choices may be misaligned with cultural, developmental, and/or gender norms (e.g., the use of the word “happy” when providing social support, the inappropriateness of the word “play” and the directionality of trust).
Based on these findings, school leaders and a subgroup of youth researchers chose to remove items 3 and 5 and revise the wording of items 1 and 9 (e.g., item 1 was changed to “I try to support someone when they are sad”). Four new items were also added to the measure that aligned with emergent themes from interview responses to questions targeting coverage on the KIVI. These included being funny, standing up for others, being complimentary or encouraging, and expressing gratitude. Of note, these qualitative insights suggest that contemporary measurement of prosocial behavior may be focusing on a restricted range of indicators, thus limiting our understanding of socially significant prosocial behaviors among diverse populations Scale up, Spread, and Sustainability of Participatory Approaches at Study Site
As previously mentioned, this case study was intended to map out the necessary processes, set up foundational structures, and troubleshoot the inevitable challenges associated with launching participatory approaches in a school setting. In keeping with that broader objective, a concerted effort was also made to document as much of the collaboration as possible (e.g., archiving agendas for each of the working sessions, generating a resource bank for all the youth-facing research method training materials, storing video recordings of both the cognitive interviewing training process, and sample interviews conducted by youth researchers with student participants).
Additionally, rather than waiting until the end of the collaboration to deliver all findings in one lump sum, the research team decided that every 2 weeks (a total of six times during the 12 weeks of collaborative working sessions), three youth researchers and one member from the university team would work on generating a “research bite” that shared how emerging knowledge had already or could potentially serve local needs. The content of the research bites was both process-oriented (e.g., new reflections on the academic-youth partnership process) and outcome-oriented (e.g., “three questions have been identified as potentially problematic on the Prosocial Behavior Scale using a tool that “grades” each one on how easy it is to understand and how important it is to ask the question”). Typically, research bites were one-page, visually appealing documents with information organized in short, clear bulleted statements.
To support continuous improvement efforts, debrief discussions were held after each collaborative session and group exit interviews were conducted at the end of the study. The qualitative questions included in the exit interviews focused on the youth researchers’ perceptions of acceptability (how they experienced the group, what were their favorite research activities, whether level of training was appropriate, how future collaborations could be improved), and their impressions of the new approach to involving students in survey selection and development. The youth researchers described feeling heard in a way that had been notably absent in their student experiences to date. In particular, they appreciated the quick turnaround from providing feedback during a collaborative working session and immediately seeing how it would be implemented into the survey edits. The also described (re)discovering their agency as they engaged in group decision-making (e.g., choosing which survey to focus on, deciding which research activities to take lead on vs. hold a supportive role), noting that it was a welcome change from being told final decisions after the fact. This contributed to a broader sense of serving as change agents for their school community. One youth researcher stated, “It was pretty cool to take the survey with everyone else but feel like we had a behind the scenes look at what went into it. I asked my friend if he thought it was any better than ones we have taken before, and he said he definitely felt like the questions just made more sense than they usually do. that made me feel like we had done a good job.” The most common benefit discussed among the youth researchers during the exit interview was the opportunity to get to know their peers better through meaningful discussions and shared goals. Even youth researchers who acknowledged struggling initially to engage in the work described how much easier it became across the sessions to begin speaking up with their ideas as they increasingly felt affirmed and supported by one another. For example, one youth researcher described the collaborative working sessions as “encouraging and non-judgmental,” while another described the work itself as “important to do together.”
Once the “pilot” case study had wrapped up, the university research team offered additional training and access to all resources to strengthen the school’s capacity to conduct their own participatory work with youth participants and without university researchers. In the year that followed, four teams of teachers (3–4 members each) led their own adult-youth collaborations with small groups of middle school students (4–7 members): two of the research teams expanded on the work conducted in the case study to examine the cross-cultural validity of additional survey measures and two explored the application of participatory approaches in revising curricula.
Discussion
Participatory approaches provide a more structured pathway to ensure marginalized communities, including students of color, are actively engaged in research from the early stages of conceptualization through implementation and dissemination (Teixeira et al., 2021). The current paper describes a collaborative process between academic and youth researchers to ensure the cultural specificity of a survey instrument developed to measure prosocial behavior. This collaboration was embedded within a larger university-school partnership seeking to improve the quality of data collected, to engage their historically underserved and underrepresented students in more inclusive practices, and to establish an ongoing process for students to clarify, reflect upon, and inform the purpose and content of survey data collection efforts. Most importantly, the academic-youth partnership offered a “first step” in overcoming the strong distrust of research held by a school community that had traditionally been the “subject” of research without reaping its benefits. The youth researchers were instrumental in steering the university team away from interpretations driven by dominant epistemologies and in grounding the work in a more accurate sociocultural reality, thus reducing bias, and informing future measurement endeavors. In return, the academic team equipped youth researchers with the backend support needed to promote a wide range of developmental competencies, including the acquisition of technical skills (e.g., critical and analytical thinking, information synthesis, and data communication) and the strengthening of interpersonal competencies (e.g., asking open-ended questions to probe for information in interactions, using disagreement sentence stems to express dissenting opinions, and engaging in collaborative problem-solving).
Lessons Learned and Key Strategies
Challenges persist in effective implementation of participatory designs, and it is often unclear how to initiate and maintain a partnership that is both research-focused and service-oriented. This may in part because participatory approaches refer more to a research paradigm than a methodology, and therefore, there are no set or prescribed methods to draw on. Instead, the methods are determined by the research context and limited guidance is available regarding application of the theoretical principles.
In an effort to bridge this “know-do” gap and to draw stronger ties between abstract recommendations and the actionable practices, the following sections take a closer look at some of the lessons learned from this case study, and name key strategies that were used to co-create knowledge and support an iterative co-learning process.
Lesson 1: Don’t “Overtrain” Youth Researchers
One thing that proved to be more challenging for the academic researchers was recognizing that the goal was not to make our young collaborators ‘experts' in a research area, but rather, to provide them with enough information to facilitate their meaningful contribution to the project. Thompson et al. (2012) warn about this risk of overtraining or “professionalizing” members of participatory groups and ultimately hindering their ability to represent their own community. This was most evident when the youth researchers were preparing to conduct the cognitive interviews. Fearing they may make a mistake or create a problem, youth researchers were reluctant to move away from the scripted language provided in the semi-structured interview protocol (which was intended to be more of a “suggestive guide,” rather than prescriptive). As the youth researchers became more comfortable with the interview process, they began using personalized language and adopting a more conversational approach when asking questions. This in turn, put the youth research participants more at ease, and both the range and quality of responses from their peers drastically improved.
Along similar lines, it is important to simplify research concepts so that they are youth-friendly and jargon-free. Reflecting on the 12 collaborative sessions, it is immediately obvious which research concepts were broken down and communicated more effectively than others. Appendix F includes a sample handout used during the training process to explain the different types of validity in youth-friendly language. As the academic and youth researchers worked together to generate new materials, it also became clear that the line is thin between what is considered “youth-friendly” and what is perceived as “condescending” or “childish.” To take a more proactive approach, new materials were previewed with at least one youth researcher prior to sharing with the larger group to solicit initial input on the difficulty of the terminology, the relevance of the examples, and/or the clarity of the messaging. Another strategy that was found to be helpful was encouraging the youth researchers to coin their own terms for the research concepts. For example, most of them referred to quantitative and qualitative analysis as “numbers” analysis and “word” analysis; and although this may not be a technically accurate, it helped the youth researchers more quickly distinguish between the two approaches during discussions about the KIVI.
Lesson 2: “Word of Mouth” Can be a Powerful Tool to Change the Narrative
In the first collaborative session with students, academic researchers asked why they thought survey data was collected at the school and what they thought was done with the information gathered. Verbatim responses included “because they [teachers] like to torture us” and “so they [administration] can tell us how bad we are at everything we do.” School leaders and staff were aware of the negative perceptions surrounding data collection efforts and were eager to use the participatory process to be more transparent about the research agenda and to demonstrate how survey findings could lead to actionable change. What ultimately had the greatest impact in accomplishing this, however, was youth researchers informally discussing their participatory experiences with peers, having the youth researchers present details about their work in a schoolwide assembly and including a brief statement in the following administrations of the student survey that acknowledged the role the youth researchers had played in revising and developing the updated items.
Lesson 3: The Level of Youth Participation Can Change at Each Stage of the Research Process
A common misperception is that the level of participation is fixed throughout the research cycle, when in fact, each step offers an opportunity to re-evaluate how best to meet the needs of both the research and those involved in the research. Ultimately, the team may draw from highly participatory strategies during some steps and rely more heavily on research-driven strategies during others. Case and point: given the wide range of activities executed in this case study, the level of participation began at “consultation” during the research design process (selecting which survey to focus on, identifying data collection instruments) and advanced to a level of co-construction during the data collection, analysis, interpretation and dissemination phases (conducting cognitive interviews with their peers, coding interview responses, generating data-informed recommendations for survey revisions).
Concluding Remarks
Over 30 years ago, Rogler (1989) was already calling for a shift away from cultural generality and toward cultural specificity, or the continuous incorporation of the values, needs, preferences, and practices of the local community at all stages of the research process. Yet existing measurement approaches in academic research have been found to reinforce stigma and sustain power imbalances. To keep pace with the rapidly changing demography, it is becoming increasingly urgent for researchers to better understand the potential role of racial and cultural differences among population groups, how such differences may impact their research study design, analysis, and interpretation, and consequently how best to engage diverse populations in research. Youth participatory research offers one promising and fast-growing approach that can both challenge the assumption of cultural homogeneity as it relates to measurement across racial and ethnic groups, as well as amplify the voice of historically marginalized youth through meaningful engagement in the research process. The collaborative nature of participatory research represents a change in existing researcher–subject power relations that may lead to generating measurement items that are more age-appropriate, culturally responsive, and psychometrically sound.
Footnotes
Appendix A
Appendix B
Appendix C
Appendix D
Appendix E
Appendix F
Author Note
The case study described in the manuscript was conducted while the author was at the University of Massachusetts Amherst.
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The preparation of this manuscript was supported by the Institute of Education Sciences, U.S. Department of Education, through Grant R305B170002 to the University of Virginia. The opinions expressed are those of the author and do not represent views of the U.S. Department of Education.
