Abstract
The current set of three studies further evaluates the validity and application of the Psychological Research Inventory of Concepts (PRIC). In Study 1, we administered the PRIC to a sample of introductory psychology students and online (Mechanical Turk) participants along with measures assessing theoretically related concepts. We found evidence of concurrent validity by demonstrating that PRIC scores relate to education, cognitive effort, and general psychological knowledge. In Study 2, we demonstrated that advanced psychology majors score higher on the PRIC than college graduates who did not major in psychology, suggesting that the PRIC assesses domain-specific knowledge and may be useful in assessment of the undergraduate psychology major. Finally, in Study 3, we demonstrated that scores on the PRIC increase from the start to the end of a research methods course. Together, these studies provide further evidence that the PRIC may be a useful index of learning research methods and statistical knowledge important for the undergraduate psychology degree.
Our previous research (Veilleux & Chapman, 2017) reported on the development and initial validation of the Psychological Research Inventory of Concepts (PRIC), a concept inventory designed to assess research methods and statistical knowledge. Our goal was to develop an inventory that would assess learning in research methods and introductory statistics courses independent of grades, similar to concept inventories developed in other fields (Anderson, Fisher, & Norman, 2002; Evans et al., 2003; Halloun & Hestenes, 1985a, 1985b; Wallace & Bailey, 2010). We developed vignettes and used narrative responses from undergraduate participants to create realistic sounding wrong answers for multiple-choice questions. We then selected items for the PRIC using Item Response Theory, a sophisticated psychometric method for evaluating item quality. The result was a 20-item measure that can be completed with a 50-min class period and can be easily scored. Our goal for the PRIC was that it could be used as a placement exam, a tool for evaluating instruction (both at the classroom level and at the department level), and for diagnosing misconceptions (Halloun & Hestenes, 1985a, 1985b).
Before the PRIC can be used for these purposes, however, additional information about the validity of the PRIC is required. Namely, we found that PRIC scores were associated with higher educational attainment and greater standardized test scores. One concern could be that the PRIC measures intellectual ability, test-taking acumen, cognitive effort, or thinking abilities acquired through any undergraduate degree, not specific to psychology. In addition, to be viable as a teaching tool, some data on the PRIC in a classroom context are necessary, such as if PRIC scores reliably increase when taking a research methods or statistics course. We would also expect that, even if the PRIC is intended to be an assessment independent from grades, PRIC scores should correspond with grades. Students who master research methods and statistical concepts in their courses should also earn good grades in those courses, although grades also often include elements other than learning (e.g., attendance, in-class activities, and timeliness of assignments).
The three studies presented in this article were designed to address some of these remaining questions about the usefulness of the PRIC. In Study 1, we administered the PRIC to both undergraduate students and a sample from the general public along with additional measures to obtain indicators of concurrent, convergent, and discriminant validity. In Study 2, we administered the PRIC to advanced psychology majors enrolled in a capstone course and compared their performance on the measure to a sample of college graduates without training in psychology. Finally, in Study 3, we administered both a pretest and a posttest PRIC to students enrolled in a research methods course to determine if PRIC scores would increase after taking a methods course and to examine the relationship between PRIC scores and course grades.
Study 1
Method
Participants and Procedure
Data were collected from two sources: a psychology subject pool and Amazon Mechanical Turk (mTurk), an online platform where people can sign up as “workers” to perform online-based studies and other tasks for small amount of money. Subject pool participants (N = 351) signed up for the study via an online course management program (Experimetrix) and received course credit, and mTurk participants (N = 245, required to be from the United States) earned US$3 for completing the set of measures.
Of the 596 participants who completed the study, the final sample was 521 after exclusions. We excluded individuals (n = 19, all subject pool) who did not meet criterion for the Instructional Manipulation Check (see measures section below), and we excluded 56 individuals (n = 39 subject pool; n = 17 mTurk) who completed the study in less than 20 min.
All participants completed the set of measures online. The items for the PRIC were administered in random order immediately after the demographics questions. The validity items followed. The average amount of time taken to complete the set of items was 50.82 min, with no significant differences in duration between subject pool (M = 52.25, SD = 28.78) and mTurk participants (M = 49.01, SD = 22.25), t(519) = 1.41, p = .16, d = .12., 95% confidence interval [CI]: [–.05, .30].
Measures
Demographics and educational history
In addition to demographic questions about gender, age, and ethnicity, we also asked participants to select their highest level of educational attainment (some high school, high school diploma, trade school certification, some college, bachelor’s degree, or advanced degree), standardized test scores (ACT and/or SAT), and if they had ever (yes or no) taken a high school psychology course, a high school statistics course, a college psychology course, a college statistics course, or a college research methods course. Due to a coding error, we neglected to obtain college major information from mTurk participants, but we asked subject pool participants their intended or actual college major.
PRIC
This is a 20-item vignette-based multiple choice measure to assess knowledge about research methods and statistics in psychology (Veilleux & Chapman, 2017). The measure includes 16 vignettes covering topics covered in undergraduate statistics and research methods courses, including replication, operationalization of variables, reliability and validity, correlation, random assignment, experimental design, factorial designs, order effects, external validity, and generalizability. Similar to concept inventories developed for other fields, wrong answers include plausible “lay answer” foils. The measure can be completed within 30 min and is easy to score, with possible scores ranging from 0 to 20 (full measure available from authors). Due to online administration, items were given in random order, where the response options for each item were also randomized.
Belief in Science Scale
The Belief in Science Scale (Farias, Newheiser, Kahane, & deToledo, 2013) is a 10-item measure assessing the belief in and values of science. The items (e.g., “The scientific method is the only reliable path to knowledge”) are given on a 1 (strongly disagree) to 6 (strongly agree) Likert-type scale, with high scores suggesting a greater reliance on scientific information.
Cognitive Reflection Test
The Cognitive Reflection Test (CRT; Frederick, 2005) is a 3-item test that assesses “miserly” processing (Evans, 2008). The 3 items on the test all have an intuitive answer that is incorrect, and to arrive at the correct answer, the easy or intuitive response must be inhibited in favor of deeper, more effortful processing. For example, 1 item is “A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?” The intuitive answer is 10 cents, but the correct answer is 5 cents. The CRT is scored on the number of correct responses (from 0 to 3) and has been shown to predict a variety of heuristics and cognitive tasks (Frederick, 2005; Toplak, West, & Stanovich, 2011).
Grit Scale
This is a 12-item scale that assesses passion and perseverance for long term goals, such as educational success (Duckwork, Peterson, Matthews, & Kelly, 2007). The items are given on a 1 (not like me at all) to 5 (very much like me) Likert-type scale. Higher scores represent higher grit (Duckworth et al., 2007).
Instructional Manipulation Check
This was a 1-item test designed to catch participants who are not paying attention to directions. In the current study, a modified version was given (Goodman, Cryder, & Cheema, 2012; Oppenheimer, Meyvis, & Davidenko, 2009) in which participants read a paragraph about decision-making and asked “what is this study about?” Embedded in the instructions was a charge to ignore the obvious answer “research methods and statistical knowledge.” Instead, the participants were instructed to select the “other” box and type “Purple Rain” in the free-answer box below. In the current study, the Instructional Manipulation Check (IMC) was given toward the beginning (after the demographics, before the PRIC) and if failed, the survey redirected the participants to a page that told the participants they were not paying attention and to try the IMC a second time (Oppenheimer et al., 2009; Study 2).
Interest in Psychological Research
This measure (Vittengl et al., 2004) is a 5-item scale assessing self-reported interest in engaging in research activities common to undergraduates. As originally given (Vittengl et al., 2004), the scale followed titles and abstracts from American Psychological Association (APA) journal articles. Here, we took the major topics generally rather than applying them to specific articles. The measure asks about five research activities: listening to a presentation at a conference, reading a research article, working on a research team, taking a class focused on research, and conducting research, each measured on a 0 (not at all interested) to 4 (extremely interested) Likert-type scale.
Psychology Print Exposure Measure
The psychology print exposure measure (PPE; Smith & Barker, 2008) is a 100-item measure that assesses memory for psychological terminology. The measure consists of 50 real psychological terms (e.g., biofeedback and semantic loop) and 50 fake psychological terms (e.g., adolescent amnesia and antisocial facilitation), where the participant is instructed to indicate whether each term is real or not. In the current study, we presented a modified version of the PPE whereby participants were given all 100 terms in random order and asked to indicate whether (a) this is a fake psychological term, (b) this is a real psychological term but I do not remember what it means, and (c) this is a real psychological term and I know what it means. Scores were determined by counting the number of correct psychological terms the participant identified (whether they could remember the definition or not) plus the number of fake psychological terms identified as fake. This measure is a brief but effective method at assessing broad psychological knowledge, which correlates with recall of key term definitions and with traditional assessments (e.g., multiple-choice tests) in general psychology courses (see Smith & Barker, 2008).
Rational/Intuitive Scale Short Form
The short form of the Rational-Intuitive Scale (RIS; Norris, Pacini, & Epstein, 1998) is a 10-item version of the original Rational Experiential Inventory (Epstein, Pacini, Denes-Raj, & Heier, 1996). The RIS measures need for cognition (Cacioppo & Petty, 1982), or the tendency to rely on rational or effortful information processing, and intuitive/experiential thinking, or the reliance on “gut” intuition. Each subscale has 5 items measured on a 1 (completely false) to 5 (completely true) Likert-type scale.
Results
Demographics of both samples are listed in Table 1. The mTurk sample was older than the Subject Pool sample, with a higher percentage of college graduates and a lower percentage of current college students, but there were no statistically significant differences in terms of gender, ethnicity, ACT, or SAT scores. Only 5.2% of the subject pool (taken from an introductory psychology course) intended to major in psychology.
Demographic Characteristics for Participants in Study 1.
Note. PRIC = Psychological Research Inventory of Concepts; LL = lower limit; UL = upper limit.
† p < .08. *p < .05. **p < .01.
The average score on the PRIC was 8.46 (SD = 2.99), or 42.12% correct, with no differences in score by sample. We also found no relationship between completion duration and score, r = .03, ns. However, surprisingly, men (M = 8.96, SD = 3.13) scored higher than women (M = 8.21, SD = 2.82), t(517) = 2.84, p < .01, d = .25, 95% confidence interval [CI]: [.08, .42], and individuals who identified as non-White scored lower (M = 7.97, SD = 3.04) than individuals who identified as White (M = 8.74, SD = 2.96), t(519) = 2.42, p = .02, d = .26, 95% CI [.05, .47]. In terms of educational differences, we compared participants without a bachelor’s degree (n = 401) to participants with a bachelor’s degree (n = 93) and individuals with an advanced degree (n = 22). A one-way analysis of variance (ANOVA) found significant differences between education groups, F(2, 513) = 3.44, p = .03, η2 = .01. Specifically, people with an advanced degree scored higher on the PRIC (M = 10.00, SD = 3.80) than people without a bachelor’s degree (M = 8.43, SD = 2.92), p = .049. Those with a bachelor’s degree (M = 8.86, SD = 3.02) scored in between, not statistically different (using Bonferroni post hoc tests) from either of the other two groups.
Similar to our prior work developing the PRIC (Veilleux & Chapman, 2017, this issue), we evaluated the relationship between PRIC scores and standardized test scores and found that higher PRIC scores were associated with greater ACT (r = .46, p < .001, n = 327) and greater SAT scores (r = .21, p < .01, n = 208). We also examined PRIC scores based on prior experience with statistics and research methods coursework and found, contrary to our previous findings (Veilleux & Chapman, 2017, this issue), no statistically significant differences in PRIC scores based on experience, although individuals who had taken a college-level statistics course had marginally higher scores (M = 8.97, SD = 3.02) compared to people who had not (M = 8.41, SD = 2.97), t(519) = 1.91, p = .06, d = .18, 95% CI [–.001, .37]. We also evaluated any differences between samples on standardized test scores or experience with coursework and found no sample differences on these variables nor in how they related to performance on the PRIC.
Zero-order correlations between all study measures are included in Table 2, with the sample-specific correlations provided in Table 3. Overall, the relationships between the PRIC and other measures were similar across both samples (see Table 3 for Fishers r-to-z transformations comparing the strength of the correlations). Higher scores on the PRIC were associated with greater cognitive effort on the cognitive reflection task and greater psychological knowledge assessed by the PPE. The weaker correlations were only significant for the full sample, namely, that higher PRIC scores were associated with more interest in research, higher need for cognition and lower intuitive thinking. PRIC scores were not related to belief in science or grit.
Correlations Between Psychological Research Concept Inventory and all Measures Included in Both Samples in Study 1.
Note. CRT = Cognitive Reflection Test; PPE = psychology print exposure measure; PRIC = psychological research inventory of concepts.
*p < .05. **p < .01.
Correlations between PRIC and other Measures only, Both in Total and by Sample and With Samples Compared in Study 1.
Note. CRT = Cognitive Reflection Test; PPE = psychology print exposure measure;
PRIC = psychological research inventory of concepts.
† p < .09. *p < .05. **p < .01.
Finally, we assessed whether the PRIC would predict psychological knowledge above and beyond other factors such as need for cognition, cognitive effort (i.e., CRT), and general ability indexed roughly by standardized test score (either ACT or SAT). We first calculated a “standardized test score” variable so that people who took either test could be included in one analysis. We calculated z scores separately for the SAT and ACT and for people who took both, and an average z score was calculated; this score was used as the standardized test score. Predicting PPE as the outcome, we used hierarchical regression by entering demographics (age, gender, and ethnicity) in the first step, with standardized test score in the second step. The third step included the CRT and need for cognition, with the PRIC entered in the final step. Regression results are presented in Table 4. The PRIC significantly predicted increased psychological knowledge even after controlling for demographics, standardized test score, and cognitive effort measured both behaviorally and attitudinally.
Hierarchical Regression Evaluating the PRIC as an Incremental Predictor of Psychological Knowledge in Study 1.
Note. PRIC = psychological research inventory of concepts.
† p < .10; *p < .05. **p < .01.
Brief Discussion
Our goal for this first study was to provide further validation information for the PRIC. Overall, we found evidence of concurrent validity, as higher scores on the PRIC were associated with greater cognitive effort on the CRT and greater psychological knowledge on the PPE. These correlations were significant for both samples separately and together, suggesting robust relationships between the PRIC and these other measures, both indicative of knowledge and performance rather than attitudes. Indeed, it also follows that individuals with greater educational achievement (e.g., advanced degrees) would outperform individuals with less education, as it is also likely that people with advanced degrees have greater exposure to research findings, regardless of the field, and practice with critical thinking and reasoning. A remaining question is whether people with a psychology major degree outperform individuals with a college degree who did not major in psychology, which would indicate discipline-specific learning or at least domain-specific learning (e.g., behavioral sciences) that the PRIC intends to capture.
We also found evidence of convergent validity: significant correlations between the PRIC and some of the attitude measures (e.g., need for cognition, interest in research, and a negative relationship with intuitive thinking) although only with the entire sample. The magnitude of these correlations was low, suggesting statistical significance was due to the large sample size. As expected, the PRIC did not correlate with belief in science or grit, a general measure of perseverance and passion for long-term goals, suggesting discriminant validity.
Unexpectedly, we found performance differences on the PRIC based on gender and ethnicity, where men outperformed women and White participants outperformed people who identified as minorities. Although this is not a unique phenomenon in the realm of academic-based testing (e.g., stereotype threat), it was unanticipated, particularly as we did not find these effects in previous work (Veilleux & Chapman, 2017, this issue), even with similar samples. It may be that these findings were unique to this particular sample; further work using the PRIC will evaluate gender and ethnicity to determine the replicability of this finding.
Finally, we found that after controlling for age, gender, and ethnicity, as well as standardized test scores (either ACT or SAT) and cognitive effort, PRIC scores incrementally predicted psychological knowledge. As a purely correlational study, we cannot make any claims as to causality or temporal order of these results; it may be that greater reasoning in the realm of research methods and statistics results in greater discrimination of psychological constructs, or it may be that in general greater knowledge about psychological processes increases scientific reasoning.
Study 2
Because Study 1 revealed that higher educational achievement was associated with greater PRIC scores, one possibility is that the PRIC simply measures education level rather than understanding of research methods and statistics learned throughout undergraduate psychology degree training. Study 2 was thus conducted to provide validation that individuals at the end of their psychological degree (e.g., graduating senior psychology majors who completed both research methods and statistics) would perform better on the PRIC than college graduates who majored in areas other than psychology.
Method
Participants and Procedure
Participants for this study were advanced psychology majors, juniors, and seniors enrolled in a required capstone course in psychology (either Advanced Research, a course focused on conducting research, or Advanced Seminar, a course focused on understanding current research in a particular topic area within psychology). These capstone students were invited to take the PRIC as part of a department assessment of the psychology major, and respondents were given extra course credit for participation. Just as with previous administrations, all participants completed the PRIC online, and the items for the PRIC were administered in random order immediately prior to a brief set of demographic questions, including overall university grade point average (GPA). The 83 capstone course individuals who completed the survey measure were 22 years old on average (M = 22.31, SD = 3.20), 83.1% women, 80.7% White, and 92.9% seniors.
The primary comparison group was a subsample of participants reported on previously: mTurk participants who completed a college degree but did not major in psychology (Veilleux & Chapman, 2017, this issue; Study 2). These participants completed the PRIC along with questions about prior educational history, including education level achieved (see Veilleux & Chapman, 2017, this issue) and for those who indicated current or past enrollment in college, their college major. For this comparison, we selected the 181 mTurk participants who indicated they did not major in psychology (mean age 35.88, 50.6% women, and 79.2% White) but had completed a bachelor’s degree (n = 142) as well as those with an advanced degree (n = 39).
To be complete, we also examined the capstone students against the mTurk individuals from Study 1 who had completed at least a college degree (n = 114; Mean age 33.27, 43.9% female, and 76.3% White). We did not assess undergraduate majors for these participants, so there are likely some psychology majors in the group.
Results
Because the comparison group of individuals with a college degree who were not psychology majors came from a different sample, we first calculated their PRIC score (M = 9.61, SD = 3.24) and percentage (M = 48.07, SD = 16.20). We then used a one-sample t-test to compare the capstone course students to the average percentage correct on the PRIC of the mTurk participants with a college degree who identified as nonpsychology majors. Results indicated that the capstone course students got a significantly higher score on the PRIC (M = 12.15, SD = 3.01) compared to the mTurk nonpsychology major participants with at least a college degree, t(83) = 7.74, p < .001, d = .84 (95% CI [0.59, 1.09]).
A one-sample t-test was also used to compare the capstone course students to the Study 1 mTurk participants with a college degree (n = 114) who obtained a PRIC score of 9.10 (SD = 3.21) and percentage of 45.48 (SD = 16.03). As expected, the capstone students also outperformed these mixed-major college graduates t(83) = 9.31, p < .001, d = 1.09, 95% CI [0.75, 1.28].
We also calculated a correlation between PRIC scores and self-reported GPA for the capstone students and found no significant relationship, r = .14, p = .21, suggesting that at least for these advanced students, greater scores on the PRIC are not related to overall academic success.
Discussion
The goal of Study 2 was to ensure that the PRIC measured understanding of research methodology and statistics specifically and was not merely a general measure of education level, educational experience, or educational performance. We found that PRIC scores were higher for upper level psychology majors enrolled in a psychology capstone course than PRIC scores for participants who had completed a nonpsychology degree or participants who had completed at least a college degree. Alongside the lack of a significant correlation between PRIC scores and overall GPA within the capstone course students, these results suggest that the PRIC does not simply measure ability in individuals with higher education nor does it measure overall academic success. Rather, the PRIC assesses ability in research methods and statistics in the field of psychology, specifically. These findings are particularly notable because our comparison samples included individuals with completed undergraduate and advanced degrees who have greater levels of education than our capstone participants.
One possible limitation of Study 2 is that our advanced psychology major sample, as juniors and seniors, are temporally closer to their methods and statistics courses than were the comparison samples. However, we previously found that the mTurk participants who had taken research methods and statistics courses a longer time ago performed better on the PRIC than those who had taken those courses more recently (Veilleux & Chapman, 2017, this issue), proximity to course completion is not a plausible alternative explanation for these results. Another limitation was that we collected only overall GPA and not GPA in the major, as it would make sense that PRIC scores related to psychology-specific coursework as opposed to overall GPA. We also recognize that results from the current study are specific to our institution and may not be generalizable to all psychology students or to all nonpsychology degree holders. In future research, it would be beneficial to compare advanced psychology students and advanced nonpsychology students enrolled at the same university and equally close to graduation. However, we believe that the evidence from this study is robust enough to address the concern that the PRIC only assesses education level and that clearly does not appear to be the case.
Study 3
One central function of the development of the PRIC was to be able to use it as an index of literacy in research methods and statistics, beyond course grades. It is notable that at our institution, students must earn at least a C in introductory statistics before they are allowed to enroll in research methods. We also specifically require introductory statistics for psychologists as a part of our major, where students cannot substitute statistics taken in the mathematics department. Thus, we would expect students at the start of the research methods semester to perform better than students in the general psychology subject pool who have not yet taken the psychology-specific statistics course.
In addition, we expected that students would improve their PRIC scores from the beginning to end of a research methods course and that the PRIC would indeed relate to course grades. Establishing the change in PRIC across a semester and relation of the PRIC to course grades was thus the function of Study 3. Because we anticipated scores would be higher than the subject pool for students entering research methods, our predictions here are thus a relatively stringent test of the effect of a research methods class on PRIC scores.
Method
Participants and Procedure
Participants were students in research methods courses who completed the PRIC online for partial class credit. These data include five sections of research methods taught in summer (1 section) and during the regular fall and spring semesters (2 sections each) and were all taught by the same instructor who is the second author of this article. Students were asked to complete the PRIC, the CRT (Frederick, 2005), and the Interest in Research Scale (Vittengl et al., 2004) during the first week of the course as an online homework assignment and received full credit if they completed the measure by the due date and time. Students were again given the Interest in Research Scale and the PRIC in the last week of the semester as another online homework assignment. As part of the assignment instructions, students were encouraged to perform at their best level and to put effort into their consideration of each item. Students’ final exam scores (which included both new and cumulative material) and total proportion of points earned across the course were also collected and used as indices of course success.
In total, 74 students completed both time points, which excludes 9 students who completed the pretest and not the posttest (i.e., people who dropped the course) and 4 students who completed the posttest but not the pretest. To ensure anonymity and to avoid potential demand characteristics and experimenter bias, the first author monitored PRIC completion and communicated whether or not each student finished the PRIC to the course instructor (the second author) who never saw any of the PRIC scores during the semester. After excluding the 2 people who completed the PRIC in less than 15 min, the final sample was 72 students (74% women and 72% White). We did not collect age data for these students, but 66.2% of them indicated they were seniors in terms of college credit.
Results
We used a one-sample t-test to compare PRIC scores for students at the start of the research methods course (M = 11.36, SD = 2.80) to the subject pool participants from Study 1. As a reminder, the 292 subject pool students were enrolled in an introductory psychology course and had average PRIC scores of 8.42 (SD = 2.86; see Table 1 for demographics). As expected, students at the start of research methods, after completing a statistics course, scored higher than subject pool students, t(71) = 8.90, p < .001, d = 1.05, 95% CI [0.70, .1.34].
A paired samples t-test indicated that students’ interest in research did not shift throughout the semester, t(72) = 1.33, ns, d = .16, 95% CI [–0.08, .39], but PRIC scores did increase from the start of the course to the end of the course (M = 12.60, SD = 2.70), t(72) = 3.91, p < .001, d = .46, 95% CI [0.21, 0.70]. We also calculated correlations between PRIC scores and all study measures (Table 5). Both pretest PRIC and posttest PRIC scores correlated with final exam grades and total course points. PRIC scores were also associated with higher cognitive effort on the CRT.
Correlations Between PRIC and all Measures Included in Study 3.
Note. CRT = Cognitive Reflection Test; PRIC = psychological research inventory of concepts. *p < .05. **p < .01.
To evaluate whether posttest PRIC scores provide any incremental prediction of course grades, we also conducted a hierarchical regression predicting grade outcomes. Pretest PRIC and CRT scores were entered in Step 1, with posttest PRIC scores entered in Step 2. When predicting total course points, the overall model was significant, F(3, 68) = 7.63, p < .001, R 2 = .23, but the only unique predictor was pretest PRIC scores, B = 9.94, standard error (SE) = 2.27, t(68) = 4.38, p < .001. The CRT was not predictive with the PRIC in the model, B = –.41, SE = 5.50, t(68) = .07, ns, and the posttest PRIC did not account for additional variance, ΔR 2 = .02, F(1, 68) = 1.83, p = .18. A second hierarchical regression evaluated final exam scores, F(3, 68) = 10.37, p < .001, R 2 = .31. Here, the pretest PRIC score was again significant, B = 9.94, SE = 2.27, t(68) = 4.38, p < .001, although not the CRT, B = .009, SE = .007, t(68) = 1.20, ns. The posttest PRIC did predict additional variability in final exam scores, ΔR 2 = .05, F(1, 68) = 4.68, p = .03, even after controlling for CRT and pretest PRIC.
Brief Discussion
We found evidence that students’ PRIC scores increased during a research methods course, providing further evidence that the measure assesses research methods and statistical knowledge. This is on top of the finding that students entered the research methods course with scores that were higher than the subject pool sample. Considering this information, the current data are notable for several reasons. First, the PRIC covers both research methods and statistical topics, so we would expect that students who completed a course in statistics would be able to answer the statistics-focused questions to a greater degree than students who have not completed statistics, such as the majority of students in general psychology. Second, even if statistics courses do not explicitly cover the nuances of research design, many statistics textbooks and courses necessarily cover basic elements of design (e.g., t-test vs. ANOVA vs. correlational designs) as well as other related topics (e.g., sampling). Thus, these data provide a relatively conservative test of PRIC validity, yet even so we see increases in the PRIC from the start to the end of a semester. Future work should assess the incremental increases in the PRIC from statistics classes alone, research methods courses alone—such as at institutions where research methods is the prerequisite to statistics, rather than the other way around—and in integrated methods and statistics courses.
Our results regarding indices of course success showed that PRIC scores positively correlated with students’ final exam scores, a proxy of their overall research methods ability. The fact that posttest PRIC scores also predict final exam scores after controlling for pretest scores indicate that the PRIC is not simply measuring general student aptitude as they enter the course. The finding that posttest PRIC scores did not predict total course points is understandable, since total points for this course also included pass/fail assignments (e.g., homework assignments, in-class assignments) and quizzes with multiple attempts. Thus, the total points index reflects students’ attendance, effort level, and assignment completion level and not just achievement-based points. The cumulative final exam, however, is an achievement-based score rather than a completion-based score, which is likely why final exam grade was predicted by posttest PRIC scores.
Limitations for the current study include the fact that the course instructor was one of the authors of this article, where many of the topics addressed on the PRIC are covered in the course. The course was designed prior to development of the PRIC such that the course was not designed to “teach to test,” but it is not clear that all research methods instructors cover the same areas with the same depth. In the future, we hope to administer the PRIC as a pretest and posttest to students in research methods courses taught by other faculty, which will also serve the function of replicating the results found here. We predict that any well-designed and well-implemented research methods course should reveal higher PRIC scores at the end of the semester compared to the start of the semester.
General Discussion
The studies in this article provide further evidence of the validity and utility of the PRIC. We replicated prior findings that higher PRIC scores were associated with greater standardized test scores (Veilleux & Chapman, 2017, this issue) and also found that higher PRIC scores were associated with greater cognitive effort and higher knowledge of psychological terminology. Additionally, we found that people with greater psychology education scored higher than nonpsychology majors with a college degree. Finally, PRIC scores increased from the start to end of a research methods course, and PRIC scores predicted performance on a cumulative final exam in a research methods course. The research presented here provides further validation of the PRIC as a standardized measure of research methods and statistical reasoning, with utility for both assessment at the department level and in the classroom.
One limitation of these studies is that all participants, both student and mTurk, completed the PRIC and other measures online. Given that they were not monitored during their completion of the study, the use of this method may unintentionally encourage and lead to low effort on the part of the participants. We attempted to control for this by excluding data from participants who completed the study in too short a time or who failed manipulation checks. However, even those participants who remained in the study may not have been as invested during the measure, as they might have been in a room with a researcher or some form of experimental supervision. Future studies would benefit from assessing whether the PRIC scores differ if administered during actual class periods supervised by an instructor or actual research sessions supervised by an experimenter.
While the PRIC does address both research methods concepts and statistical literacy, the majority of the items are geared toward understanding methodology rather than interpreting results of statistical analyses. This was an intentional decision, given the requisite brevity of the final version of the measure and extant work suggesting that integration of statistics and research methods may increase overall learning (Barron & Apple, 2015; Pliske, Caldwell, Calin-Jageman, & Taylor-Ritzler, 2015). Indeed, we also found that students entering a methods course, having completing a statistics course, scored higher than students in an introductory psychology course. However, we did not compare PRIC scores for those who just completed a statistics course compared to those who just completed a research methods course. An alternative solution would be to develop a separate, though complementary, concept inventory to assess pure statistics concepts and application, potentially with a greater emphasis on probability, variability, hypothesis testing, and advanced concepts such as multiple regression. This would likely be a more effective and practical strategy than adding more statistics-specific questions to the PRIC given that the goal is to keep it brief enough to administer in a single class period.
A final limitation of this study is that it is possible that the PRIC simply assesses critical thinking rather than assessing psychological research reasoning. We acknowledge that the relationship between the PRIC and the CRT and the PRIC and standardized testing scores may suggest that the PRIC relates to intellectual ability and academic effort more generally. However, in support of the PRIC as a psychology-specific measure, we found that scores on the PRIC predicted psychological knowledge (PPE) above and beyond other measures of cognitive effort (CRT, need for cognition) and measures of ability (standardized test scores). Given that the PPE assesses memory for and familiarity with psychological terminology, we believe that the incremental variance attributed to the PRIC suggests more than just critical thinking. We also found that advanced psychology students enrolled in a capstone course performed higher on the PRIC than a mixed sample of individuals with college degrees and more advanced degrees in other fields. This provides strong support to the idea that the PRIC assesses more than just critical thinking or general intellectual ability. However, we acknowledge that this link is not definitive and warrants future testing, just as it would be useful to compare individuals with advanced research degrees (e.g., doctorates) in the sciences with individuals who received advanced degrees in areas that might require critical thinking but not methodological reasoning (e.g., humanities, arts). It would also be useful to evaluate the PRIC alongside the Psychological Critical Thinking Exam (Lawson, 1999; Lawson, Jordan-Fleming, & Bodle, 2015) to determine whether the PRIC is associated with psychology-specific thinking.
The three studies described here represent initial uses of the PRIC. As one of the central goals of concept inventory development is to assess learning beyond grades, future work can evaluate the value of the PRIC as a tool for assessing the retention of research methods concepts after the course has ended and some amount of time has passed, both within the undergraduate degree program and after completion of the bachelor’s degree. Our own program has begun using the PRIC as a measure of Goal 2 (Scientific Inquiry and Critical Thinking) in terms of assessing the strength of our program using APA’s guidelines for the psychology major (American Psychological Association, 2013). Some of our program assessment data were included here (Study 2) as evidence that our majors have more research methods and statistical knowledge at the end of the program than the beginning. We are also using the PRIC to identify areas where growth is needed. We know that despite decreases in misconceptions about psychological concepts over an undergraduate program of study, significant misconceptions still remain at the end of a psychology major’s undergraduate career (Hughes, Lyddy, & Lambe, 2013; Lyddy & Hughes, 2012). That is true with our data as well; many students are graduating with low PRIC scores, suggesting they may have difficulty recognizing quality scientific evidence that they read about in articles or books. We hope that future work will investigate how PRIC scores predict long-term retention of research methods and statistical reasoning after graduation and beyond.
We also hope future studies will use the PRIC to evaluate whether instruction in research methods and statistics is adequate and where specific misconceptions are still evident. The students can likely do better, and we can be doing better as instructors of statistics and research methods. We can engage in item analysis to understand which concepts the students seem to have a harder time understanding (e.g., factorial designs, confounds, the definition of a p value) and then put additional attention onto these topics. We also hope the PRIC can then be used as an outcome measure for testing improved research methods teaching strategies. In fact, we are already planning future studies to evaluate how various approaches to teaching these courses may increase scientific literacy.
As educators of undergraduate students in psychological science, we aimed to create a practical measure that we would use with our students. Ideally, other researchers and educators in research methods and statistics in psychology and the behavioral sciences more generally will adopt the PRIC and use it as a tool in their research laboratories and in their classrooms. Considering the negative opinion students continue to hold of research methods and statistics alongside the centrality of these subjects in our field, any strategies that might improve learning or valuing the scientific method are sorely needed. We hope that the PRIC might have a place in helping to advance these teaching and learning goals in the future.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was funded by a “Research in Teaching” grant from the University of Arkansas Teaching and Faculty Support Center, and from a Society for the Teaching of Psychology/Psi Chi Assessment Grant.
