Abstract
Quantitative synthesis of data from single-case designs (SCDs) is becoming increasingly common in psychology and education journals. Because researchers do not ordinarily report numerical data in addition to graphical displays, reliance on plot digitizing tools is often a necessary component of this research. Intercoder reliability of data extraction is a commonly overlooked, but potentially important, step of this process. The purpose of this study was to examine the intercoder reliability and validity of WebPlotDigitizer (Rohatgi, 2015), a web-based plot digitizing tool for extracting data from a variety of plots, including XY coordinates of interrupted time-series data. Two coders extracted 3,596 data points from 168 data series in 36 graphs across 18 studies. Results indicated high levels of intercoder reliability and validity. Implications of and recommendations based on these results are discussed in relation to researchers involved in quantitative synthesis of data from SCDs.
Quantitative synthesis of data from single-case designs (SCDs) are being published with increasing regularity in psychology and education journals. This trend is largely related to the inclusion of SCDs in What Works Clearinghouse (WWC) evidence reviews and accompanying advancement of statistical methods for SCDs (Shadish et al., 2014) to identify evidence-based practices (EBPs) within these and related disciplines (e.g., school psychology, special education). Because quantitative synthesis of data from SCDs is in its infancy, there is a great deal of variability in methodological and statistical procedures utilized by researchers. This research requires a number of steps (i.e., literature search, application of inclusion criteria, appraisal of methodological quality, coding individual- and study-level variables, data extraction, and calculation of effect sizes), each of which influences whether a practice is determined evidence-based.
Because SCD researchers do not ordinarily report numerical data in addition to graphical displays, reliance on plot digitizing tools to determine precise XY coordinates when calculating certain effect sizes is necessary. A plot digitizing tool, broadly, is a software that allows users to extract numerical data from different types of plots or graphs. There are several plot digitizing tools available that vary in price, platform compatibility, and overall usability (Moeyaert, Maggin, & Verkuilen, 2016). To extract XY coordinates, these tools require users to import graphs, calibrate axes by clicking known values for the tool to interpolate a coordinate system, and manually click each data point in a data series (Shadish, Brasil, Illingworth, Nagler, & Rindskopf, 2009). These tools often have an automatic data extraction function, but in the researchers’ experience, precision is often compromised by multiple data series in the same panel and other “noise” in graphical displays. Ultimately, the utility of automatic data extraction functions should be examined empirically, as suggested by Shadish and colleagues (2009).
Although researchers regularly examine intercoder reliability of number of studies meeting inclusion criteria, appraisal of methodological quality, and coding of individual- and study-level variables, reliability and validity of data extraction has been consistently overlooked in recent quantitative syntheses of SCDs; however, it is possible that systematic error could affect results of data analysis and conclusions drawn by researchers about the effectiveness of an intervention. Most quantitative syntheses of SCDs report the plot digitizing tool used to extract data, but this is not always the case (e.g., Bowman-Perrott et al., 2013; Bowman-Perrott, Burke, Zhang, & Zaini, 2014). Two existing quantitative syntheses of SCDs were identified that discuss intercoder reliability of data extraction in detail. Dart, Collins, Klingbeil, and McKinley (2014) reported exact agreement and mean proportional agreement in a meta-analysis of peer-management interventions for a subset of articles included. Exact agreement, or the percentage of extracted y-values that were identical across coders, was 66.2%. In this study, mean proportional agreement was calculated by dividing one coder’s extracted value by the second coder’s. When there was noncorrespondence between coders, the smaller value was always the dividend, so that quotients were less than 1. The mean of these values was calculated and multiplied by 100% resulting in a mean proportional agreement of 93.9%. In a meta-analysis of behavioral interventions for adolescents and adults with autism spectrum disorder, Roth, Gillis, and Reed (2014) reported mean proportional agreement of 96.3% for a subset of the articles included. This study used a point-by-point approach to defining proportional agreement in that an agreement was defined as “two data points being identical or one unit apart” (Roth et al., 2014, p. 265). It would be considered an agreement, for example, if one coder extracted a y-value of 2 and a second coder extracted a y-value of 3, but a disagreement if the second coder extracted a y-value of 4.
There are four existing investigations of intercoder reliability and validity of plot digitizing tools with respect to quantitative syntheses of SCDs (Boyle, Samaha, Rodewald, & Hoffman, 2013; Flower, McKenna, & Upreti, 2015; Rakap, Rakap, Evran, & Cig, 2016; Shadish et al., 2009). These studies have provided evidence for the intercoder reliability and validity of the following plot digitizing tools: DataThief III (Flower et al., 2015; Tummers, 2006), DigitizeIt (Bormann, 2012; Rakap et al., 2016), GraphClick (Arizona Software Inc., 2010; Boyle et al., 2013; Flower et al., 2015), and Ungraph (Biosoft, 2004; Rakap et al., 2016; Shadish et al., 2009). Comparisons of DigitizeIt, GraphClick, and UnGraph (Rakap et al., 2016) and DataThief III and GraphClick (Flower et al., 2015) resulted in equally reliable and valid data extraction. In existing investigations of plot digitizing tools, it does not appear one outperforms another.
Intercoder reliability, or whether two coders independently extract similar data, is critical in quantitative synthesis of data from SCDs. Table 1 summarizes extant investigations of intercoder reliability. Existing studies have examined intercoder reliability at three levels. First, researchers have examined whether coders extract the same number of data points per graph (Rakap et al., 2016; Shadish et al., 2009) or across all studies coded (Boyle et al., 2013). This has been analyzed as both percentage of agreement (92.3%-98.3% agreement across two studies) and the correlation between the number of data points extracted by both coders (r = .999 across all three studies). Discrepancies across coders were usually attributed to qualities of graphs, that is, overlapping data points due to multiple data series in the same panel or data points within the same data series that were difficult to distinguish from one another because they were very close to each other (Rakap et al., 2016; Shadish, 2009).
Summary of Intercoder Reliability Statistics Calculated Across Four Studies of Plot Digitizing Tools.
Second, all studies examined whether coders extracted similar y-values. This has been analyzed as percentage of exact agreement, percentage of proportional agreement, and the correlation between y-values extracted by both coders. Boyle et al. (2013) reported exact agreement across coders for 32% of data points. Percentage of proportional agreement, which has been operationalized as y-values within 1% of the y-axis range, has been much higher across studies. Boyle et al. reported 89.5% proportional agreement and Rakap et al. (2016) reported percentage proportional agreement by plot digitizing tool used, ranging from 90.6% to 91.3%. Correlations between y-values extracted by both coders have been almost perfect across all four studies. Researchers have identified three primary sources of error in identifying noncorrespondence between y-values extracted by coders: incorrectly calibrating axes, following an incorrect data series, and clicking slightly different locations on a data point (Boyle et al., 2013; Rakap et al., 2016). In the most systematic investigation of this noncorrespondence to date, Rakap and colleagues found clicking slightly different locations on a data point was the reason for the disagreement between 83.2% and 85.5% of the instances of noncorrespondence, depending on the plot digitizing tool used.
Finally, both Rakap et al. (2016) and Shadish et al. (2015) investigated the reliability of the mean difference between baseline and intervention calculated from both sets of extracted data. This is an important step in calculating some effect sizes, thus provides some idea as how unreliability in data extraction may affect later effect size calculation. Both studies reported high correlations ranging from r = .954 to .965; however, these correlations are slightly lower than other reported reliability indices.
Validity, or how well extracted data corresponds to known values, has been studied using two methods. First, researchers have correlated mean scores reported in the original research report with mean scores calculated from extracted data. These analyses have yielded near-perfect correlations (Rakap et al., 2016; Shadish et al., 2009). However, usually only a small set of primary sources included in these articles have reported means by phase or case. Second, researchers have systematically generated hypothetical data sets with corresponding graphs and correlated known values with the values extracted from the graphs by the coders. This method of examining validity has also yielded very high correlations (Boyle et al., 2013; Flower et al., 2015). Because SCD researchers are supplementing visual analysis with effect size calculation with increasing regularity, it is possible to investigate validity by comparing effect sizes calculated from extracted data with those reported in original research reports, which presumably are calculated from known numerical values.
The purpose of the current study was to examine the intercoder reliability and validity of WebPlotDigitizer (Rohatgi, 2015) in extracting graphed data from SCDs. The current study extends existing investigations of intercoder reliability and validity of plot digitizing tools in several ways. First, we utilized WebPlotDigitizer, a web-based application for extracting data from a variety of plots, including XY coordinates of interrupted time-series data. This tool has not been subject to existing investigations of intercoder reliability and validity of extracting graphed data from SCDs. Because it appears to have been used in at least one existing meta-analysis of SCD (Burke, Boon, Hatton, & Bowman-Perrott, 2015), it is important to establish its intercoder reliability and validity. Second, this study incorporates all indices of intercoder reliability used in existing studies, that is, intercoder reliability of the number of data points extracted, relationship of y-values, and relationship between a common index of non-overlap, Tau-U (Parker, Vannest, Davis, & Sauber, 2011), calculated from extracted data. Third, because SCD researchers have increasingly reported Tau-U in research reports, this study examines validity in a novel way by comparing Tau-U values reported in original research reports with those calculated from extracted data.
Method
Article Identification
In December 2015, the first author searched Google Scholar using the keyword Tau-U with the words behavior, education, autism, or psychology in the title of the publication. This search strategy led to the identification of 81 unique results. All studies were peer-reviewed journal articles. To be included, studies must have included deliberate introduction of an intervention. Quantitative syntheses of SCDs utilizing Tau-U and conceptual articles discussing use and properties of Tau-U were excluded. The study must have utilized a SCD (i.e., reversal, multiple baseline including multiple probe, alternating treatment, or a combination of these designs). Studies must have calculated and reported Tau-U. Only studies utilizing school-aged participants were included. Included studies are denoted with an asterisk in the reference list. Eighteen studies published between 2012 and 2015 in 16 different journals met the inclusion criteria. In all, 3,596 data points were extracted from 168 data series in 36 graphs across 18 studies.
Hardware/Software
To extract data and conduct data analysis, the primary coder and one secondary coder utilized a HP ProDesk 600 G1 Tower PC. The other secondary coder utilized an HP Compaq 8200 Elite Convertible Minitower PC. All computers had an Intel Core i5 processor and ran Windows 7 Professional. All coders used an HP USB Optical Mouse to extract data.
WebPlotDigitizer is a web-based plot digitizing tool available to users free of charge. After launching the digitizer, users upload a screenshot of an XY chart and are prompted to calibrate the axes. After calibrating, the user manually clicks each data point within the data series and downloads the extracted coordinates as a Microsoft Excel spreadsheet.
Coders
The primary coder (the first author) extracted data from all data series. Two secondary coders, one doctoral student in school psychology and one undergraduate psychology major, extracted data from approximately one half of the data series. All coders had prior experience with WebPlotDigitizer in conducting a meta-analysis of school-based behavior reduction interventions.
Training
Prior to extracting data, the primary coder modeled each step of a task analysis of data extraction using a graph from a study not included in the present analyses (see Table 2 for the task analysis). Next, the secondary coders practiced data extraction with feedback using the same graph. The duration of the training was approximately 30 min. Secondary coders were not trained to a specific mastery criterion.
Task Analysis of Data Extraction Using WebPlotDigitizer.
Procedures
Data extraction was completed in accordance with the task analysis in Table 2. All coders had access to the task analysis throughout data extraction. Studies were randomly assigned to secondary coders to approximately equate the number of data series. Secondary coders completed all data extraction within 3 days of training. Secondary coders were instructed not to communicate in regard to data extraction. Coders were also instructed not to edit extracted data. With the exception of Bouck, Bouck, and Hunley (2015), Rispoli et al. (2013), and Rispoli et al. (2015), y-values were rounded to the nearest whole number. For Bouck et al. (2015), several data points appeared to fall halfway between the nearest whole numbers (i.e., participants appeared to be given half credit for the accuracy of certain math problems). Coders visually inspected these graphs and changed appropriate values to reflect half credit. For Rispoli et al. (2013) and Rispoli et al. (2015), information in the text indicated the dependent variable was a rate (i.e., frequency of behavior per unit of time); therefore, these values were rounded to the nearest 10th using WebPlotDigitizer. Secondary coders exported data from WebPlotDigitizer to a preformatted Microsoft Excel spreadsheet for data management. Spreadsheet columns were labeled with the participant’s pseudonym, the dependent variable, and phase (i.e., A, B).
Data Analysis
Extracted data were imported into SPSS for analysis of intercoder reliability and validity. To examine intercoder reliability, a variety of descriptive statistics as well as Pearson’s r were calculated to determine the strength and direction of the association between y-values and Tau-U values calculated from the coders’ extracted data. To examine validity, we calculated descriptive statistics and Pearson’s r to examine the strength and direction of the association between Tau-U values reported in original research reports and those calculated from extracted data.
Tau-U is an index of non-overlap that can control for undesirable baseline trend (see Parker et al., 2011, for a conceptual introduction). Tau-U values range from 0 to 1 and are interpreted as the proportion of improvement in data from baseline (A) to intervention (B). For instance, a Tau-U of .80 for a given phase contrast suggests 80% of data improved between the phases or, similarly, that 80% of data do not overlap. Phase contrasts can be combined for participants as well as over several participants either within or across studies, and thus has utility for researchers conducting quantitative syntheses of SCDs (Parker, Vannest, & Davis, 2014). In interpreting the magnitude of effects, Vannest and Ninci (2015) suggested benchmarks of 0 to .20 as small, .20 to .60 moderate, .60 to .80 large, and greater or equal to .80 as large or very large, though benchmarks vary slightly across other studies and typically are accompanied by the caveat that the Tau-U values should be evaluated against the social validity of intervention effects and client needs.
The free calculator at www.singlecaseresearch.org (Vannest, Parker, & Gonen, 2011) was used to calculate Tau-U. To calculate Tau-U, articles were read to determine which phase contrasts (e.g., A/B), and combinations of phase contrasts (e.g., A/B contrasts combined for all participants) were reported. Unless a phase contrast involved a maintenance phase (e.g., Huskens, Verschuur, Gillesen, Didden, & Barakova, 2012) or generalization probes (e.g., Miller, Dufrene, Olmi, Tingstrom, & Filce, 2015), all reported Tau-U values were entered into SPSS by study, participant, dependent variable, and phase contrast.
Because there is variability in how researchers calculate Tau-U using the online calculator at www.singlecaseresearch.org, articles were read to determine whether decision rules were explicitly stated as to the conditions under which baseline trend was controlled and whether Tau-U values resulting from the combination of phase contrasts were combined or weighted. When these rules were explicitly stated, they were followed. If they were not explicitly stated, Tau-U was calculated without controlling for trend or weighting values resulting from the combination of phase contrasts. If these values were discrepant from those reported in the article, then Tau-U was recalculated by controlling for trend and/or weighting values resulting from the combination of phase contrasts, and the value nearest that reported in the article was entered in SPSS. Although Tau-U is a proportion, researchers sometimes reported negative values, likely because the online calculator outputs negative values when data points in a B phase are lower than an A phase (i.e., behavior reduction interventions). If researchers reported negative Tau-U values, we did as well. These steps were taken to minimize discrepancies in Tau-U values due to differences in their calculation and/or reporting and maximize our ability to interpret any degree of noncorrespondence as attributable to data extraction.
Results
Intercoder Reliability
Intercoder reliability was investigated by comparing the primary coder’s data with the secondary coder’s data in terms of the number of data points extracted, the relationship of the y-values extracted, and the relationship of Tau-U values calculated from the extracted data.
In all, 3,596 data points were extracted from 168 data series in 36 graphs across 18 studies. In terms of the agreement of the number of data points extracted across coders, 17 of 18 (94%) studies were in exact agreement. For the one study where there was disagreement, the secondary coder extracted values from two data points that were generalization probes (i.e., 146 data points extracted by the primary coder and 148 data points extracted by the secondary coder). This discrepancy was reconciled prior to examining the correspondence between y-values.
The relationship of y-values extracted across coders was examined in three ways. First, we investigated the exact correspondence of y-values. In all, 2,461 of 3,596 (68%) of y-values were in exact agreement. The percentage of exact agreement averaged across studies was 72% (SD = 30%, range = 11%-100%). The percentage of exact agreement was 100 for six of 18 studies.
Second, we examined proportional correspondence of y-values. Values within 1% of the y-axis range were considered to be in agreement. For example, if the primary coder extracted a value of 9 and the secondary coder extracted a value of 8 and the y-axis ranged from 0 to 100, then this was considered an agreement (i.e., [9 − 8] / 100 × 100% = 1%). However, if the primary coder extracted a value of 9 and the secondary coder extracted a value of 8 and the y-axis ranges from 0 to 10, then this would not be considered an agreement (i.e., [9 − 8] / 10 × 100% = 10%). Overall, 3,308 of 3,596 (92%) of y-values were in proportional agreement. The percentage of proportional agreement averaged across studies was 92% (SD = 15%, range = 46%-100%). The percentage of proportional agreement was 100 for nine of 18 studies.
Third, we examined the correlation between y-values extracted across coders by study. The average Pearson correlation was r = .999 (SD = .001, range = .998-1.000) which indicates a near-perfect relationship of the y-values extracted across coders.
As an extension of the intercoder reliability of the mean difference between baseline and intervention analyses described in Rakap et al. (2016) and Shadish et al. (2009), we examined the intercoder reliability of 139 Tau-U values calculated from both sets of extracted data that mirrored those reported in original research reports. The correlation was strong and near-perfect, r = .997, p < .001.
Validity
Validity was investigated by comparing Tau-U calculated from the primary coder’s extracted data with Tau-U values reported in original research reports. Tau-U values reported in original research reports were considered the standard by which to evaluate the validity of extracted data. As stated above, 139 Tau-U values were calculated and 54% were in exact agreement. About 10% (n = 14) of Tau-U values were discrepant by greater than .10. There was a near-perfect relationship between Tau-U calculated from the primary coder’s extracted data and Tau-U values reported in original research reports, r = .989, p < .001.
Discussion
With quantitative syntheses of SCDs being published regularly in the fields of psychology and education, particular attention should be paid to the intercoder reliability and validity of plot digitizing tools. Rarely do researchers report the numerical data needed to calculate various effect sizes within their manuscripts; therefore, data must be extracted from the included graphical displays. The purpose of this study was to evaluate the intercoder reliability and validity of WebPlotDigitizer, a web-based plot digitizing tool.
Results of the present study indicate WebPlotDigitizer is a reliable and valid tool for extracting data from single-case graphs. In terms of intercoder reliability, more than half of the extracted values were in exact agreement and more than 90% were in proportional agreement. Correlations between the two coders resulted in a near-perfect relationship. These data are highly consistent with other investigations of intercoder reliability of plot digitizing tools (Boyle et al., 2013; Flower et al., 2015; Rakap et al., 2016; Shadish et al., 2009).
This study expanded the current literature by comparing Tau-U values reported in the original studies with those calculated from the coders’ extracted data. Slightly over half of the calculated Tau-U values perfectly matched what was reported in the original research. However, around 10% of the calculated effect sizes were discrepant from reported values by more than .10. Although steps were taken to minimize the degree of noncorrespondence between the two sets of Tau-U values, it is possible that procedural differences in calculation and/or reporting may have accounted for some degree of noncorrespondence. Resulting correlations between the two sets of Tau-U values was near-perfect.
Implications and Limitations
There are accumulating data to suggest data extraction is reliable and valid for a number of plot digitizing tools. As indicated by Boyle et al. (2013), high levels of intercoder reliability may be partially accounted for by previous experience of coders with plot digitizing tools. Also, systematic procedures for training coders as well as supports for extracting (i.e., task analyses) and managing data (i.e., preformatted spreadsheet labeled with the participants’ pseudonym, dependent variable, and phase) likely bolster intercoder reliability. Whether researchers conducting quantitative syntheses of SCDs utilize these sorts of systematic procedures for training coders as well as supports for data extraction and management is often unknown.
As with other steps of quantitative synthesis of SCDs (e.g., appraisal of methodological quality, coding individual- and study-level variables), there is a degree of unreliability in data extraction, particularly when examining percentage of exact and proportional agreement. Consequently, it may be beneficial for researchers to examine intercoder reliability of data extraction for a subset of articles included in quantitative syntheses of SCDs (e.g.,Dart, Collins, Klingbeil, and McKinley 2014; Roth et al., 2014). Researchers might do well to report information about plot digitizing tool used, training provided to coders, supports for data extraction and management, percentage of studies in which intercoder reliability was examined, degree of exact and proportional agreement of extracted data, and whether any degree of unreliability impacts subsequent effect size calculations.
As a number of plot digitizing tools have been shown to yield reliable and valid data, with none appearing to outperform others, researchers and practitioners may consider different dimensions of usability in making a selection. Among other research questions, Moeyaert et al. (2016) examined usability of four plot digitizing tools in extracting data from SCDs. Four coders rated DataThief, Ungraph, XYit, and WebPlotDigitizer on a scale of software usability and a single-item reflecting overall user friendliness. The authors also reported the average time it took to extract data from graphs utilizing a multiple baseline design with one dependent variable. Results of the scale of software usability and the single-item reflecting overall user friendliness favored Ungraph and WebPlotDigitizer. Data were extracted using Ungraph, XYit, and WebPlotDigitizer in an average of about 15 min per graph with DataThief taking almost twice as long. Although we did not record duration of coders’ data extraction per graph, this aligns, at least anecdotally, with our experience for the current study. It should be noted that the graphs included in the current study varied by design and number of dependent variables, which would likely introduce additional variability. Plot digitizing tools were additionally compared on cost: WebPlotDigitizer, as mentioned previously, is free whereas DataThief costs US$25, Ungraph costs US$300, XYit costs US$89.
In terms of training provided to coders, extant studies suggest brief trainings involving a model with opportunities for practice with feedback are typically sufficient to yield high levels of intercoder reliability, even with novice or minimally experienced coders. Access to supports for data extraction and management discussed previously likely assist coders through the process. Training coders to a mastery criterion may be an additional strategy to increase intercoder reliability of data extraction further. Moeyaert et al. (2016) showed experienced coders were able to extract data more quickly than novice coders, which is an important consideration for many researchers and practitioners. In the current study, training might also have explicitly addressed examples and nonexamples of generalization probes depicted in graphs given the primary and secondary coder disagreed on the number of data points included in one study, perhaps because this was unacknowledged during training.
Although not a specific aim of this study, a good deal of variability in the calculation and/or reporting of Tau-U values was noted and thus provide for discussion points. Rather than comment on how Tau-U should be calculated, it seems prudent for researchers to be more explicit about certain aspects of their analytic plan. First, it would be beneficial if researchers were more specific about decision rules for controlling for undesirable baseline trend. Decision rules, when stated, hinged on a variety of factors including Tau values for a within phase contrast or associated p values. Sometimes all baseline trends were controlled, regardless of the associated Tau or p value. Usually, however, there were no stated decision rules stating the conditions under which undesirable baseline trend was controlled. Second, in combining phase contrasts using the referenced free online calculator (Vannest et al., 2011), there appears to be variability in whether Tau-U values are simply combined or weighted. In replicating Tau-U values, it would be helpful for researchers to report whether values were combined or weighted.
One limitation of this study dealt with the aforementioned variability in the calculation and/or reporting of Tau-U. As stated, steps were taken to minimize the degree of noncorrespondence between the two sets of Tau-U values; however, procedural variations in calculation and/or reporting cannot be completely ruled out as a source of noncorrespondence. Thus, future research may examine validity in yet another novel way, perhaps by intentionally searching for and identifying original research reports that include numerical data in additional to graphical displays. As noted in the introduction, another area for future research may involve investigating the intercoder reliability, validity, and efficiency of automatic data extraction methods incorporated into plot digitizing tools in relation to manual methods utilized in this study.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
