Abstract
There are few assessments that gather valid, highly detailed data on short-term (i.e., weekly) symptom frequency/severity retrospectively. In particular, methodologies that provide valid data for research investigating symptom changes are typically prospective, expensive, and burdensome. The purpose of this study was to evaluate a new interactive and graphical assessment tool for gathering detailed information about eating-related symptom frequency/severity retrospectively over a 3-month period. A mixed eating disorder sample (N = 113) recruited from the community provided symptom data once weekly for 12 weeks and completed the Interactive, Graphical Assessment Tool (IGAT) assessing eating disorder symptoms on three occasions to determine the test–retest and concurrent validity of the IGAT. The IGAT performed marginally better than other measures for retrospective symptom frequency assessment in the eating disorders and did so at a greater level of detail than other available tools. Future research should evaluate the IGAT with other behaviors of interest.
Uncovering causal, maintaining, and prognostic factors influencing the fluctuation of symptoms in eating disorders first requires accurate measurement of these symptoms over time. Ideally, to avoid errors committed when recalling and reporting on past events (detailed below), individuals would be continuously measured in real-time, providing an uninterrupted and unbiased picture of their symptoms. Ecological momentary assessment (EMA; Shiffman, Stone, & Hufford, 2008) is a method that approximates this ideal by sampling individuals in their natural environments repeatedly (e.g., several times per day) over generally short periods of time (e.g., 1-2 weeks). This method is best for capturing short-term symptom change and the temporal association of symptoms with other variables of interest. Most commonly, palm-top computers or smartphones are used to query individuals as they go about their everyday lives, providing an externally valid sample of their experiences. This methodology is powerful for examining relationships that are often central to theoretical models of the maintenance of behavior (e.g., emotion regulation models of binge eating; Smyth et al., 2007). In the eating disorders field, evidence indicates that EMA provides useful data for answering specific questions that are often of interest (Wonderlich et al., 2015).
As exciting as EMA is, it is not ideal or practical for all situations. It is a prospective methodology, capturing information from protocol initiation to termination. In many research and clinical situations, individuals present for the first time (e.g., study entry or treatment admission), and information about their recent past is highly valued. In others, they can only be sampled on a few occasions over a long period of time (e.g., a few follow-up assessments over many years) when uninterrupted pictures of their symptoms are desired. In these situations, retrospective methods are required.
The Eating Disorder Examination interview (Fairburn, Cooper, & O’Connor, 2008) is the most widely used retrospective measure that captures symptom frequency fluctuations, providing data on symptom frequency at the month level. In addition, this interview measures five other constructs that may not be of interest to those studying symptom course and diagnostic status, although they are meaningful for other purposes. It is also time-consuming and expensive to administer, with research finding that relatively few clinicians (15%-36%) use this or any other reliable and valid measure as part of the treatment intake, progress, or outcome process (Anderson & Paulosky, 2004; Towne, De Young, & Anderson, in press). Thus, for those interested in characterizing symptom change over the recent past (e.g., the past 3 months) in the eating disorders to understand diagnostic change, individual symptom change, and the relationship between these symptoms over time, there is a need for a retrospective measure that can efficiently capture symptom fluctuation information, particularly at a higher resolution than is currently available. Such a tool must minimize the biases inherent to the retrospective reporting process.
The aim of the present study was to evaluate a new computer-based assessment tool that does not address the same research questions as EMA, for which EMA is better suited. Instead, this retrospective measure may be appropriate for contexts in which an individual first presents or cannot engage in frequent reporting over longer periods of time. It captures symptom fluctuation at the week level, a level of detail appropriate for capturing symptom fluctuations over periods of time (e.g., months or even years) that influence diagnostic status, course, and outcome. The present study tests this measure’s test–retest validity and concurrent validity.
Symptom Fluctuation in Eating Disorders
Eating disorders are often long standing (Herzog et al., 1999). Review studies have found that only 46% of women with anorexia nervosa achieve full remission 4 to 10 years after presentation (Steinhausen, 2002), and only 48% of women with bulimia nervosa achieve this state 5 to 10 years after presentation (Keel & Mitchell, 1997). Even so, marked symptom fluctuation both over the short (i.e., week-to-week; Edler, Lipson, & Keel, 2007; Lester, Keel, & Lipson, 2003) and long term (i.e., years; Eddy et al., 2008; Milos, Spindler, Schnyder, & Fairburn, 2005; Tozzi et al., 2005) characterizes eating disorders. For instance, relapse occurs in about 30% of individuals with bulimia nervosa (Keel & Mitchell, 1997) and 10% of individuals with anorexia nervosa (Kordy et al., 2002; Strober, Freeman, & Morrell, 1997). Early response to treatment (i.e., symptom change) is one of the most robust predictors of success at end of treatment (i.e., remission; Thompson-Brenner, Shingleton, Sauer-Zavala, Richards, & Pratt, 2015). Furthermore, fluctuations in the frequency of binge eating and purging in women with bulimia nervosa are associated with menstrual cycle phase (Edler et al., 2007; Lester et al., 2003), which demonstrates how obtaining knowledge about patterns of symptom fluctuation can illuminate possible maintaining and contributory factors.
Studies of long-term symptom fluctuation in the eating disorders are often studies of diagnostic crossover (i.e., meeting diagnostic criteria for one eating disorder at one point in time and meeting diagnostic criteria for a different eating disorder at another point in time). These studies indicate that approximately one-third of individuals whose initial eating disorder diagnosis is anorexia nervosa later develop bulimia nervosa, and 14% to 27% of those initially diagnosed with bulimia nervosa later develop anorexia nervosa (Eddy et al., 2008; Tozzi et al., 2005). Another study found that only 28.6% of participants never crossed over to another diagnosis or experienced a remission of symptoms across three assessment points spaced over 2.5 years (Milos et al., 2005). In short, these studies indicate that symptom fluctuation is the norm rather than the exception in eating disorders over the long term.
Outcome and diagnostic crossover research has greatly informed the study of eating disorder course by characterizing heterogeneous change over time. However, their reliance on categorical outcomes (e.g., remission and relapse) and diagnoses to describe individuals’ status at different points in time limits their findings to characterizing symptom course at the syndrome level. In addition, the use of different definitions for outcome terms (Clausen, 2004; Couturier & Lock, 2006; Keel, Mitchell, Miller, Davis, & Crow, 2000; Kordy et al., 2002) makes the comparability of findings across studies unclear. Thus, examining symptom change at the syndrome level obscures the phenomena of interest, as there are many ways that an individual may experience diagnostic crossover or a change in outcome (e.g., relapse). Information at the symptom level is required for the characterization of individual symptom change.
Retrospective Assessment of Symptom Frequency in the Eating Disorders
The Eating Disorder Examination (EDE; Fairburn et al., 2008), which is an interview, and the Eating Disorder Examination–Questionnaire (EDE-Q; Fairburn & Beglin, 2008) are among the most widely used assessments for eating disorder psychopathology (Anderson & Paulosky, 2004; Towne et al., in press). Both measures retrospectively assess symptom frequency (i.e., binge eating, self-induced vomiting, laxative misuse, and compulsive exercise). The EDE assesses the previous 3 months; the EDE-Q assesses the previous 28 days, and both result in symptom frequency estimates at the month level. A study using the EDE, EDE-Q, and daily self-monitoring in a sample of 82 individuals with binge eating disorder found that the EDE-Q appeared to accurately assess the frequency of objectively large binge eating episodes (i.e., eating episodes involving both a subjective sense of loss of control and an objectively large amount of food), as the frequencies obtained from the EDE-Q and from self-monitoring were significantly correlated and were not significantly different (Grilo, Masheb, & Wilson, 2001a). However, the EDE-Q underestimated the frequencies of both subjectively large binge eating episodes (i.e., eating episodes with a sense of loss of control but not an objectively large amount of food) and objective overeating episodes (i.e., eating episodes that are objectively large and are not accompanied by a sense of loss of control). The results of this study were replicated in a separate sample of 47 individuals with binge eating disorder (Grilo, Masheb, & Wilson, 2001b). Thus, it appears as though the EDE and EDE-Q assess the frequency of objectively large binge eating episodes at the month level over the past 1 month (EDE-Q) or 3 months (EDE) reasonably well.
The only other study that tested a method of retrospective report for eating disorder symptom frequencies came to a different conclusion about the validity with which individuals report the frequency of binge eating (Bardone, Krahn, Goodman, & Searles, 2000). In this study, a timeline follow-back approach was used to measure binge eating and binge drinking in college students in a sample of 38 college-age women who endorsed binge eating and drinking. Bardone et al. compared data obtained daily via interactive voice response technology over the telephone, and the same data were collected retrospectively via 12-week timeline follow-back. The timeline follow-back approach (Sobell, Maisto, Sobell, & Cooper, 1979) uses a calendar to facilitate recall for assessing alcohol use over the preceding 90 days, and an interviewer aids the respondent in recalling his daily alcohol consumption by going back in time, 1 day at a time. When alcohol use is stable, the interviewer may simply inquire about “change points” to accelerate the process. The alcohol timeline follow-back method has been found to have good test–retest reliability in clinical samples, students, and individuals from the community (Sobell, Sobell, Leo, & Cancilla, 1988). Indeed, Bardone et al. (2000) found that the timeline follow-back approach demonstrated good concurrent validity with the daily reports of drinking; however, they found the timeline follow-back method to be “woefully inadequate” for the assessment of binge eating (p. 9). Overall, the frequency of binge eating obtained from the timeline follow-back vastly underestimated that obtained from the daily reports. Unfortunately, no other studies to the authors’ knowledge have compared the frequencies of eating disorder behaviors assessed retrospectively with those obtained prospectively. Thus, even less is known about individuals’ ability to accurately recall self-induced vomiting, laxative, fasting, and exercise episodes.
The Retrospective Reporting Process
A number of errors have been documented in the responses of individuals who are asked to report on their past behavior (Schwarz & Sudman, 1994). These errors may occur at one or more of the four main parts of this process: comprehension of question content, retrieval of relevant information, estimation and judgment of the retrieved information, and formulation of a response (Friedenreich, 1994). Retrieval and estimation are particularly relevant to retrospective assessment.
Early theories suggested that when asked for information about the frequency of a behavior, individuals use a pure “recall and count” strategy (Brewer, 1994); however, later theories posited “reconstructive” strategies that involve decomposing the behavior into manageable units that are then reconstructed to estimate frequency over time (Means, Swan, Jobe, & Esposito, 1994). According to Bradburn, Rips, and Shevell (1987), there are two parts to the estimation process. First, individuals decompose the event into parts, such as by figuring out how often an event occurs during a week and then multiplying it by the number of weeks in the reference period. This appears to be most relevant to events that occur with some regularity such that an “average week” is easy to bring to mind. Second, individuals allow the amount of information about the event that they can recall to influence their estimate. If it is easy to think of events, then most individuals reason that the event must have occurred more frequently than events that are harder to retrieve.
This process is also known as the availability heuristic and can result in overestimating particularly salient events through forward telescoping and underestimating less salient events through omission. Forward telescoping refers to reporting that an event occurred within the reference period when it actually occurred outside of the reference period such that the event has been “telescoped” forward in time (Brennan, Chan, Hini, & Esslemont, 1996). Omission occurs when individuals forget that an event occurred or forget when the event occurred so that they do not report that it occurred during the reference period (Brennan et al., 1996). Finally, although not specifically related to saliency, rounding can affect all estimates and refers to the tendency to report numbers in commonly used increments, such as intervals of 5 or 10.
Solutions for mitigating these sources of error include bounding, in which individuals are provided with the dates of the reference period (e.g., How many times has X occurred since April 16th?) rather than specifying the amount of time in the reference period (e.g., How many times has X occurred over the past 6 months?). This has been found to reduce forward telescoping (Brennan et al., 1996). Also effective is the use of landmark events to bind a reference period (Loftus & Marburger, 1983). Landmark events may be either public (e.g., New Year’s Day, September 11th) or private (e.g., the first day of a new job, the birth of a child). Thus, they may be both assessor- and respondent-generated. Finally, placing events and dates in a personal timeline on a calendar can help individuals create a context for the behaviors that are of interest and aid in the accurate retrieval of relevant information (Friedenreich, 1994).
The Present Study
The primary aim of this research was to evaluate the performance of a short-term (i.e., 3-month interval) retrospective assessment tool that was designed specifically to address the issues raised here. In addition to using assessor- and respondent-generated landmark events and a calendar to bind the reference period, and a response interface that does not lend itself to rounding, this assessment aims to capitalize on the visual representation of data to reduce respondent burden.
Visual elements in graphical representations use spatial proximity and placement to represent other spatial relations or are visuospatial metaphors (Tversky, Bauer Morrison, & Betrancourt, 2002). For instance, distance most often conveys dimensions such as time or cost. Evidence indicates that children as young as 5 years old and individuals from many different cultures interpret graphical representations of information similarly (Tversky et al., 2002). Possibly the most important aspect of graphical representations is that they follow the congruence principle, which states that “the structure and content of the external representation should correspond to the desired structure and content of the internal representation” (Tversky et al., 2002, p. 257). When this principle is followed, graphical information is quickly and relatively easily interpreted. Because “graphics externalize internal knowledge” (Tversky et al., 2002, p. 248), the use of graphics in assessments can relieve the respondent of some short-term memory and processing burdens. Indeed, decreasing subject burden increases the accuracy of retrospective report (van der Vaart, Van der Zouwen, & Dijkstra, 1995).
The present assessment tool uses line graphs to illustrate the relationship between time and symptom frequency/severity. Line graphs may be particularly useful for depicting change over time, since individuals have an inherent tendency to map changing rates onto changing slopes (Gattis & Holyoak, 1996). This study tested the test–retest reliability and concurrent validity of this Interactive, Graphical Assessment Tool (IGAT) that obtains eating disorder symptom data at the week level for a 12-week period. This assessment tool is self-administered and requires relatively few resources to employ, which may increase the likelihood of its adoption.
Method
Participants
Participants (N = 113) were recruited using advertisements inviting individuals with eating disorders to participate in a study to further the development of a questionnaire to assess eating disorder behaviors. A sample of individuals heterogeneous in their presentation of eating disorder symptoms was sought to represent a broad range of eating disorder symptom presentations. Advertisements were posted in the Albany, New York, area and on the Internet (e.g., craigslist.org). Participants resided in 34 states and the District of Columbia. The majority of participants were women (n = 93; 82.3%), and participants’ ages ranged from 18 to 72 years, with a mean (SD) of 32.91 (12.65) years. A total of 69.9% of participants identified their ethnicity as Caucasian, 9.7% as Asian/Pacific Islander, 8.0% as Black/African American, 7.1% as Hispanic, 0.9% as Native American, and 2.6% chose not to provide this information.
Measures
Eating Disorder Diagnostic Scale
The Eating Disorder Diagnostic Scale (EDDS; Stice, Telch, & Rizvi, 2000) is a 22-item self-report questionnaire that assesses eating disorder psychopathology for the purposes of deriving Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition (DSM-IV) eating disorder diagnoses. The EDDS has been found to be highly specific and sensitive for eating disorder diagnoses (Krabbenborg et al., 2012). Two supplemental items were added to this scale for the present study to assess the magnitude of distress over the presence of purging behaviors and binge eating separately on a 7-point scale from “not at all distressed” to “extremely distressed.” This measure was used to screen participants in the present study. Cronbach’s alpha for the symptom composite scale was .85.
Eating Disorder Examination–Questionnaire
The EDE-Q (Fairburn & Beglin, 2008) is a 28-item self-report questionnaire that assesses the frequency of binge eating, purging, and driven exercise over the previous 4 weeks in addition to providing indicators of eating-related psychopathology on four subscales and a global score. The behavioral frequencies obtained for binge eating days and episodes and purging (laxative plus self-induced vomiting) episodes were the primary data used in this study. Cronbach’s alpha for the items comprising the global score was .91.
Positive and Negative Affect Schedule
The Positive and Negative Affect Schedule (PANAS; Watson, Clark, & Tellegen, 1988) assesses affect along two dimensions. Individuals rate the degree to which 20 affect-laden words describe how they have felt using a 5-point scale from “Very slightly or not at all” to “Extremely.” In the present study, participants completed the PANAS with instructions for how they feel “in general” or how they have felt over “the past week,” depending on the assessment occasion. Internal consistency for the PANAS measured by Cronbach’s alpha averaged across time points was .95 for positive affect and .92 for negative affect using “the past week” instructions.
Interactive, Graphical Assessment Tool
The IGAT (Figure 1) is a computer-based self-report instrument that collects data at the week level for the previous 12 weeks. In this illustration and testing, the IGAT was configured to measure body weight, binge eating, purging (i.e., self-induced vomiting, laxative, enema, or suppository misuse, insulin or other medication misuse, and diuretic misuse), fasting, exercise undertaken in response to feeling bad about body shape or weight, and stress. Stress was added as a noneating disorder symptom and nonbehavioral symptom to examine the extent to which the IGAT could collect meaningful data on less discrete constructs and constructs not specific to the sample used in this study. Pilot testing with focus groups of individuals with eating disorders and college students indicated that it was an acceptable medium through which to query this information. The IGAT can be viewed at the following web address: http://goo.gl/8kk2Ov.

A screenshot of the Interactive, Graphical Assessment Tool (IGAT), showing an example of the customized calendar on the bottom, the week for which the participant is providing a rating above it, and the line graph interface at the top.
The IGAT begins by displaying a calendar that depicts the previous 12 weeks and with the current day highlighted. The respondent is instructed to personalize the calendar by selecting events from a list of 14 provided events (e.g., birthday, death of a loved one, birth of a loved one, and first day of work/school). Individuals may also enter their own events. This checklist format is used again later for the reporting of methods of purging, because checklists have been found to improve the accuracy of retrospective report (van der Vaart et al., 1995). This personalized calendar later serves to bind the reference period being assessed. Participants then respond to a series of items that assess their highest and lowest weight over the past 12 weeks and the presence of binge eating (individually indicating the presence of objectively large eating episodes and subjective loss of control), purging behaviors, fasting, and exercise. A specific start date is provided with each of these questions to bind this 12-week interval. Their answers to these questions are used to tailor the assessment to meet their specific symptom presentation.
Next, a tutorial describes the task participants are asked to perform. Individuals are asked to indicate the frequency of their symptoms (i.e., binge eating, purging, fasting, and exercise) and their body weight and stress level (from 0 to 7) over the past 12 weeks by manipulating (i.e., dragging) dots on a line graph. Twelve dots are present, each representing one of the past 12 weeks. The x-axis of the line graph represents time and ranges from “12 weeks ago” on the left to “last week” on the right. The y-axis represents number of days ranging from a low of zero to a high of seven, except when rating body weight when the y-axis changes to range from the respondent’s low body weight to his or her high weight or for stress when it represents low to high. For each week on which binge eating or purging is reported to have occurred (i.e., at least 1 day), participants are queried with a pop-up window to input the number of individual episodes that occurred. The dots on the lines are connected to one another by a colored line. Tabs are located at the top of the line graph; each tab represents one domain to be rated. The color of this tab corresponds to the color of the line connecting the dots. Clicking on these tabs changes the domain to be rated. Importantly, as a respondent moves from one tab (i.e., symptom/feature) to another, the lines she or he has created for other symptoms remain visible so that she or he may use them to contextualize his or her knowledge about the symptom she or he is currently rating. Finally, the calendar the respondent personalized during the first part of the assessment is visible below the graph. When the respondent places the mouse pointer over a dot or clicks and drags the dot, the week that the dot represents is displayed in detail to further aid the respondent in contextualizing his or her symptoms in time.
Weekly Self-Monitoring Questionnaire
Body weight in pounds and the frequency of binge eating, purging, fasting, and exercise motivated by feeling bad about one’s weight or shape were assessed with seven questions that were each answered with a number. Four questions required a response from zero to seven regarding the number of days over the previous week that various eating disorder symptoms were experienced. The question about body weight required a response in pounds. Binge eating and purging were also assessed in terms of the number of individual episodes experienced. The wording of these items was virtually identical to the wording of the items on the IGAT, in diagnostic manuals (e.g., DSM-IV-TR; American Psychiatric Association, 2000), and on other eating disorder assessments (e.g., the EDDS).
Design and Procedures
Interested participants were screened over the telephone using the EDDS for the following inclusion criteria: (1) at least 18 years of age and (2) the presence of an eating disorder as indicated by at least one of the following: (a) body weight below a body mass index of 18kg/m2 and undue influence of body weight or shape on self-evaluation; (b) the presence of purging (i.e., self-induced vomiting, laxative, diuretic, or enema misuse, or the abuse of medication such as insulin) at least once every 2 weeks and undue influence of body weight or shape on self-evaluation or marked distress about purging; (c) the presence of binge eating episodes at least once per week and undue influence of body weight or shape on self-evaluation or marked distress about binge eating. Participants were excluded if they were unable or unwilling to take part in the study procedures as described below.
Participants completed a battery of self-report questionnaires via the Internet, including the PANAS, EDE-Q, and IGAT (in the order listed) on three separate occasions that were each 6 weeks apart (i.e., at baseline, Week 6, and Week 12; Figure 2). In addition to these three assessment points, participants completed the Weekly Self-Monitoring Questionnaire (WSMQ) and the PANAS weekly for the 12 weeks of the study. Participants were compensated with $5 for completion of the baseline questionnaires, $10 each for the Weeks 6 and 12 assessments, and $1.25 for each of the 12 weekly assessments for a total possible compensation of $40.

The timing of the administrations of the Interactive, Graphical Assessment Tool (IGAT) (baseline, Week 6, and Week 12) and Weekly Self-Monitoring Questionnaire (weekly from Weeks 1 through 12) over the course of participation and the time periods they assessed.
Data Analyses
To examine test–retest reliability, the most proximal 6 weeks assessed at the baseline administration of the IGAT (IGATB) were compared individually to the most distal 6 weeks assessed at the Week 6 administration (IGAT6). In addition, the most proximal 6 weeks from the IGAT6 were individually compared to the most distal 6 weeks assessed at the Week 12 administration (IGAT12). Because each administration of the IGAT assesses the previous 12 weeks, each pair of IGAT administrations (i.e., IGATB/IGAT6 and IGAT6/IGAT12) assessed an overlapping 6-week interval. Thus, test–retest reliability for this instrument represents the measurement of symptoms that occurred at a single point in time (i.e., a particular week) but were assessed twice, 6 weeks apart. Most applications of test–retest reliability purport to measure the repeatability of the measurement of a construct that is presumed to be invariant across the test–retest interval; however, in the present application, this is not assumed, because the construct being assessed is identical at both time points. There is no true score change. Although this would ostensibly increase the purity of the test–retest estimate, this introduces error attributable to the limits of human memory and retrieval processes, which likely worsen as the period being assessed moves further into the past.
For this reason, in the concurrent validity analyses, the WSMQ data were compared to both the most proximal 6 weeks of the IGAT6 and all 12 weeks of the IGAT12, with the WSMQ data used as the criterion. Not only does this arrangement allow for estimates of the validity of the data reported on the IGAT12 for all 12 weeks recalled, but it also allows for a comparison between the strength of the associations between the 12-week and 6-week recall of the first 6 weeks of data collection. This comparison provides some evidence for whether the concurrent validity of the IGAT depends on the length of the period being assessed. In addition, the IGAT was compared with the EDE-Q at three time points (i.e., baseline, Week 6, and Week 12). For these comparisons, only the most recent 4 weeks assessed by the IGAT were used, as this is the same time frame assessed by the EDE-Q. Furthermore, these 4 weeks were aggregated, as the EDE-Q assessed the prior 28 days without any additional delineations of the time interval.
The test–retest reliability and concurrent validity of the IGAT were examined using an intraclass correlation coefficient (ICC), which can vary from 0 to 1 and can be interpreted as the proportion of variance accounted for in one trial by another trial or the ratio of variance explained to variance explained plus error variance (Weir, 2005). As described by McGraw and Wong (1996), the ICC used in this study is labeled ICC (A,1) because the present application fits the 2-way random model in which a single score is provided. More specifically, this ICC is a measure of agreement rather than a measure of consistency. This means that disagreements adversely affect the coefficient even if they differ by a constant value. In other words, in the present study, both systematic and random errors were of interest when estimating reliability and validity. In addition, these data require a 2-way model because all participants provided data for each time point being compared, which is the typical test–retest configuration (Weir, 2005). They require a random effects model because both the raters (i.e., participants) and the ratings (i.e., the weeks of symptoms) are samples from larger populations. No universally accepted standards exist against which to compare an ICC to quantify its strength, because the merits of reliability and validity are relative to existing measures and the area of study (Weir, 2005). However, to facilitate discussion, we used the verbal descriptors suggested by Shrout (1998) for reliability: 0.81-1.0 is substantial, 0.61-0.80 is moderate, 0.41-0.60 is fair, 0.11-0.40 is slight, 0.00-0.10 is virtually none. Finally, paired samples t tests or Wilcoxon Signed Rank tests, depending on whether the variables were normally distributed, were used to compare weekly symptom frequency (and body weight) means between the IGAT and the WSMQ and the IGAT and EDE-Q to assess if, and in what direction, any systematic bias existed in the IGAT. Stress was compared to weekly negative affect ratings as measured by the PANAS. Although stress and negative affect are different constructs, fluctuations in stress are associated with PANAS negative affect (e.g., at a Pearson r of .44) but not positive affect (e.g., Pearson r of −.09; Watson, 1988; see also Clark, Watson, & Leeka, 1989). Thus, in this case, a statistically significant relationship would provide evidence of convergent rather than concurrent validity. Spearman’s Rho correlation coefficients were calculated to quantify the strength of the association between stress and PANAS negative affect. Statistical significance was set at p < .05 for all tests, as sensitivity to differences was of the utmost importance, and it is interpreted in light of the large number of tests performed.
Receiver operating characteristic curve analyses were conducted to quantify the performance of the IGAT at identifying the presence of eating disorder behaviors at diagnostic threshold levels. The area under the curve indicates the proportion of randomly drawn pairs of cases/noncases that would be correctly identified as such, with .50 indicating chance-level performance and 1.0 indicating perfect performance (Hanley & McNeil, 1982). Confidence intervals that do not include .50 indicate statistically significant case identification. Other diagnostic efficiency statistics were calculated using case status as derived from WSMQ data as the criterion as follows: sensitivity = true positives / (true positives + false negatives); specificity = true negatives / (true negatives + false positives); negative predictive value = true negatives / (false negative + true negatives); positive predictive value = true positive / (true positive + false positive; Akobeng, 2006).
Power Analyses
Statistical power for the test–retest reliability and concurrent validity analyses was calculated a priori using G*Power (Version 3.0; Erdfelder, Lang, & Buchner, 2007). For 80% power to detect small–medium sized mean differences (i.e., Cohen’s d = .30) with alpha set at .05 using two-tailed paired samples t tests, 90 individuals were needed to complete the 12 weeks of weekly self-monitoring and 2 consecutive administrations of the IGAT. These same 90 individuals would provide statistical power to detect statistically significant correlation coefficients measuring bivariate relationships that share at least 4.5% of variance.
Results
Eight individuals (7.1%) met inclusion criterion (a) but not (b) or (c), resembling anorexia nervosa, restricting type (ANr). Ten (8.8%) met inclusion criterion (a) and either or both (b) and (c), resembling anorexia nervosa, binge eating-purging type (ANbp). Fifty-five (48.7%) met inclusion criteria (b) and (c) but not (a), resembling bulimia nervosa, purging type (BNp). Thirty (26.5%) met inclusion criterion (c) but not (a) or (b), resembling binge eating disorder (BED). Finally, 10 (8.8%) met inclusion criterion (b) but not (a) or (c), resembling purging disorder (PD).
Participant Retention
A total of 119 individuals were eligible for the study and agreed to participate. Six individuals never began the study, withdrawing before providing data. Of the 113 who began participation, 97% (n = 110) completed the baseline assessment, 84% (n = 95) completed the Week 6 assessment, and 78% (n = 88) completed the Week 12 assessment. Dropout at the Weeks 6 and 12 assessments was not associated with gender, χ2(1) = 0.06, p = .81 and χ2(1) = 0.72, p = .40, 6 and 12 weeks, respectively; eating disorder diagnosis, χ2(4) = 3.44, p = .49 and χ2(4) = 0.58, p = .97, 6 and 12 weeks, respectively); illness severity as measured by the EDE-Q Global score, t(108) = 0.31, p = .758 and t(108) = −0.08, p = .933, 6 and 12 weeks, respectively); or mood as indicated by PANAS positive, t(109) = 0.61, p = .542 and t(109) = 0.28, p = .779, 6 and 12 weeks, respectively) and negative, t(109) = 0.28, p = .782 and t(108) = −0.29, p = .774, 6 and 12 weeks, respectively) affect at baseline.
For the weekly assessments, a total of 1,320 assessments were possible (i.e., 12 assessments for each of the 110 individuals who completed the baseline assessment). A total of 1,066 weekly assessments were completed for a completion rate of 81%. The mean (SD) number of weekly assessments completed was 9.69 (3.96) and ranged from 0 to 12 but was extremely positively skewed with 68 of the 110 individuals completing all 12 assessments. The number of weekly assessments completed was not associated with eating disorder diagnosis, F(4, 108) = 0.74, p = .57, or gender, t(109) = −1.03, p = .30; however, the number of weekly assessments completed was significantly associated with baseline body mass index, r(108) = −0.22, p = .02 and binge eating, r(72) = 0.23, p = .04; and fasting, r(59) = 0.35, p = .006, frequency. The median times to complete the IGAT at the baseline, Week 6, and Week 12 assessments were 13.6, 10.3, and 8.3 minutes, respectively.
Body Weight
Week-by-week test–retest reliability and concurrent validity results for the IGAT’s assessment of body weight can be found in Table 1. In sum, the 6-week test–retest reliability for assessing body weight was substantial. The mean ICC for IGATB / IGAT6 ratings was .996, and the mean ICC for IGAT6 / IGAT12 ratings was .995. There were no mean differences between assessments. Furthermore, the IGAT was highly valid as indicated by the mean ICC for the IGAT12 being .993. The mean ICC for the 6 overlapping weeks of the WSMQ and the IGAT6 was also .993. There were no statistically significant mean differences.
Body Weight: Test–Retest Reliability and Concurrent Validity of IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; ICC = intraclass correlation coefficient; WSMQ = Weekly Self-Monitoring Questionnaire. ICCs significant at p < .001. Mean differences are the first assessment listed minus the second evaluated with Wilcoxon signed rank tests. Boldfaced means are significantly different at p < .05. IGATB, 6, 12 = Baseline, Week 6, and Week 12. See Figure 2 for help interpreting assessment weeks.
Binge Eating
Days
Week-by-week test–retest reliability and concurrent validity for binge eating days can be found in Table 2. Overall, the 6-week test–retest reliability was fair, with a mean ICC for IGATB / IGAT6 ratings of .584 and a mean ICC for IGAT6 / IGAT12 ratings of .549. One of 12 mean differences was statistically significant in the direction of indicating that binge eating days were estimated to be fewer by the second IGAT administration.
Binge Eating Days: Test–Retest Reliability and Concurrent Validity of IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; ICC = intraclass correlation coefficient; WSMQ = Weekly Self-Monitoring Questionnaire. ICCs significant at p < .001. Mean differences are the first assessment listed minus the second evaluated with paired samples t tests. Boldfaced means are significantly different at p < .05. IGATB, 6, 12 = Baseline, Week 6, and Week 12. See Figure 2 for help interpreting assessment weeks.
The validity for binge eating days is indicated by a mean ICC for the IGAT12 of .595. The mean ICC for Weeks 1 through 6 was somewhat lower than Weeks 7 through 12 (.546 vs. .643). The mean ICC for Weeks 1 through 6 as measured by the IGAT6 was .597, which is slightly higher than the value obtained by the IGAT12 for these same weeks (i.e., .546). Together, these may suggest that the accuracy of the IGAT’s assessment of binge eating days decreases as the time period being queried becomes more distant. Mean differences were present for Weeks 2 and 6, and both indicated that IGAT12 underestimated the frequency of binge eating days.
Comparing the IGAT and EDE-Q at baseline, Week 6, and Week 12 on binge eating days indicated variability, with low agreement at baseline, ICC = .323, F(98, 98) = 1.94, p < .001, but strong agreement at Week 6, ICC = .848, F(91, 91) = 12.01, p < .001; and Week 12, ICC = .828, F(84, 84) = 10.70, p < .001. The means were not significantly different between them at baseline, t(98) = −1.66, p = .100); Week 6, t(91) = 0.17, p = .867; or Week 12, t(84) = −1.39, p = .170.
Episodes
Week-by-week test–retest reliability and concurrent validity for binge eating episodes can be found in Table 3. The 6-week test–retest reliability was fair, with mean ICCs for the IGATB / IGAT6 and IGAT6 / IGAT12 ratings of .440 and .443, respectively. One of 12 mean differences for binge eating episodes was statistically significant and, as with binge eating days, indicated that episodes were estimated to be fewer by the second IGAT administration.
Binge Eating Episodes: Test–Retest Reliability and Concurrent Validity of IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; ICC = intraclass correlation coefficient; WSMQ = Weekly Self-Monitoring Questionnaire. All ICCs significant at p < .001, except when indicated by *, which is significant at p < .05. Mean differences are the first assessment listed minus the second evaluated with Wilcoxon signed rank tests. Boldfaced means are significantly different at p < .05. IGATB, 6, 12 = Baseline, Week 6, and Week 12. See Figure 2 for help interpreting assessment weeks.
The validity evidence for binge eating episodes is indicated by a mean ICC for the IGAT12 of .639. The mean ICC for Weeks 1 through 6 was slightly lower than Weeks 7 through 12 (.615 vs. .662). The mean ICC for Weeks 1 through 6 as measured by the IGAT6 was .557, which is lower than the value obtained by the IGAT12 for these same weeks (i.e., .615), providing conflicting evidence for the effect of time. Mean differences were present for Weeks 3, 5, 6, and 12, all indicating that the IGAT underestimated the frequency of binge eating episodes.
Comparing the IGAT and EDE-Q at baseline, Week 6, and Week 12 on binge eating episodes indicated that episodes were not particularly consistent across measures at baseline, ICC = .323, F(98, 98) = 1.94, p < .001, but were much better at Week 6, ICC = .665, F(90, 90) = 5.15, p < .001, and Week 12, ICC = .709, F(84, 84) = 5.86, p < .001. The means of binge eating episodes were not significantly different between these assessments at baseline, t(98) = −0.13, p = .894, or Week 12, t(84) = 0.88, p = .381, but the EDE-Q estimated approximately 3.5 fewer binge eating episodes than the IGAT6, t(90) = 2.24, p = .028.
The sensitivity and specificity of the IGAT12 for identifying individuals who reported an average of two or more binge eating episodes per week over the 12-week study period using the data from the WSMQ as the criterion were calculated to facilitate comparison between the performance of the IGAT and the Timeline Follow-back for binge eating conducted by Bardone and colleagues (2000). The sensitivity of the IGAT was .81, and the specificity was .88. The positive predictive value was .91, and the negative predictive value was .77. Additional information on the performance of the IGAT for identifying the presence of binge eating episodes at threshold level according to the DSM-5 (American Psychiatric Association, 2013) can be found in Table 4.
Sensitivity, Specificity, Negative and Positive Predictive Value, and Area Under the Curve of the IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; NPV = negative predictive value; PPV = positive predictive value; AUC = area under the curve; CI = confidence interval. Values are for the IGAT’s performance of detecting the behaviors listed at a frequency of greater than or equal to once per week (i.e., the DSM-5 frequency cut-off) using the WSMQ as the criterion.
Purging
Days
Week-by-week test–retest reliability and concurrent validity for purging days can be found in Table 5. The 6-week test–retest reliability was moderate, with a mean ICC for IGATB / IGAT6 ratings of .726 and a mean ICC for IGAT6 / IGAT12 ratings of .769. None of the mean differences were statistically significant.
Purging Days: Test–Retest Reliability and Concurrent Validity of IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; ICC = intraclass correlation coefficient; WSMQ = Weekly Self-Monitoring Questionnaire. ICCs significant at p < .001. Mean differences are the first assessment listed minus the second evaluated with Wilcoxon signed rank tests. Boldfaced means are significantly different at p < .05. IGATB, 6, 12 = Baseline, Week 6, and Week 12. See Figure 2 for help interpreting assessment weeks.
The concurrent validity evidence for purging days was indicated by a mean ICC for the IGAT12 of .808. The mean ICC for Weeks 1 through 6 was slightly lower than Weeks 7 through 12 (.787 vs. .829). The mean ICC for Weeks 1 through 6 as measured by the IGAT6 was .834, which is higher than the value obtained by the IGAT12 for these same weeks (i.e., .787), both suggesting that the validity of the IGAT for purging days decreased the further back in time that information was queried. Mean differences were present at Weeks 6 and 12, suggesting that the IGAT underestimated purging days on these weeks.
Episodes
Week-by-week test–retest reliability and concurrent validity for purging episodes can be found in Table 6. The 6-week test–retest reliability was fair, with mean ICCs for the IGATB / IGAT6 and IGAT6 / IGAT12 ratings of .461 and .592, respectively. Three of the 12 mean differences were statistically significant, such that the second IGAT administration resulted in lower estimates than the first; however, this pattern only exists for the IGAT6 / IGAT12 and not the IGATB / IGAT6.
Purging Episodes: Test–Retest Reliability and Concurrent Validity of IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; ICC = intraclass correlation coefficient; WSMQ = Weekly Self-Monitoring Questionnaire. ICCs significant at p < .001. Mean differences are the first assessment listed minus the second evaluated with Wilcoxon signed rank tests. Boldfaced means are significantly different at p < .05. IGATB, 6, 12 = Baseline, Week 6, and Week 12. See Figure 2 for help interpreting assessment weeks.
The validity evidence for purging episodes is indicated by a mean ICC for the IGAT12 of .725. The mean ICC for Weeks 1 through 6 was lower than Weeks 7 through 12 (.679 vs. .771). The mean ICC for Weeks 1 through 6 as measured by the IGAT6 was .622, which is slightly lower than the value obtained by the IGAT12 for these same weeks (i.e., .679), providing inconsistent evidence for the effect of time. Mean differences were present at Weeks, 2, 6, 7, 10, and 12, all in the direction of indicating that the IGAT underestimated purging episodes on these weeks.
Comparing the most recent 4 weeks assessed by the IGAT and EDE-Q at baseline, Week 6, and Week 12 for purging episodes (i.e., self-induced vomiting plus laxative misuse episodes for the EDE-Q) resulted in high ICCs for baseline, .641, F(99, 99) = 4.54, p < .001; Week 6, .788, F(89, 89) = 8.40, p < .001; and Week 12, .765, F(85, 85) = 7.45, p < .001. The means were not significantly different for baseline, t(99) = −0.33, p = .743; Week 6, t(89) = 0.80, p = .424; or Week 12, t(85) = 0.52, p = .605. Information on the performance of the IGAT for identifying the presence of purging episodes at threshold level according to the DSM-5 (American Psychiatric Association, 2013) can be found in Table 4.
Fasting
Week-by-week test–retest reliability and concurrent validity for fasting days can be found in Table 7. The 6-week test–retest reliability was variable, ranging from fair to moderate. The mean ICC for IGATB / IGAT6 ratings was .546, and the mean ICC for IGAT6 / IGAT12 ratings was .764. One of the 12 mean differences was statistically significant, with the second IGAT administration resulting in a lower estimate than the first.
Fasting: Test–Retest Reliability and Concurrent Validity of IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; ICC = intraclass correlation coefficient; WSMQ = Weekly Self-Monitoring Questionnaire. ICCs significant at p < .001. Mean differences are the first assessment listed minus the second evaluated with Wilcoxon signed rank tests. Boldfaced means are significantly different at p < .05. IGATB, 6, 12 = Baseline, Week 6, and Week 12. See Figure 2 for help interpreting assessment weeks.
The validity for fasting days was variable, with a mean ICC for the IGAT12 of .648. The mean ICC for Weeks 1 through 6 was lower than Weeks 7 through 12 (.539 vs. .757). The mean ICC for Weeks 1 through 6 as measured by the IGAT6 was .567, which is slightly higher than the value obtained by the IGAT12 for these same weeks (i.e., .539), both differences suggesting that the validity of the IGAT for assessing fasting decreased the further back in time that information was queried. There were significant mean differences found for Weeks 1, 2, 3, 6, 8, 9, and 12, all in the direction of indicating that the IGAT was underestimating the frequency of fasting days. Information on the performance of the IGAT for identifying the presence of fasting at threshold level according to the DSM-5 (American Psychiatric Association, 2013) can be found in Table 4.
Exercise
Week-by-week test–retest reliability and concurrent validity for exercise days can be found in Table 8. The 6-week test–retest reliability was variable, ranging from fair to moderate. The mean ICC for IGATB / IGAT6 ratings was .564, and the mean ICC for IGAT6 / IGAT12 ratings was .752. One of the 12 mean differences was statistically significant, with the second IGAT administration resulting in a lower estimate than the first.
Exercise: Test–Retest Reliability and Concurrent Validity of IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; ICC = intraclass correlation coefficient; WSMQ = Weekly Self-Monitoring Questionnaire. ICCs significant at p < .001. Mean differences are the first assessment listed minus the second evaluated with Wilcoxon signed rank tests. Boldfaced means are significantly different at p < .05. IGATB, 6, 12 = Baseline, Week 6, and Week 12. See Figure 2 for help interpreting assessment weeks.
Validity evidence for exercise days is indicated by a mean ICC for the IGAT12 of .616. The mean ICC for Weeks 1 through 6 was only slightly lower than Weeks 7 through 12 (.612 vs. .620). The mean ICC for Weeks 1 through 6 as measured by the IGAT6 was .601, which is slightly lower than the value obtained by the IGAT12 for these same weeks (i.e., .612), providing inconsistent evidence for the effect of time. There were significant mean differences at Weeks 1, 2, 4, 5, 6, 7, 8, 9, 10, and 12, all in the direction of indicating that the IGAT underestimated the frequency of exercise. Table 4 contains information on the performance of the IGAT for identifying the presence of exercise at threshold level according to the DSM-5 (American Psychiatric Association, 2013).
Stress
Week-by-week test–retest reliability and concurrent validity for stress can be found in Table 9. The 6-week test–retest reliability was variable, ranging from slight to fair. The mean ICC for IGATB / IGAT6 ratings was .316, and the mean ICC for IGAT6 / IGAT12 ratings was .442. However, only 1 of the 12 comparisons of mean differences was statistically significant, with the second IGAT administration resulting in a lower estimate than the first.
Stress: Test–Retest Reliability of IGAT.
Note. IGAT = Interactive, Graphical Assessment Tool; ICC = intraclass correlation coefficient. All ICCs significant at p < .001 except those indicated by ǂ, which are statistically significant at p<.05. Mean differences are the first assessment listed minus the second evaluated with paired samples t tests. Boldfaced means are significantly different at p < .05. IGATB, 6, 12 = Baseline, Week 6, and Week 12. See Figure 2 for help interpreting assessment weeks.
The Spearman’s Rho values quantifying the strength of the relationship between the PANAS negative affect scale assessed on each of the 12 weeks and stress assessed via the IGAT12 averaged 0.361 with a standard deviation of 0.101 and a minimum and maximum of 0.173 and 0.516, respectively. P values ranged from .125 to <.001; all but 2 of the 12 weeks were statistically significant at p < .05. Similarly and more parsimoniously, a pseudo R2 calculated by comparing an intercept-only generalized linear model predicting stress from the IGAT12 to the same model with PANAS negative affect indicates that 15.8% of the variance in stress can be accounted for by negative affect, B = 0.86, Wald χ2(1) = 26.16, p < .001. The covariance structure was specified as independent based on the best model fit as indicated by the quasi likelihood under independence model criterion.
Discussion
Test–Retest Reliability
The results of the test–retest reliability analyses indicate that the IGAT reliably assesses the domains examined in this study. In particular, body weight was extremely reliably assessed. Purging days were very reliably assessed, and binge eating days and purging days were more reliably assessed than episodes were for either behavior. Similarly, fasting and exercise days were fairly reliably assessed. All of the ICC values were statistically significant, and mean differences were in the direction of higher reports on initial compared with subsequent assessments. Thus, as the time period being queried became more distal, individuals’ estimates of the frequencies of their behaviors decreased; however, no obvious patterns emerged in the ICCs that indicate the presence of linear trends for the test–retest reliability over the 6-week periods. Any inferences drawn from these differences must be done so with caution, as only 3.5% of the tested mean differences were statistically significant, which is not more than should be expected by chance.
Concurrent Validity
The findings regarding the concurrent validity of the IGAT correspond closely to those supporting its reliability, but the corresponding ICCs are lower in magnitude, which is to be expected, with reliability providing the ceiling for validity. Approximately 30% of the mean differences between the IGAT and the WSMQ were statistically significant, and all of the differences were in the direction of the IGAT underestimating symptom frequencies. This pattern, though unfortunate, is consistent with previous research on the retrospective assessment of binge eating frequencies (e.g., Bardone et al., 2000; Grilo et al., 2001a), finding lower estimates with retrospective recall compared to self-monitoring. The omission of behaviors on the IGAT may be due to forgetting, error in estimation, or some other motivated cause, but the present study cannot shed light on which of these possibilities is most likely. These omissions are likely responsible for the reduced diagnostic sensitivity of the IGAT, as the omission of behaviors push some individuals below the criterion threshold.
In addition, the IGAT assessed binge eating and purging frequencies very similarly to the EDE-Q when limited to the past 4 weeks and aggregated across this period, which is the time period assessed by the EDE-Q and the method by which it is conducted. These findings indicate that the IGAT appears to be at least as valid at assessing the frequency of these behaviors as the EDE-Q over this time frame. Furthermore, in one comparison, the IGAT resulted in less underestimation than the EDE-Q, which has often been found to result in lower binge eating estimates than the EDE interview (e.g., Grilo et al., 2001a, 2001b), suggesting that the IGAT may have gathered more valid data on binge eating.
Few other options are available for comparing the performance of the IGAT to existing measures, because few measures exist that are comparable to the IGAT. The one exception is the timeline follow-back procedure used by Bardone and colleagues (2000) to retrospectively assess binge eating over a 12-week period in a group of college students who also provided the frequency of their binge eating daily via interactive voice response telephone calls. The sensitivity and specificity of the timeline follow-back for identifying individuals who binge ate at or above the DSM-IV-TR diagnostic cutoff frequency for bulimia nervosa (i.e., at least 2 times weekly over 3 months; American Psychiatric Association, 2000) were 1.00 and 0.33, respectively. In the present study, the sensitivity and specificity of the IGAT for the same purpose were 0.81 and 0.88. Thus, the IGAT was less sensitive than the timeline follow-back but far more specific and approached levels of acceptability for the diagnosis of mental disorders. These differences are likely both due to methodological differences in the timeline follow-back and the IGAT and differences in the samples (i.e., nonclinical college students, many of whom did not binge eat, and a community sample of individuals with eating disorders, most who did binge eat). The diagnostic efficiency statistics further indicate that the IGAT generates comparable data to weekly self-monitoring, while slightly underestimating behavioral frequencies. The receiver operating characteristic curves indicated that approximately 90% of DSM-5 threshold/nonthreshold randomly drawn pairs would be accurately classified as such for binge eating and purging by the IGAT. Although the IGAT could be used as a screening instrument, there are lower burden measures (e.g., extremely brief paper and pencil measures such as the EDDS) that may be more appropriate to use as screening instruments. Instead, the IGAT may be best at monitoring symptom fluctuations at or surrounding diagnostic thresholds over time.
Validity evidence for the stress assessment by the IGAT was less impressive than the assessments of behavior. Although the negative affect scale used in this study only correlates moderately strong with measures of perceived stress, and the strengths of this relationship observed in this study were only slightly less, the test–retest reliability of stress was weak. Future research should use a more appropriate measure of perceived stress to establish concurrent validity alongside refining the description of “stress” in the IGAT to improve its test–retest reliability. In the current implementation, “stress” was left to the respondent to define.
Limitations
Other limitations are also worth noting. First, the sample size fell slightly below the a priori goal for many individual tests as a result of missing data due to attrition. While ICCs were all statistically significant and would likely not have been affected by more data, it is possible that more mean differences would have emerged. Thus, there is an increased risk of Type II error among the mean difference analyses. Second, this study suffers from many of the limitations associated with internet-based research. While all of the participants were screened over the telephone prior to entry into the study, their subsequent participation took place over the internet. Participants may have been distracted at times while completing assessments, as the assessment environment could not be controlled. Not meeting face-to-face with participants limited our ability to make behavioral observations to “fact check” screening information, such as body weight. However, as the maximum reimbursement was $40 for participating in this 12-week study, the incentive to fabricate information to gain entry into the study was likely low. Third, although participants completed the IGAT in a relatively short period of time given the great deal of information they provided, and attrition was low, the IGAT requires substantial attentional resources to complete and may not be tolerated well by all. Feedback from focus groups early in the testing of the tool supported that it is not only challenging but also generally well tolerated. Even so, it is unknown how much of the attrition in this study may have been due to participants finding the IGAT too difficult to complete. Fourth, a definition of binge eating consistent with that in the DSM-5 was provided to participants when they completed the IGAT, but an example of a binge eating episode was not provided, which research suggests improves the accuracy of self-reported binge eating (Goldfein, Devlin, & Kamenetz, 2005). Future research should experiment with providing such an example with the IGAT. Future research may also investigate whether the IGAT performs differently in certain diagnostic groups, which this study was not powered to test. Finally, the assignment of verbal descriptors to ranges of ICCs representing the test–retest reliability of the IGAT may become less accurate in the future as better measures are created. As the state of the field changes, so too should the evaluation of what amount of agreement is “substantial.”
Conclusions
The IGAT appears to be a potentially useful assessment tool for gathering data at the week level over the previous 12 weeks using body weight and symptom frequencies for binge eating, purging, fasting, and exercise in this study. How it performs with other features (e.g., self-harm behaviors, drinking or drug use episodes, panic attacks, etc.) is unknown and should be the subject of future research. The IGAT is not a substitute for prospective data collection, and it does not collect data at the daily or momentary level, as in EMA. When momentary or day-to-day fluctuations are of interest for more fine-grained analysis of the temporal relationship between things such as mood and behavior, EMA is more appropriate. However, in applications for which prospective data collection is impossible (e.g., when patients first present for treatment and the interest is in their recent past) or impractical (e.g., when an uninterrupted view of symptom frequencies over the course of many months or years using multiple follow-ups is the goal), the IGAT may provide a reasonable estimation of symptom course.
The IGAT may provide advantages over conventional retrospective measures due to its efficiency at collecting a great deal of information, ease of administration (i.e., over the Internet), and cost-effectiveness. Nevertheless, it would benefit from continued development aimed at identifying which components effectively aid in the collection of reliable and valid data and which are superfluous or even interfere with this process. Those components identified as effective could then be enhanced to maximize the utility of the instrument, and those that are not helpful removed.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Funding provided in part by The University at Albany Graduate School’s Dissertation Research Fellowship, the Psychology Department’s Edward Blanchard Dissertation Award, the Graduate Student Association, and the Benevolent Association.
