Abstract
The current study evaluated the accuracy of the Structured Interview of Reported Symptoms, Second Edition (SIRS-2) in a criterion-group study using a sample of forensic psychiatric patients and a community simulation sample, comparing it to the original SIRS and to results published in the SIRS-2 manual. The SIRS-2 yielded an impressive specificity rate (94.3%) that exceeded that obtained using the original SIRS scoring method (92.0%) and approached that observed in the SIRS-2 normative data (97.5%). However, changes in scoring resulted in markedly lower sensitivity rates of the SIRS-2 (36.8% among forensic patients and 66.7% among simulators) compared with the SIRS (47.4% and 75.0%, respectively). The removal of the Total Score from the SIRS-2 further hindered identification of feigning. Analyses also evaluated the additive value of the new RS-Total and MT Index scales in the SIRS-2. Implications of these results for forensic psychologists are discussed.
The Structured Interview of Reported Symptoms (SIRS; Rogers, Bagby, & Dickens, 1992) has been labeled the “premier measure for the assessment of feigned mental disorders” (Rogers, Sewell, & Gillard, 2010, p. 1) and the “gold standard” among measures of feigning (Boccaccini, Murrie, & Duncan, 2006). Furthermore, several surveys have characterized the SIRS as the most well-known (Boccaccini et al., 2006) and most frequently administered (Archer, Buffington-Vollum, Stredny, & Handel, 2006) measure of symptom exaggeration. Comprised of 172 items, the SIRS assesses multiple strategies of feigning, tapping both unlikely symptoms (i.e., those atypical or improbable among genuine psychiatric patients) and amplified symptoms (i.e., realistic complaints that are reported with exaggerated frequency or severity; Rogers, Payne, Berry, & Granacher, 2009).
Original SIRS validation studies used simulation and criterion-group designs, as well as clinical and nonclinical samples, yielding an overall specificity rate of 99.5% and sensitivity rate of 48.5% (Rogers et al., 1992). Results of a recent meta-analysis of the SIRS yielded large effect sizes for averaged primary scales (d = 1.53) and total scores (d = 2.02), demonstrating the measure’s effectiveness in differentiating feigning from genuine responding (Green & Rosenfeld, 2011). However, classification accuracy rates differed significantly from those described in the SIRS manual. For example, findings from the meta-analysis indicated that the average specificity rates for studies published after the SIRS was commercially released was 78.1% (range = 58% to 100%) and were, on average, significantly lower in clinical samples compared with nonclinical samples (73.3% vs. 94.1%, respectively). Although classification accuracy of genuine respondents was much lower than reported in the SIRS manual, the SIRS identified suspected feigners and simulators more often than was found in the original validation studies. Overall, sensitivity rates across studies published between 1992 and 2010 averaged 70.7% (range = 46% to 100%), compared with a sensitivity rate of 48.5% reported in the SIRS manual. Furthermore, the SIRS more accurately identified feigning in samples from simulation designs than those obtained in criterion-groups designs (74.4% vs. 62.3%, respectively).
Rogers and colleagues (Rogers et al., 2010) recently published a second edition of the measure (the SIRS-2). The SIRS-2 retains the identical questions, administration format, and division of items into primary and supplementary scales as did its predecessor but uses modified scoring and classification methods. Specifically, a respondent is classified as feigning on the SIRS when he or she scores in the Definite feigning range on one primary scale or in the Probable feigning range on three primary scales. In Indeterminate cases (i.e., when a respondent scores in the Probable feigning range on one or two scales and/or in the Indeterminate range on multiple scales), the Total Score can also be used to classify feigning. In contrast, classification of examinees using the SIRS-2 involves a flow chart of decision rules that results in multiple possible outcomes (Genuine, Feigning, and three categories of Indeterminate classification). A classification of feigning requires elevations on primary scales and on either the newly introduced Rare Symptoms Total (RS-Total) scale or the Modified Total Index (MT Index; a variation on the Total Score used in the original SIRS). In addition to more stringent criteria for classifying feigning, the SIRS-2 divides the Indeterminate grouping of the SIRS into Indeterminate-Evaluate, Indeterminate-General, and Disengaged, based on the score of other indices (e.g., MT Index and Supplementary Scale Index). According to the SIRS-2 manual, more extensive assessment is advised for those cases in the Indeterminate-Evaluate range, as this group has a base rate of feigning exceeding 50% (Rogers et al., 2010); the Indeterminate-General is less likely to indicate feigning (with an estimated base rate of 34.3%).
The SIRS-2 manual provided updated normative data for the primary scales using original SIRS validation research (i.e., conducted prior to 1992) and research conducted after the publication of the SIRS. However, validation data for the new scoring guidelines of the SIRS-2 were derived from only 522 cases; 314 cases drawn from the original validation studies and 208 cases drawn from a sample of “multiply traumatized inpatients” (Rogers et al., 2010, p. 37). The SIRS-2 manual reported an impressive specificity rate (97.5%). However, the classification accuracy rates reported in the SIRS-2 manual excluded 120 examinees (23% of the total sample) that were classified as Indeterminate, resulting in a sensitivity rate of 80% compared with 49% if Indeterminate cases were considered unidentified (DeClue, 2011). Thus, sensitivity rates are overestimated when calculated without including Indeterminate cases, as only the most blatant cases of genuine and feigning presentations are retained for analyses (Rubenzer, 2010).
DeClue (2011) and Rubenzer (2010) noted several additional concerns regarding the SIRS-2. For example, they expressed concern regarding the normative sample, and particularly the predominance of “presumed genuine” highly symptomatic patients, half of whom were diagnosed with Dissociative Identity Disorder. Given the low rate of Dissociative Identity Disorder in forensic settings, the generalizability of decision rules based on this sample is questionable (particularly in the absence of any cross-validation of new scales, indices, or decision rules). Further concerns raised by DeClue (2011) and Rubenzer (2010) included the small sample size (n = 36) of suspected feigners and lack of clarity about how criterion groups were formed.
Given the evidence that the original validation data for the SIRS differed from the accuracy rates found in subsequent research, independent research using the SIRS-2 is needed to provide necessary cross-validation data regarding the classification accuracy of the SIRS-2. In addition, little research to date has directly compared classifications made by the SIRS-2 with the original SIRS, for which far more extensive data are already available. The current study sought to fill this void, using both criterion and simulation groups in order to minimize inherent limitations in each of these designs (see Rogers & Gillard, 2011).
Method
Participants
The current study involved a combination of criterion-group and simulation designs, with samples of pretrial criminal defendants and community members, respectively. Archival records of 144 pretrial criminal defendants consecutively admitted to Kirby Forensic Psychiatric Center in New York City for restoration of competency to stand trial between July 2008 and June 2009 were reviewed for this study. Approximately one in five patients (n = 30; 20.8%) who were admitted to this facility were excluded from analysis because they were not administered the SIRS, typically because they did not speak English (n = 10; 6.9%) or had refused or were too symptomatic to complete psychological testing (n = 14; 9.8%). The final sample size included 114 patients; 79.2% of those admitted to the maximum-security forensic hospital.
The patient sample was predominantly male (n = 101; 88.6%) and averaged 37.0 years of age (SD = 11.1; range = 18-65 years). Most were ethnic minorities, with 78 (68.4%) identifying themselves as Black/African American, 23 (20.2%) as Hispanic, and 1 (0.9%) as Asian. The remaining 12 patients (10.5%) identified as Caucasian. On average, patients reported 10.8 years of education (SD = 2.5 years; range = 6-16 years). Treating psychiatrists diagnosed most patients at the time of testing with either a psychotic disorder (n = 85; 74.6%) or a mood disorder (n = 24; 21.1%). Three patients (2.6%) were diagnosed with a substance abuse or dependence disorder without a comorbid Axis I diagnosis, and two patients (1.8%) were deemed to have no diagnosis and were labeled as malingering. Most patients reported lengthy histories of mental illness, including multiple previous hospitalizations (M = 6.7, SD = 11.1; range = 0-75) and computerized criminal history reports revealed numerous arrests (M = 13.6, SD = 17.5; range = 1-145).
Community members who were over the age of 18 years and fluent in English were eligible to participate in the simulation portion of the study. Males were deliberately oversampled to maintain a similar gender composition as found in the forensic sample. The demographic and background characteristics of the community participants also mirrored those of the forensic sample in terms of ethnicity, socioeconomic status, and history of prior arrests. Of the 41 community members who provided informed consent to participate in the study, 5 participants were excluded after they reported that they had failed to follow the study instructions to malinger incompetency to stand trial for several of the measures. Therefore, the final analyses included data from 36 community participants.
The majority of community participants were male (n = 32; 88.9%) and ranged in age from 18 to 74 years (M = 35.9, SD = 15.8). Almost half identified their ethnicity as Black/African American (n = 17; 47.2%), with 7 (19.4%) Caucasian, 7 (19.4%) Hispanic, and 4 (11.1%) Asian; 1 participant (2.8%) identified as being of mixed ethnicity. All the participants reported that they had completed at least 11 years of education (M = 14.2; SD = 1.8). More than one quarter (n = 10; 27.8%) of community participants stated that they had been diagnosed with a mental disorder in the past and 30.6% (n = 11) reported that they had been previously arrested. The community participants did not significantly differ from the forensic patient sample in terms of gender, age, or ethnicity. However, participants in the community sample had completed significantly more years of education, t(146) = 9.08, p < .001, and had significantly fewer arrests, t(148) = 7.20, p < .001, and were less likely to have a mental disorder diagnosis than the patient sample, χ2(1, N = 150) = 99.60, p < .001.
Procedure
Within a few weeks of hospitalization (M = 2.7 weeks; SD = 1.8), all patients admitted to the forensic hospital were referred for psychological testing to assist in clarifying diagnoses and treatment planning. Testing included a standardized battery of measures, including the SIRS, administered to all patients undergoing an evaluation of competency to stand trial. Results of some of the additional measures administered are described in recent publications (e.g., Pierson, Rosenfeld, Green, & Belfi, 2011; Stimmel, Green, Belfi, & Klaver, 2012). Prior to testing, treating psychiatrists informed clinicians (psychologists or psychology graduate students) conducting the evaluations (via questions contained on a referral form) whether or not they believed patients to be feigning at least some of their symptoms. Psychiatrists provided dichotomous ratings of “genuine” or “feigning,” as well as ratings along a 7-point scale of the degree of suspected feigning (1 = no indication of malingering, 4= unclear; some problems appear exaggerated or fabricated/others appear genuine, 7 = completely malingering). For the purpose of analyses, psychiatrists’ dichotomous ratings were used. Psychiatrists’ opinions were based on multiple sources of information including their clinical observation of patients’ behavior while in the hospital, intake and follow-up interviews, mental status examinations, social history, available psychiatric records, legal information (e.g., previous competency evaluations and computerized criminal history reports), and input from other members of the treatment team. In short, psychiatric judgments were based on all available information except current psychological testing. Because these “preliminary” impressions reflect an inherently imperfect standard, we also examined the reports regarding competency to stand trial conducted prior to patients’ discharge (hereafter referred to as predischarge reports). As reported below, records from eight patients were removed because of discrepancies between the opinions of the treating psychiatrist and the predischarge report. The latter are, of course, potentially confounded by the results of the SIRS since clinicians may have changed their opinion about the veracity of the patient’s presentation after learning the results of this or other psychological tests. Additionally, it is possible that some patients modified their presentation of symptoms over the course of hospitalization. However, these data provide an important check on the accuracy of clinician assessments of the genuineness of the patient’s clinical presentation and offer a rare perspective on the ecological utility of psychological test data.
Community simulators were administered the same set of measures as the patients, using standardized instructions. However, they were informed that they would receive a monetary incentive of $40 if they completed all the measures. They were also offered an additional incentive to optimally simulate incompetence to stand trial to enhance the external validity of the simulation (Rogers, 1997). Specifically, they were told that they could earn an additional $100 if they convincingly feigned symptoms of a severe mental illness but evaded detection by the measures administered. They were then provided with a lay-person’s explanation of symptoms and behaviors that might render a defendant incompetent to stand trial and instructed to fake incompetency. A copy of the script used can be obtained by request from the first author.
Forensic psychiatric patients and community simulators were administered a battery of measures including the SIRS. The SIRS was scored using the original cutoffs and guidelines and then, for the purpose of the current analyses, rescored using the SIRS-2 decision rules. To ensure comprehension of and compliance with all instructions, community participants were administered a poststudy questionnaire. They were asked to summarize the instructions that they were given and, using a numerical rating scale (1 = no effort to 5 = maximum effort), were asked to rate the amount of effort that they exerted to follow these instructions. All participants were able to accurately summarize the instructions. However, one participant was removed from further analyses after reporting an effort rating lower than three out of five and four participants were removed after reporting that they had forgotten the instructions for at least part of testing.
Statistical Analyses
The first set of descriptive analyses compared the proportion of forensic patients classified by the SIRS and SIRS-2 as Genuine, Indeterminate, and Feigning. Because Genuine and Indeterminate classifications were not explicitly defined in the original SIRS, we used guidelines described in the test manual. Specifically, Genuine respondents were identified as those who obtained six or more primary scale elevations in the “Honest” range, whereas Indeterminate classifications were used for those individuals who did not fall into either the Genuine or Feigning groups. Group classifications based on the SIRS-2 followed the decision rules outlined in the test manual. Criterion group determinations were based on a combination of psychiatrists’ referrals as well as the clinical/diagnostic impressions communicated in predischarge reports (described in more detail below). These classifications served as the criterion for identifying genuine and suspected feigning patients. Sensitivity and specificity rates were calculated for both scoring methods (SIRS/SIRS-2). Because traditional measures of classification accuracy require a dichotomous classification (genuine vs. feigning), accuracy rates for the SIRS-2 were calculated first by collapsing those classified as Indeterminate-Evaluate, Indeterminate-Genuine, Disengaged, and Genuine classification categories into Genuine and all others into Feigning. However, given the estimated 50% base rate of feigning in the Indeterminate-Evaluate group (Rogers et al., 2010), accuracy rates were then recalculated by combining the Indeterminate-Evaluate patients with those classified as Feigning on the SIRS-2. Chi-square goodness of fit analyses were used to compare results of the SIRS-2 from the current study to that reported by Rogers et al. (2010) in the SIRS-2 manual and to calculations provided by DeClue (2011), which included Indeterminate cases in the calculation of sensitivity and specificity. Identical analyses were conducted with the simulation sample. Finally, exploratory analyses were conducted to determine the additive value of the RS-Total scale and the MT Index of the SIRS-2.
Results
More than one fifth (n = 23; 20.2%) of forensic patients were suspected of feigning symptoms by their treating psychiatrist prior to the initiation of psychological testing; 91 patients (79.8%) were believed to be genuinely symptomatic. All patients underwent an evaluation to determine whether they were competent to stand trial (and thus could be discharged) after testing was completed (M = 8.8 weeks; SD = 11.3). The time elapse between testing and evaluation of competency was unrelated to whether treating psychiatrists believed patients to be genuine or feigning, t(112) = 1.03, p = .31. Overall, the predischarge reports matched treating psychiatrists’ classifications of genuine versus suspected feigning in 106 (93.0%) cases. Specifically, reports gave no indications of exaggeration or fabrication of symptoms in 87 (of 91) cases considered by psychiatrists to be genuine. Furthermore, predischarge reports corresponded with psychiatrists’ suspicions of feigning in 19 (of 23) cases. Thus, subsequent analyses included these 106 patients for whom classification was agreed on by treating psychiatrists and independent evaluators; the remaining 8 cases where initial impressions and predischarge reports conflicted were omitted because these cases may have been confounded by the results of psychological testing (and the SIRS, in particular). The predischarge reports of suspected feigners indicated that treatment staff suspected patients of feigning psychiatric symptoms in nine cases, a combination of psychiatric symptoms and cognitive deficits in eight cases, cognitive deficits in one case (this patient was identified as feigning by the SIRS and SIRS-2), and unspecified symptoms in one case.
Classification of Forensic Patients
Regardless of whether they were believed to be genuinely symptomatic or were suspected of feigning symptoms, forensic psychiatric patients were less likely to be classified as Genuine using the scoring method of the SIRS (n = 66; 62.3%) than the SIRS-2 (n = 76; 71.7%), χ2(1, N = 106) = 42.64, p < .001. As a result, more patients were classified as Indeterminate or Feigning by the SIRS (n = 24, 22.6%, and n = 16, 15.1%, respectively) than the SIRS-2 (n = 18, 17.0%, and n = 12, 11.3%, respectively). See Table 1 for concordance and discrepancy rates between the SIRS and SIRS-2. Of those classified as Indeterminate by the SIRS-2, seven were classified as Indeterminate-Evaluate, eight were classified as Indeterminate-General, and three were classified as Disengaged.
Concordance and Discrepancies Between SIRS and SIRS-2 Classifications of Pretrial Forensic Patients
Note. SIRS = Structured Interview of Reported Symptoms. Bolded values indicate concordance between SIRS and SIRS-2 classifications. SIRS classification: Feigning = 1 scale in the Definite range or ≥3 in the Probable range; Genuine = ≥6 scales in the Honest range; Indeterminate = remaining cases. SIRS-2 classification: Decision tree.
All Indeterminate cases were combined with Genuine classifications to initially calculate accuracy rates of the SIRS and SIRS-2 (i.e., distinguishing those classified as Feigning based on the SIRS/SIRS-2 from all others). The SIRS scoring method yielded a lower specificity rate than the SIRS-2 (92.0%, 95% confidence interval [CI] = 84.1%, 96.7%; and 94.3%, 95% CI = 87.1%, 98.1%, respectively); however, the overlapping confidence intervals demonstrate that the difference in specificity rates was not significant. However, as displayed in Table 2, sensitivity rates for both scoring methods were below 50% (47.4% for the SIRS, 95% CI = 24.5%, 71.1%; and 36.8%, 95% CI = 16.3%, 61.6% for the SIRS-2) when only cases clearly identified as feigning were considered indicative of accurate detection. When the forensic patients classified as Indeterminate-Evaluate were combined with those classified as Feigning on the SIRS-2, sensitivity increased substantially (from 36.8% to 52.6%, 95% CI = 28.9%, 75.6%); however, specificity also decreased, from 94.3% to 89.7% (95% CI = 81.3%, 95.2%).
Classification Accuracy of the SIRS Versus SIRS-2 in the Pretrial Defendant Sample
Note. SIRS = Structured Interview of Reported Symptoms.
SIRS Total Score ≥76 is used to classify suspected feigning in Indeterminate cases.
According to the SIRS manual, a Total Score cutoff can be employed in Indeterminate cases (i.e., when an examinee has primary scales in the Indeterminate and Probable ranges, but does not meet the threshold for feigning). To evaluate the additive value of the Total Score, we calculated the accuracy rates of the SIRS using either scale elevations or, in the case of Indeterminate classifications, the Total Score to identify feigning versus genuine responding. This approach identified six additional patients suspected of feigning (over that identified by scale elevations alone), increasing sensitivity to 78.9% (95% CI = 54.4%, 94.0%). However, specificity decreased to 88.5% (95% CI = 79.9%, 94.4%).
Chi-square analyses indicated that classification of the two patient groups (excluding Indeterminate cases) in the current study differed significantly from classification rates reported in the SIRS-2 manual, χ2(3, N = 88) = 8.71, p = .03; these results largely reflected the different sensitivity rates (58.3% in the present study vs. 80.0% in the SIRS-2 manual, with Indeterminate cases excluded; see Table 2). However, when these data were contrasted with calculations presented by DeClue (2011), which included the entire SIRS-2 validation sample (without excluding Indeterminate cases), results of the current study no longer differed significantly from the SIRS-2 normative data, χ2(3, N = 106) = 4.86, p = .18.
Classification of Community Simulators
Consistent with findings from the patient sample, community simulators were more likely to be classified as feigning using the original SIRS scoring algorithm (n = 27; 75.0%; 95% CI = 57.8%, 87.9%) compared with the SIRS-2 (n = 24; 66.7%; 95% CI = 49.0%, 81.4%), although not significantly so (based on the overlapping confidence intervals). However, there was no difference in the proportion of simulators classified as Genuine by the two methods (n = 5; 13.9%); these five “false negative” classifications were the same individuals. Of note, all five of these simulators were identified as feigning on others scales administered, suggesting that they were indeed “false negatives” on the SIRS/SIRS-2. Fewer simulators fell in the Indeterminate categories when scored using the SIRS decision rules (n = 4; 11.1%) compared with the SIRS-2 algorithm (n = 7; 19.5%). Only two of the simulators were classified as Indeterminate-Evaluate by the SIRS-2 (the category thought to be more likely indicative of feigning); the remaining five simulators classified as Indeterminate fell into the Indeterminate-General category.
Chi-square analyses indicated that the sensitivity rate (excluding Indeterminate cases) of the SIRS-2 among simulators (82.8%; 95% CI = 64.2%, 94.2%) was not significantly different from results reported in the SIRS-2 manual, χ2(1, N = 29) = 0.14, p = .71. However, the rate was significantly higher than expected based on DeClue (2011), χ2(1, N = 36) = 4.41, p = .04. Both classifications based on the original SIRS (primary scale elevations and Total Score) yielded higher sensitivity rates (75.0% and 83.3%, respectively) than those obtained using the SIRS-2 when Indeterminate cases were retained (i.e., Indeterminate cases where collapsed with Genuine classifications).
Exploratory Analysis of the SIRS-2 Decision Tree
The SIRS-2 scoring method introduces two new scales intended to improve classification accuracy over interpretation of primary scale elevations alone. Figure 1 details the proportion of patients and simulators classified at each branch of the SIRS-2 decision tree. The added threshold of the RS-Total, intended to reduce false positive classifications, eliminated two of seven genuine patients with elevated primary scales (i.e., eliminated two “false positives”). However, two patients suspected of feigning and three simulators who exceeded the threshold for primary scale elevations did not obtain scores above the cutoff on the RS-Total scale and were eventually classified as either Indeterminate or Genuine. Thus, the addition of the RS-Total reduced false positive classifications but at an obvious cost in reduced sensitivity.

SIRS-2 decision tree
Scores on the MT-Index are reviewed for those cases in which equivocal evidence of feigning exists (i.e., those participants with one Definite or three Probable elevations who did not elevate the RS-Total scale and those with one or two scales in the Probable feigning range). As displayed in Figure 1, the MT Index did not categorize any participants as likely feigning, suggesting that this index may not adequately identify feigners who were not identified by the scale elevations and an elevated RS-Total score.
Discussion
The publication of the SIRS-2 introduced revised algorithms and updated normative data for this widely used instrument. These revisions are timely, particularly in the context of an ever-growing number of studies that rely on the SIRS for validation of alternative measures of feigning (an approach Rogers refers to as “bootstrapping”; Rogers & Gillard, 2011). Moreover, forensic evaluators require reliable accuracy rates to inform clinical opinions. For example, a recent meta-analysis indicating that classification accuracy rates in studies conducted since the SIRS was made commercially available 20 years ago have diverged significantly from the original validation studies (Green & Rosenfeld, 2011). Despite the need for robust accuracy data, the results published in the SIRS-2 manual have yet to be cross-validated in samples that were not used to develop the revised algorithms. In the current study, SIRS-2 classifications were compared with the results of the SIRS and to the findings reported in the SIRS-2 manual. Overall, the results are mixed.
Among hospitalized forensic psychiatric patients, the SIRS-2 yielded an impressive specificity rate (94.3%) that exceeded that obtained using the SIRS scoring method (92.0%). Furthermore, accuracy exceeded average findings of studies of the SIRS conducted since 1992 (e.g., Green & Rosenfeld, 2011). Although the specificity rates obtained in the current study include genuine patients classified as Indeterminate, the majority fell in the Genuine range (72.4% of genuine responders for the SIRS classifications and 81.6% based on the SIRS-2). These rates, however, are substantially lower that the specificity rate described in the SIRS-2 manual (97.5%). In contrast to impressive specificity rates, neither scoring method was particularly effective at identifying feigning, with lower sensitivity rates among patients observed using the SIRS-2 (36.8%) than the SIRS (47.4%). Both versions yielded higher sensitivity rates among simulators than patients suspected of feigning (66.7% vs. 75.0%, respectively). These findings indicate that detection of feigning may be expected to increase in settings where examinees have less experience with severe psychopathology, as might be true of correctional institutions.
Differences in sensitivity rates between the SIRS-2 and SIRS were even more pronounced when patients’ and simulators’ Total Scores were incorporated. The application of the Total Score in Indeterminate cases (examinees who scored in the Probable feigning range on one or two scales or in the Indeterminate range on multiple scales) identified 30% more forensic patients who were suspected of feigning than when identification was based solely on scale elevations, yielding a sensitivity rate of 78.9%. Likewise, inclusion of the Total Score to identify feigning increased sensitivity among simulators, albeit to a lesser extent than among the suspected feigning patient sample (i.e., an increase of only 3% for the SIRS and 16% for the SIRS-2 to 83.3%). These findings are also consistent with results from Green and Rosenfeld’s (2011) meta-analysis of the SIRS, in which suspected and simulating feigners were more accurately identified in studies using the Total Score compared with the primary scale elevations. In the current study, the trade-off for using the Total Score was slight, with a reduction in specificity from 92.0% (for the SIRS approach) and 94.3% (using the SIRS-2 scoring algorithms) to 88.5%. These results suggest that the removal of the Total Score from classification decisions in the SIRS-2 may be premature.
The SIRS-2 essentially replaced the Total Score with the MT Index. However, in the current study, the MT Index did not assist in identifying any patients suspected of feigning or simulators. Future research should determine whether the MT Index is useful in differentiating feigners from genuine responders, as well as identifying optimal cut-offs for these classifications.
The SIRS-2 also introduced the RS-Total scale to reduce false positive classifications. Results from this study indicated that the addition of the RS-Total scale served this purpose, reducing the false positive rate among genuine patients (from 8.0% to 5.7%). However, the “cost” was substantial, as five feigners (two psychiatric patients and three community simulators) were misclassified as Indeterminate or Genuine based on the RS-Total scale. Given that no additional questions are needed to score this scale, even a small reduction in false positive classifications is likely worthwhile given the significant adverse implications of erroneously labeling a genuinely mentally ill examinee as feigning.
The high specificity rate of the SIRS-2 in the current study was only slightly lower than that reported in the SIRS-2 manual (97.5%), provided Indeterminate classifications were treated as indicative of genuine responding rather than feigning. This finding is all the more striking given the marked differences in sample characteristics between the normative data and the current study. Specifically, the SIRS-2 manual indicates that half of the genuine sample was diagnosed with Dissociative Identity Disorder (DID). In light of previous research on the SIRS that indicates low accuracy rates among DID patients (e.g., 65%; Brand, McNary, Loewenstein, Kolos, & Barr, 2006), it is notable that these results generalize to the forensic psychiatric population assessed in the current study (most of whom were diagnosed with a psychotic disorder).
Despite providing some support for the efficacy of the SIRS-2 in classifying genuine patients, the results of the current study did not substantiate Rogers et al.’s (2010) findings regarding sensitivity. Several factors may account for this disparity. Some degree of shrinkage is expected in cross-validation research. This may be even more likely given the small sample size of suspected feigners (n = 36) in the SIRS-2 normative data and lack of clarity about how criterion groups were identified in the studies that formed the basis for the normative data (DeClue, 2011; Rubenzer, 2010). The differences in observed sensitivity rates may also be an artifact of how accuracy rates were calculated, as Rogers and colleagues excluded 120 examinees (23% of all cases) that were classified as Indeterminate. This decision to exclude all but the clearest cases greatly inflates the apparent sensitivity of the instrument (reported as 80% in the SIRS-2 manual), although the present study found far lower rates of detection even when this same approach was used (sensitivity of 52.6% among psychiatric patients suspected of feigning). When Indeterminate cases were retained, results of the current study appear to match DeClue’s (2011) recalculations of the SIRS-2 normative data (when forensic patients formed the criterion group).
It is unclear how evaluators (or researchers) should proceed when faced with Indeterminate SIRS/SIRS-2 profiles. In the current study, rates of Indeterminate classifications in the patient and simulation samples were lower (17.0% and 19.5%, respectively) than those reported in the SIRS-2 manual (23%). Although researchers have begun looking at optimal methods of combining measures of feigning in psychiatric settings (e.g., Rosenfeld, Green, Pivovarova, Dole, & Zapf, 2010), results remain preliminary, providing little guidance as to which measures would best enhance classification accuracy over the SIRS-2. More typically, the SIRS is recommended as a follow-up to be given in cases of possible feigning as identified by screening measures. Consequently, research is needed to identify measures that accurately supplement SIRS-2 results, particularly in ambiguous cases. Furthermore, additional data are needed to understand possible differences between genuine responders and suspected feigners who are classified as Indeterminate, to aid in interpretation of those who fall in this category.
Despite the strengths of the current study in evaluating the new scoring methods of the SIRS-2, these results are potentially limited by the study design. Specifically, any study that uses a criterion-groups methodology to classify individuals as feigning versus genuine is subject to an unknown amount of error. In this study, we aimed to minimize any such errors by requiring consensus between psychiatrists’ classifications (made at least 1 week after the start of treatment with patients) and independent evaluators’ classifications (made prior to discharge). Although the latter ratings may have been influenced to some extent by test results, reports often provided detailed examples of feigning and genuine clinical presentations, derived from hospital progress notes and observations made during the predischarge evaluations. Furthermore, any errors using this consensus approach would influence observed accuracy rates of both the SIRS and SIRS-2, and thus should not affect comparisons between the two scoring approaches. Limitations of the criterion-groups design were also addressed through inclusion of a simulation sample that was comparable to the patient sample in terms of age, gender, and ethnicity (and included multiple individuals with psychiatric and criminal histories). The inclusion of this latter sample also increases generalizability to populations with less severe (or no) psychiatric illness, as might be expected in correctional institutions.
In accordance with the American Psychological Association’s (APA’s) Ethical Guidelines (APA, 2002), psychologists are advised to adopt updated testing measures within one year of their publication. Thus, it may be assumed that clinicians will be expected to implement the scoring methods of the SIRS-2 within the current calendar year. However, both the APA guidelines (in Standard 9.02(b)) and the Specialty Guidelines for Forensic Psychologists (in Guideline 10.02) limit use of “assessment instruments [to those] whose validity and reliability have been established for use with members of the population assessed.” Furthermore, both of these ethical guidelines stipulate that “when such validity or reliability has not been established, psychologists describe the strengths and limitations of test results and interpretation” (APA, 2002). The limited settings in which the SIRS-2 has been validated, coupled with differences between sensitivity rates observed in the current study versus those published in the manual, raises concerns for both the reliability and validity of the SIRS-2 scoring algorithms. The current results and recent critical reviews of the SIRS-2 by DeClue (2011) and Rubenzer (2010) suggest that it may be premature to replace the SIRS without further study, or in keeping with ethical practices, without acknowledgement of the outstanding limitations of the SIRS-2. Furthermore, additional research is warranted before the Total Score is excluded from interpretation; studies using the SIRS Total Score have yielded higher classification accuracy rates than those applying the SIRS scoring algorithms (Green & Rosenfeld, 2011). Finally, given the unique context in which the current study was conducted (i.e., with forensic patients previously found incompetent to stand trial and subsequently hospitalized for treatment), additional research is necessary in settings where rates of feigning, as well as the blatant nature of feigned presentations might be expected to be higher (e.g., in initial evaluations of insanity, presentencing, and/or initial evaluations of competency).
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The first author received funding from the American Psychology and Law Society (Grants-in-Aid) and the American Academy of Forensic Psychology (Dissertation Award) for this research.
