Abstract
Conducting behavioral analysis and administering questionnaires at the start of psychotherapy is time consuming and costly. This study explores a new approach linking diagnostic information across disorders to behavioral response categories (BRCs) used in behavioral analysis addressing the research gap concerning the integration of behavioral analysis into questionnaire-based diagnostic methods. Clinical psychologists assigned SCL-90-R items to one of four response variables (emotional, cognitive, physiological, behavioral). Three groups emerged: items unequivocally assigned to one BRC, items assigned to multiple BRCs (combined Mvar = 296.15), and items that couldn't be assigned due to expert disagreement, ambiguous formulations, or enumerations in items (Mvar = 109.97). The feasible approach enables receiving multiple types of diagnostic information while only executing one questionnaire, potentially enabling cost-effective diagnostics, reducing therapy time, and making treatment more accessible to individuals with lower socioeconomic backgrounds. The findings emphasize the importance of semantic clarity in item formulations.
Introduction
Current State of the Art of Gathering Information in Clinical Psychology
In our daily lives we routinely gather, assess, and evaluate information from ourselves and others (Bröder and Hilbig 2017). In psychotherapy, we routinely examine and systematize human behavior as a foundation for decision-making in the selection of appropriate interventions to speed up the patient`s recovery (BMSGPK 2020).
Cognitive behavioral therapy (CBT), the most practiced form of psychotherapy, frequently employs behavioral analysis as a reliable method for obtaining treatment-relevant information (David et al. 2018; Hofmann et al. 2012). These standardized methods include patients’ symptoms, maintaining conditions, and behaviors. Among various behavioral analyses, the SORKC model (situation, organism, response, contingency, consequence), based on Kanfer and Saslow (1965), is the most widely used (Tuschen-Caffier and van Gemmeren 2018). A key aspect of SORKC is the individual’s response to a selected situation, encompassing emotional, cognitive, behavioral, and physical reactions.
Alongside behavioral analysis, standardized diagnostics, such as questionnaires, are used to assess patients’ perceptions. It is common to conduct a broad diagnostic process across disorders to obtain a comprehensive symptom overview and minimize overlooked symptoms, which may affect diagnosis and interventions (BMSGPK 2020). For economic reasons, most clinicians use validated questionnaires. In clinical practice, the “Symptom Checklist-90-R” (SCL-90-R) is frequently employed, as it offers a structured, concise overview of psychological distress (Franke 2002).
SCL-90-R: Statistical Quality
The SCL-90 assesses subjective impairments across 90 psychosomatic symptoms in the last seven days. Available in standard and revised versions (SCL-90-S/R), it has been used in more than 1,000 studies and translated into over 26 languages (Hildenbrand et al. 2015). It has received favorable international assessments and is widely valued by clinicians (Schmitz et al. 2000; Shafique et al. 2017). The 90 items form nine factors (somatization, obsessive–compulsive, interpersonal sensitivity, depression, hostility, anxiety, phobic anxiety, paranoid ideation, psychoticism) and one general factor (General Severity Index). Completion takes about 10–15 minutes (Franke 2002). Overall, international findings on validity, reliability and objectivity indicate high robustness and generalizability across populations (Carrozzino et al. 2023; Chen et al. 2017; Franke 2002; Sereda and Dembitskyi 2016).
Regarding reliability, the SCL-90-R shows high person separation (Elliott et al. 2006) and good-to-very-good test–retest reliability (Tomioka et al. 2008). The General Severity Index (GSI), reflecting all symptoms, consistently indicates psychological distress, explaining 28–39.7% of variance (Grande et al. 2014; Hardt et al. 2000; Hessel et al. 2001; Holi 2003; Sereda and Dembitskyi 2016). Recent results for all subscales and the GSI on internal consistency, ranging from good to very good, confirm earlier findings (Sereda and Dembitskyi 2016; Shafique et al. 2017; Siqveland et al. 2016). Convergent validity was found for depression, anxiety, somatization, and the GSI (Ardakani et al. 2016; Lee et al. 2018; Shafique et al. 2017).
Discriminant validity has been demonstrated for depression (Bianciardi et al. 2020), anxiety, interpersonal sensitivity (Bech et al. 2013), paranoid ideation, psychoticism (Chen et al. 2017) and the GSI (Shafique et al. 2017). The validity of the remaining subscales remains debated (Ardakani et al. 2016; Grande et al. 2014; Hildenbrand et al. 2015). The SCL-90 differentiates stress levels of inpatient and outpatient treatments (Ransom et al. 2010) and psychiatric versus community samples (Holi 2003; Müller et al. 2009; Ransom et al. 2010; Schmitz et al. 2000; Shafique et al. 2017; Tan et al. 2015).
In contrast, its factorial validity is debated. While some research teams formed new factors (Elliott et al. 2006), others found fewer than the original nine (Franke 2002; Grande et al. 2014). Some observed multidimensionality (Carrozzino et al. 2023), while most did not (e.g., Ardakani et al. 2016; Hessel et al. 2001), thus questioning the structure proposed by Derogatis et al. (1976). Franke (2002) speculated that multiple factors only emerge if the sample consists of individuals with distinct diagnoses. Alternatives to the SCL-90-R/S include the less researched VDS90 (Sulz and Grethe 2005) or instruments with fewer dimensions, e.g., the PHQ with five submodules (Spitzer et al. 1999), the GHQ-60 with six (Vázquez-Barquero et al. 1988), and its short versions GHQ-28/30 with four or fewer (Gnambs and Staufenbiel 2018; Shafer 2024).
Introducing Behavioral Analysis and the SORKC Model
Wilhelm Wundt (1874) was the first who classified human responses to situations into four categories: physiological, behavioral, cognitive, and emotional responses. Many decades later, Lindsley (1964) proposed the initial draft for standardized behavioral analysis (SRKC), and Kanfer and Saslow (1965, 1969) refined it to its current form (SORKC). Since then, the behavioral analysis model, known as “SORKC,” has consisted of five key components. The letter “S” signifies a “specific situation” or “stimulus” that leads to specific behaviors and consequences. “O” represents the “organism” and its relevant conditions, which interact with the specific situation and the corresponding behavior and consequences. The letter “R” represents all “responses” to a situation that can be divided into four aspects: physiological, behavioral, cognitive, and emotional. The letter “K,” which stands for “contingency,” represents the temporal proximity between responses and their corresponding consequences, serving as a connecting link between “R” and “C,” determining the predictive value of “R.” Finally, the letter “C” summarizes the relevant short- and long-term “consequences” of the responses.
SORKC Model: Applications
Behavioral analysis is widely used beyond clinical contexts to support behavioral exploration, differential diagnosis, prognosis, case conceptualization, intervention planning, and treatment evaluation in preventive, curative, and palliative scenarios (Radha and Raj 2020). Clinically, SORKC can assess whether physical or neurological findings affect reported psychological symptoms, including those linked to physiological correlates (Wilson 2002). In workplaces, SORKC has promoted behavior modifications such as work ability, problem-solving skills, daily structure, work-life balance, and networking (Hautzinger 2022).
Currently, no restrictions for applying SORKC or its components are known based on diagnosis (Böse 2007; Hennings 2020; Kuss and Lopez-Fernandez 2016; Linton et al. 1984; Malhotra et al. 2009; Schulz et al. 2017; Tretter and Loeffler-Stastka 2021; von der Heiden et al. 2019; Witthöft and Hiller 2010). Patients with complex symptoms such as derealization and depersonalization have benefited from its sparse, versatile structure, which allows therapists and patients to decide whether such phenomena should be treated as stimulus, cognitive, physiological, or emotional response, or as a consequence (Heidenreich et al. 2006). However, the SORKC model does not include semantic or general language models (Wilson 2002).
This study investigates a new assessment approach linking diagnostic information across disorders to behavioral response categories from behavioral analysis, thereby obtaining two types of diagnostic data with one procedure. The key question is in which cases this method outperforms behavioral analysis and standardized questionnaires in terms of time and suitability for patients.
It is important to streamline diagnostic procedures to make psychotherapy more efficient and accessible, particularly for those most in need but with limited financial resources. Therefore, the objectives of our study are as follows:
A. Establish a standardized procedure to connect questionnaire items referring to symptoms to the response categories of a behavioral analysis.
B. Investigate the potential use of various variables to estimate the suitability of an item for inclusion in this process.
C. Evaluate feasibility of our proposed method based on statistical data.
Methods
To explore the general implementation of the behavioral analysis response categories of SORKC into an already existing diagnostic tool, we integrated it into the well-examined SCL-90-R in an expert consensus mapping study.
Welding the R-variable of the SORKC Model to the SCL-90-R
During the initial phase of the investigation, we presented specialists in clinical psychology with the SCL-90-R items and requested that they assign each item to one of the four SORKC response variable categories. Multiple choices were not provided to the participants. This is depicted in Figure 1.

Process of linking items to behavioral response categories.
The sample comprised 40 psychologists, among whom 34 identified as female and six as male. The average age of the participants was M = 29.4 years (SD = 4.25). Seven of the experts possessed a master’s degree in clinical psychology, two had a diploma in psychology, 24 were psychotherapists in training, and five were already certified psychotherapists, leaving two experts with other clinical psychological backgrounds. Thirty-eight experts were from Europe, one was from the Middle East, and one was from Central Asia. The average length of relevant work experience among participants was M = 3.05 years, ranging from less than a year to 7.83 years.
The research was conducted online in the third and fourth quarter of 2023. To be eligible to participate, psychologists were required to have a background in clinical psychology and state that they would take the questionnaire seriously.
We applied a standardized procedure to classify items into behavioral analysis categories (emotional, cognitive, behavioral, and physiological). Given that no established framework exists for this type of classification, we developed a partly bottom-up evaluation approach (Figure 2). This means that the results of some items made additional aspects of the algorithm in Figure 2 necessary, which later allowed for the controlled handling of the remaining items. In the first step, items were screened for consistency in expert ratings to determine whether they exhibited sufficiently homogeneous classification patterns. When expert ratings indicated ambiguity across parameters, item wording and semantic structure were examined to clarify the behavioral response domain. Items that remained semantically ambiguous after this step were excluded from the classification.

Decision algorithm for linking items to BRCs based on experts in the field.
Analysis
Following the completion of the survey we employed the R software (R Core Team 2013; version 4.4.0) and its corresponding user interface RStudio (RStudio Team 2024; version 2024.04.1) to analyze the collected data. As we adopted an exploratory approach, we routinely examined whether the analyzed variables adhered to the assumption of a normal distribution (see Results section). Our analysis revealed that non-parametric bootstrapping was the appropriate method (Efron and Tibshirani 1994; Hesterberg 2015). To perform this method, we utilized the software provided by the R package “boot” (Canty et al. 2024).
Agreement Rate among Experts
The use of the two-thirds agreement criterion is based on qualitative research rather than quantitative research (Dixon 2011; McGann 2004). No preliminary sensitivity analysis was conducted on another sample to further examine the alternative fitting agreement rates for our case. There are items that could be expected to have an agreement rate of more than two-thirds, which in decision-making is usual and favorable (Dixon 2011; McGann 2004), but still deviates in terms of parameters from the mean. To reach a balance between a high consensus between experts without an excessive preemptive exclusion of items from the process, the two-thirds criterion was implemented and combined with the “deviation from descriptive or inferential parameters” criterion.
Results
To appraise the method of linking items to specific behavioral response categories, we will explore sample characteristics as well as agreement rates among experts and compare well-designed items to ambiguously designed ones by examining their statistical characteristics and scrutinizing the implications of applied criteria. Ambiguous items could not be considered for our analysis (see Figure 2). Our method proved to be feasible and low in cost.
Following the sequence depicted in the flowchart (Figure 2), in 67 out of 90 instances, 66% or more of the experts concurred with their evaluation of the behavioral analysis category for a particular item. The remaining 23 cases failed to reach a consensus on their evaluation. This meant that experts assigned items to multiple categories. Among these 23 cases, 19 items contained two clusters of words that were linked with connectives such as “and” or “if.” Following the flowchart (Figure 2), these items were assigned to both categories that the experts most frequently assigned the item to. The remaining four items contained two clusters of words that were linked with the connective “or,” thus they were excluded from being assigned to any behavioral analysis category. Two groups were formed: one group of items that were assigned mostly to one behavioral analysis category (consensus group) and one group of items that were assigned to more than one behavioral analysis category (non-consensus group). These groups were then analyzed separately. All Bonferroni corrected p-values of chi-square were significant with p < .0001 and ranged between 7.72 x10-26 and 2.19 x10-3. Further outcomes of the consensus subgroup examinations are illustrated in Table 1. The results indicate that heterogeneity was present even in the items experts agreed on, thus showing that further grouping into distinct behavioral response categories was still necessary.
Statistical results for the subgroup of 67 items agreed on by 66% of the experts.
Note. RIC = relative information content. Means, standard deviations, and confidence intervals were estimated using non-parametric bootstrapping (NPB; columns 2 and 3). NPB for all estimated means was performed with 10000 resamples. The simple difference between the two behavioral analysis categories that most experts agreed on was depicted, in contrast to the weighted cumulative difference (variance) per item (see row 2 versus row 5). The p-values of the chi-square statistics were adjusted using Bonferroni correction.
The following section summarizes the results for the 23 items in the non-consensus group: the mean variance was M = 109.97, SD = 6.62, with a 95% confidence interval of [96.30, 123.80], whereas the mean chi-square was M = 32.99, SD = 1.99, with a 95% confidence interval of [28.87, 37.12]. It is important to note that the bootstrapped means and their corresponding standard deviations were identical to the non-bootstrapped means and standard deviations for three or more decimal places. The p-values ranged from 0.197 to < .001. In total, only Items 23 and 71 received p-values higher than .05.
The combination of the “deviation from distribution or means” with the “two-thirds majority” criteria in the agreement rate among experts resulted in another distinct group of items. These are items with a two-thirds majority in the agreement rate but also a deviation from the distribution or mean. They showed higher mean of RIC with M = 0.63, with a 95% confidence interval of [0.58, 0.67], but still significant Bonferroni-corrected p-values for all chi-squared tests.
However, the 66% agreement rate among experts proved to be an insignificant factor in determining a deviation from the 95% confidence interval of the non-parametric bootstrapped means, regardless of the parameter chosen, if the item contained elements that did not suggest inseparable connectedness. The results of the variance (weighted difference between means and observed values due to squaring the difference), simple difference, relative information content, chi-square, and the Bonferroni corrected version of p-values of these chi-square tests (Table 1) all yielded the same conclusion for those cases: If less than 66% of experts agreed on an item and the item contained elements that did not suggest inseparable connectedness, all of these parameters deviated from the 95% confidence interval.
The Shapiro-Wilk test was conducted for each variable to evaluate the normality of the sample distribution for all variables. The results indicated that the data were not normally distributed, with a test statistic of W = 0.92 and p-values of < .001 for the variance and chi-square tests, and W = 0.15 and p < .001 for both the Bonferroni corrected and uncorrected p-values of the chi-square statistics. Additionally, the results showed a p-value of < .001 for the relative information content, with a test statistic of W = 0.90.
Adhering to all the procedures outlined in the flowchart, we arrived at the distribution of items per category, as shown in Table 2.
Number of item assignments based on item type.
Note. IC = connective suggesting inseparable connectedness of different words of an item e.g., “and” or “if.” Items with “or” connective were not assigned (hence missing in the “cumulative” row).
In the case of item 14, “Feeling low in energy or slowed down,” 63 “Having urges to beat, injure, or harm someone,” 67 “Having urges to break or smash things,” and 72 “Spells of terror or panic,” we were unable to allocate them to any specific category because of their phrasing, and thus they were omitted from the analysis.
The consequences of wording can be observed in their respective distributions. For item 14, experts were torn between the physiological (23) and behavioral categories (13). For item 63, the most selected categories were behavioral (26) and cognitive (9). For item 67, the most frequent categories were behavioral (25) and cognitive (8). And for item 72, the categories of emotional (20) and physiological (14) were most often chosen.
In the case of item 70, which reads, “Feeling uneasy in crowds, such as shopping or at a movie,” 62.5% of the experts agreed on its proper categorization. The distribution of the item was 25 times emotional, 8 times cognitive, and 7 times behavioral. This item exhibits connecting elements (“such as”) which made it difficult to determine whether the following examples were part of an intuitive enumeration of the premise, leading to disagreement among the experts regarding its category. Given the unique and ambiguous design of the item, we were forced to make a decision based on a literature review (e.g., Damasio 2003; Immordino-Yang and Damasio 2007) and assigned it to both the emotional and cognitive categories. In light of this outcome, we discuss alternative approaches for resolving such ambiguous cases.
Discussion
Implications
Developing a strategy to integrate SORKC behavioral analysis into the psychometric assessment of the SCL-90-R proved challenging. We faced constraints and serendipities and identified topics for future research. The prevalence of psychosomatic symptoms (Leclerc et al. 2022) creates social and economic pressure on insurance and welfare systems (Knapp and Wong 2020), leading to high expenditures worldwide (European Agency for Safety and Health at Work 2014). Our approach offers a faster diagnostic method, reducing costs of psychosomatic diagnostics in time and money. This feasibility is currently under investigation. Shorter diagnostic time benefits not only institutions but also patients, who receive treatment earlier.
Methodological Reflection
The findings emphasize semantic criteria. The semantic criterion in our evaluation plan predicted deviations from the 95% CI more accurately than any other parameter. Future item and questionnaire design should therefore place stronger emphasis on semantic coherence, as both clinical populations and patients may struggle with item interpretation and rating.
Limitations
Our sample showed homogeneous ecological validity. To better assess feasibility, future research should include experts from diverse cultural backgrounds and ethnicities, balanced by gender, age, and clinical psychology experience. Further studies may also examine additional sociodemographic factors such as socioeconomic status and personality traits, which could then be considered and statistically controlled.
Bootstrapped methods served as a safety precaution. The consistent outcomes across consensus and non-consensus groups suggest that analyzing one variable may suffice. Identifying the variable best suited for this role should be a focus of future research. As noted, if fewer than 66% of experts agreed on an item and it lacked inseparable connectedness, all parameters deviated from the 95% CI. This was not predictable due to non-parametric bootstrapping and random distributions. Nevertheless, our sample size was adequate and indicates that fewer experts may generate sufficiently sensitive data. Once refined, fewer than 40 experts may be sufficient, provided the sample is diverse enough to ensure validity, thereby expediting and reducing costs.
Unexpectedly, items excluded from analysis (Figure 2) showed lower mean variance. For example, Items 63 and 67 (“Awakening in the early morning” and “Having ideas or beliefs that others do not share”) had high agreement rates but failed to reach the 66% criterion, leading to lower variance.
Regarding item 70 (“Feeling everything is an effort”), an alternative would be a dimensional approach, allocating items to every category with >0 responses based on percentages. We refrained from this to maintain the agreement rate criterion, opting for a dichotomous evaluation. Future research could compare these two approaches to determine advantages and disadvantages. Empirical comparisons would allow scientists to select the most suitable approach rather than relying on educated guesses, as in our exploratory case.
Limitations of the utilized classification system SORKC, the statistical implications of standardization (Radha and Raj 2020), and the debility to depict interactive effects (Pietrzak et al. 2018) are well known. Thus far, the most frequently cited drawback is the little representation of social variables and the limited influence of emotions and psychosocial variables in contrast to other models (e.g., Tretter and Loeffler-Stastka 2021; Wilson 2002).
Our proposed approach for integrating these measures serves as a basis for future research. We emphasize that this method does not render behavioral analysis obsolete in every situation. Future research should focus on investigating the potential positive and accelerating effects of supplemental initial information on therapy duration and outcomes. Moreover, our approach introduces a new way to operationalize the comparison between patients’ self-assessments of BRCs and therapists’ observed BRCs on a behavioral analysis level. This enables future research to analyze perceptions in the therapeutic process in new ways. Currently, it is known that the implementation of superior psychometric assessments has a favorable impact on therapy (BMSGPK 2020), however, there are no data available regarding whether there is an upper limit of diagnostic processes at which patients no longer profit from it in any form. Integrating multiple measures to attain maximum utility could potentially reduce therapy time and make it more financially accessible to individuals with lower socioeconomic backgrounds or in countries with limited social health care systems.
Conclusion
Our explored method presents a viable approach for integrating existing psychometric assessments by extending the analysis without expanding the test itself. This process was effective and without any statistical disadvantages. Our findings indicate that fewer experts are required to investigate this issue, as demonstrated in our study. However, in addition to mathematical considerations, achieving acceptable content validity requires larger sample sizes in studies that evaluate the heterogeneity of experts. Nonetheless, our approach offers the potential for patients, therapists, and other clinical personnel to obtain valuable information from situational analyses in the form of cross-situational perception analysis at the outset of a clinical journey. Preliminary analysis showed that using established questionnaires and interpreting item answers in terms of behavioral response categories holds potential for reducing costs in terms of time and money needed for an adequate diagnostic process.
Footnotes
Author Note
They had full access to all data and take full responsibility for its integrity and accuracy.
Ethical Considerations
This study was approved by the Ethics Committee of the University of Ulm on 07.09.2023 (No. 298/23). All participants provided written informed consent for participation and publication.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
