Abstract
We used the model of prospective memory and habit development to derive recommendations for designing behavior-change campaigns that used prompts or household visits as reminders. We followed an exemplary procedure comprising the calibration of the model, based on 48 time series gathered during a campaign promoting recycling habits and a systematic exploration of the solution space. For the parameter estimation, an algorithm was developed that worked at two levels. A higher level algorithm optimized parameters that were set to equal values for all agents, whereas a lower level algorithm estimated the values of agent-specific parameters for each agent separately, using the parameter values of the higher level algorithm for the other parameters. This procedure resulted in an excellent fit of the model to the data (R 2 = 75%) For the systematic exploration, an indicator expressing campaign effects in one value was defined and the following findings could be derived. Activities should focus on the first week of a campaign. Follow-up visits or refreshing of prompts should be done within 4 days after the initial visit. Later activities, such as additional visits or refreshing of prompts, bring little further effects. Investing heavily in the design of the prompts for improving their salience is only worthwhile in populations with a low commitment to perform the behavior. Furthermore, covering more than 10% of the places where the target behavior should be performed with prompts mostly does not lead to additional effects to make it worthwhile.
Keywords
Introduction
Routine behaviors are simple actions that are frequently performed without thinking much about them. Nevertheless, due to the high number of performance events, the effects of these actions can add up to large impacts. For example, regularly taking prescribed medication can save one’s life, and regularly separating waste for recycling or composting can conserve a large amount of resources. Thus, routine behaviors are critical for protecting health and the environment, but each single action is of such low importance for an individual that it is performed mostly unconsciously or habitually. How then can such behaviors be changed? What measures can be taken if people are already convinced and intend to perform a behavior but at the critical moment, simply forget to do so? Neither arguments nor norms are effective in such situations (because the people are already convinced), but a number of techniques (so-called memory aids) exist to remind people, right at the moment when the behavior should be performed, to do so (Intons-Peterson & Fournier, 1986). These authors also distinguish between external memory aids or written prompts—which are setup in situations where the behavior should be performed—and social memory aids or verbal prompts—which are other persons’ reminders for the person to perform the behavior (e.g., in the form of household visits). In this article, we refer to external memory aids simply as “prompts” and social memory aids as “household visits.”
Right after being reminded, one starts forgetting again, but if the actions are performed frequently enough, habits develop that prevent forgetting the execution of the behavior (Tobias, 2009). Thus, changing routine behaviors by reminding a person to perform certain actions is a dynamic process that unfolds over weeks and months. Consequently, it is not at all obvious how memory aids should be best deployed within a campaign that aims to change a routine behavior. Which and how many techniques should be used together to maximize efficiency? How often should they be used and at what intervals? Is it better to invest in designing highly effective measures or invest in a larger coverage if resource limitations do not allow maximizing both?
An established method for investigating complex dynamic problems and optimizing interacting processes is computer simulation. This approach allows conducting a large number of experiments in silico that would be impossible to do empirically. However, most models are designed without considering how their parameters could ever be estimated, based on time-series data representing real events. For optimizing campaigns that change routine behaviors, Tobias (2009) presented a model that can be calibrated to represent real-world scenarios. The model of prospective memory and habit development performed well in an empirically based system analysis but was never applied to optimizing campaigns. We designed a hierarchical optimization algorithm to calibrate this model and explored the effects of systematic parameter variations to determine the optimal design of campaigns that could change routine behaviors with interventions that would prevent forgetting.
This article has a twofold contribution. On one hand, we derive praxis-relevant recommendations on how to design campaigns that promote intended but not yet habituated routine behaviors. On the other hand, we demonstrate how computer simulations can be used in applied social–scientific research comprising the calibration of the model to specific settings and its systematic exploration by simulation experiments. By doing so, we also uncover some problems that might be inherent to this approach but have not yet received much attention. We first present the model and then explain the parameter estimation algorithm. Finally, we discuss the design of the simulation experiments and present the key results.
Model of Prospective Memory and Habit Development
Overview
The model of prospective memory and habit development investigated here was presented by Tobias (2009) with an extensive theoretical and empirical foundation. The model represents processes of forgetting and reminding and habit development due to frequent performance of a behavior. It simulates how the frequency with which a behavior is performed changes over time due to various forms of being reminded and the dependence on psychological variables and parameters. The model was designed to explain the dynamics of individual behaviors in cases, where remembering their performance in critical situations is the main determinant for behavior execution. Therefore, it is suited for explaining the changes in simple and frequently performed everyday behaviors that people already intend to perform but to which they are not yet habituated.
Preferences (intentions, attitudes, norms, etc.) are not considered here because to change them, completely different campaigns have to be implemented. In cases where people are not yet convinced to perform a behavior, first, a campaign targeting preferences is usually implemented and only then and separately, a campaign using memory aids to support the development of habits is added (e.g., Kraemer & Mosler, 2012). Combining all processes into one model would add almost no explicative power (as shown later, the prospective memory model can explain 75% of the variance in the empirical data) but dramatically increase the complexity of the model, most probably to a degree that would make calibration of the model impossible. However, the model considers external (e.g., prompts) or social (e.g., home visits) memory aids and techniques that increase the persons’ commitment to perform the behavior (e.g., self-commitment), such as the campaigns of Inauen, Tobias, and Mosler (2014). Since preferences are not considered in the model, there is also no need for modeling the interaction among individuals, 1 which reduces the complexity of the model even further and considerably simplifies the parameter estimation.
The model is a parsimonious integration of empirical findings regarding memory processes and habit development. Its particular strength is the form of considering situation-specific differences in behavior performance. As explained later in more detail, the effects of reminders and the development of habits depend much on the cues present in specific situations or the distance from these cues when a behavior is to be performed. Considering all characteristics of all possible situations in which a behavior could be performed would lead to an extremely complex model and would require an enormous amount of data to initialize the model. Two concepts are used to avoid such problems.
First, the main output variable of the model is behavior frequency (BF), defined as the fraction of opportunities that were taken to perform a behavior (e.g., the number of times that waste was separated for recycling among all the times that waste was disposed). To reach a certain BF, the behavior has to be performed in different situations that might have different psychological characteristics. Consequently, different BFs have different attributes, such as habit strengths or accessibilities. For example, a person can have a strong habit to separate waste in his or her kitchen but not on the street. The resolution of BF (i.e., how many situations are distinguished) can be set to any value; for this investigation, the resolution is 0.05, resulting in up to 21 different situations in which the behavior can be performed.
The other concept used to abstract from specific situations, in which a behavior can be performed, constitutes similarity functions. The idea is to express all possible situations in the form of their similarity to a reference situation, in which characteristics of the situation (e.g., cues present) and its consequences for the behavior (e.g., whether bins for collecting waste separately are present) are considered. These similarity functions can be defined to represent each situation separately; however, the concept’s strength is that these functions can also be defined in a rather simple form to abstract from single situations and by doing so dramatically simplify the model.
A critical simplification applied here is the assumption that a behavior is performed in all situations, in which it is performed to achieve a lower BF, also when it is performed in any higher BF. This assumption is justified in cases where behaviors and situations can be ranked on one dimension regarding the difficulty to perform or, more specifically, remember the behaviors. In cases where two or more dimensions are required for ranking the behaviors, more complex similarity functions have to be defined accordingly. However, the similarity functions should be kept as simple as possible, so that the model can be calibrated with little empirical data and systematically explored to optimize real-world campaigns.
The model is implemented as a discrete-time simulation, and the only entity considered is an agent representing the psychological processes within a human being. In each time step and agent, six processes are updated, which are presented in Figure 1 and explained in more detail in the following subsections. In the first step, in case a prompt is setup, its salience (i.e., how much the prompt catches one’s eye) has to be updated. If in the current round, a prompt is setup, the salience is set to 1 (i.e., its maximum value). The accessibilities and habit strengths have to be updated for each BF separately. Increases and decays in accessibility and habits are synchronously updated (i.e., all changes are summed up and then added to the previous value, considering the bounds of the variables). Next, the BF executed in the current simulation step is determined. Table 1 compiles all parameters, variables, and functions used in the model as reference. The parameter values presented are the result of the calibration to the empirical time series.

Process overview for the model of prospective memory and habit development.
Parameters, Variables, and Functions of the Model and the Estimated Parameter Values.
Note. The range shows the bounds for the values used for estimating the parameters. With the exception of the slope and the turning-point parameters of the similarity function, which have theoretical ranges of [0.0, ∞), all parameters have theoretical value ranges of [0.0, 1.0]. The value shows the estimated parameter values used for the experiments. The values for the agent-specific parameters are not presented because they have been varied for the experiments.
Submodel 1: Selection of Executed Behavior Frequency
The processes related to forgetting and remembering intended behaviors are investigated in prospective memory research (see McDaniel & Einstein, 2007 for an overview). Three factors determine whether an intended behavior is remembered at a critical moment: (1) the accessibility of the behavior, (2) cognitive resources in the situation, and (3) the habits developed for that behavior. Accessibility describes how easy it is to remember a behavior (e.g., Higgins, 1996). Depending on the situation, accessible behaviors may be forgotten; at other times, even behaviors that are difficult to remember are remembered. Cognitive resources are the most investigated situational influences on remembering an intended behavior (e.g., Loft & Yeo, 2007; Marsh, Hancock, & Hicks, 2002). Cognitive resources describe how much a person is able to think or actively remember, which is mostly determined by the number of distractions. On the other hand, situational effects can also facilitate remembering by activating habits, which are understood as the associative connections between a situation and a behavior that is often performed in a given situation (Aarts, Verplanken, & van Knippenberg, 1998; Danner, Aarts, & de Vries, 2007). These findings can be formalized by comparing the accessibility to situation-specific influences on cognitive resources and habit activation. In the model, the behavior is remembered in all situations related to the BF, for which the accessibility of AccBF is:
C AT is the constant of the accessibility threshold, which determines the relative weighting of the influences of distractions and habits; the greater the constant, the stronger the influence of habits and the weaker the influence of distractions. While habit strength (HSBF, weighted with WHAT) reduces the threshold (i.e., facilitates remembering), distractions increase the threshold (i.e., make remembering more difficult). Because distractions are situation specific, they are a function of BF; if a behavior is seldom observed, it is mostly done in situations where it can be remembered easily. With increasing frequency of the behavior performance, the behavior also has to be performed in situations with more distractions. Ultimately, a linear function is sufficient for replicating the present data; therefore, BF is simply weighted by WBFAT.
The distraction effect is limited to
Submodel 2: Forgetting and Remembering the Execution of Intended Behaviors
Accessibilities are increased by (1) certain events (e.g., being reminded by a person; Schaefer & Laing, 2000), (2) performing the behavior (Hacker, Herrman, Pakossnik, & Rudolf, 1998), and (3) certain cues such as reminders (Intons-Peterson & Fournier, 1986). Without these influences, the accessibilities decay over time (Higgins, 1996). Regarding the accessibility-increasing influences, three moderating factors have to be considered: (1) the importance of the behavior for the person or the commitment of the person to perform the behavior (Gollwitzer, 1999), (2) the spatial distance of the behavior performance from a reminding cue (Guynn, McDaniel, & Einstein, 1998), and (3) the salience of the reminder (Marsh, Hicks, & Hancock, 2000), which decays over time. The accessibility dynamics can then be defined as follows:
In each time step, the accessibilities (AccBF) decay (AccDecay) and may be increased by certain events (AccGainEvent), the performance of the behavior (AccGainBeh) or a reminder (AccGainRem). The accessibility decays proportionally to the accessibility decay parameter (ADP):
The effect of events on accessibility is partially independent (AGCEvent) of the commitment intensity (CI) and partially dependent on it (weighted by WCIEvent). The effect of reminding events on the accessibility of the behavior is formalized as:
The range of the effect that depends on the CI is limited to the range that is not explained by the constant. This facilitates the interpretation of the parameter values and improves the identifiability. For the data-based investigation, two types of events are considered: (1) Event 1—setting up the reminder (investigated, e.g., by Goschke & Kuhl, 1993) and (2) Event 2—being reminded by other individuals (investigated, e.g., by Schaefer & Laing, 2000), particularly the approximately weekly interviewer visits.
The increase in accessibility due to behavior performance is modeled to be directly proportional to the frequency of performing the behavior:
The effect of a reminder is formalized similarly to the effect of events with a constant (AGCRem) and by weighting CI (WCIRem). Additionally, the distance of the behavior performance from the reminder is considered by multiplying the similarity function Sf R (BF):
Again, as in Equations 1 and 4, the range of the effect that depends on the CI is limited to the range that is not explained by the constant to prevent an overlap of the constant and commitment-dependent effect. The similarity function Sf R (BF) models the effects of differences in the situations in which the behavior is performed. To formalize how many situations are very similar to the one in which remembering the behavior is easiest and how many and how different the other situations are, a logistic function is used, which is defined by three parameters specifying range (DP), slope (SS ), and turning point (TS ):
The similarity function used for modeling habit development (see Equation 12) uses the same function, with the same parameter values (except experiments on coverage, where the turning-point parameter for prompts, TP , is varied independently). Figure 2 plots the function with the calibrated values for T and the largest value used in experiments.

Plot of the similarity function with turning point parameter T as calibrated and T = 1.0 (largest value used in experiments).
SalienceRem models the decaying effect of a reminder; at the moment of setting up the reminder, the parameter is set to 1 and then decays proportionally at the rate set by the salience decay parameter SDPRem:
Submodel 3: Development of Habits
When a behavior is performed frequently, habits develop. These associations between situational cues and the behavior foster the ability to remember performing the behavior. Models of phenomena related to habits (e.g., Botvinick & Plaut, 2004; Krushke, 2001) can be summarized as three fundamental principles: (1) habit strength decays over time, (2) habit strength increases by performing the behavior in a specific situation, and (3) the similarity of the current situation and the situations in which the behavior had been performed earlier determines the increase in habit strength. The dynamics of habit strength (HSBF) are defined as follows:
Habit decay is again modeled as proportional decay with the habit decay parameter (HDP):
The increase in habit strength depends on the frequency of the executed behavior (BFExe). To scale the habit strength for any BF between 0 and 1, the increase in habit strength further depends on the HDP and on the current habit strength (HSBF). Thus, for any BF, a habit with a value of 1can be reached, but for lower BFs, this process takes much longer:
Equation 11 describes the increase in habit strength at all frequencies up to the executed BF (i.e., for all situations in which the behavior is actually performed). A number of studies, from behavior shaping (e.g., Skinner, 1953) to data-based modeling (e.g., Botvinick & Plaut, 2004), show that performing a behavior in a specific situation also affects associations between similar behaviors and similar situations. The higher the degree of similarity is, the stronger the effects will be. For the model presented here, it is assumed that habit strength increases for BFs higher than BFExe if the additional behavior performances occur in similar situations. Thus, for BFs higher than the executed frequency, habit strength increases by:
The smaller the difference is between the value of the similarity function for the executed behavior Sf H (BFExe) and that of the higher BF, Sf H (BF), the more is habit strength increased for the higher BF.
Implementation of the Simulation
In this subsection, we provide technical information about how the simulation was implemented. This information completes the model description above, and it allows implementing the simulation in any platform and replicating our investigation.
The model was implemented in Java as a completely deterministic time-discrete simulation. One step of the simulation represented one day in the empirical data. For the parameter estimation, the simulations were conducted for 40 steps, setting up the only prompt at Step 10. The experiments were performed for 103 steps, with the first visit at Step 3. Since the agents would not interact, the number of agents would have no effects on the outcomes. The parameter estimation was done with 48 agents, one for each data set available. The simulation experiments were performed with only one agent, for which the parameters were systematically varied.
The simulation is initialized as follows: For all parameters and variables, the values are set based on an input file, with the exceptions of the start values of AccBF and AGCUnkn. The start values of AccBF are calculated based on the start value of BFExe and a parameter that specifies how close the accessibility is to neighboring BFs (both values are given in the input file). Even if the start BF is defined, the start value of the accessibility can still be set in a range that is given by the value, for which this BF is just remembered. Thus, the slightest reduction in the accessibility will reduce BFExe up to the value, for which the next higher BF is just not remembered, and thus, the slightest increase in the accessibility will increase BFExe. The additional initialization parameter defines the start value of AccBF together with the start value of BFExe, and this approach makes it impossible to set the start value of AccBF to implausible values.
AGCUnkn is calculated before the start of the simulation and for each agent separately, based on the assumption that the start behavior is stable. This means that if no reminding event takes place and no prompt is setup, the agents will show the BF set as the start behavior for the rest of the simulation. To reach this, the increase in the accessibility without any reminding intervention should have exactly the same value as that of the decay in the accessibility. Without reminding interventions, the accessibility is increased by the behavior performance and unknown influences, which can be understood as an agent-specific base rate for remembering. Thus, AGCUnkn can be calculated as:
AccDecay and AccGainBeh were already explained (Equations 3 and 5, respectively). In the presented investigations, the agents are always initialized with HSBF = 1. In the case of setting the start BF to 0, there can be no habits for the behavior. However, for calculating AGCUnkn also in this case, a habit strength of 1 for the lowest behavior intensity larger than 0 is assumed. Otherwise, implausibly large jumps to high BFs would be observed if an agent would experience an increase in the accessibility of the behavior.
Recalibration of the Model Using a Hierarchical Parameter-Optimization Approach
Methods for Estimating the Parameters
This study builds on the data presented by Tobias (2009). His study built on 48 time series gathered during a pilot campaign promoting waste recycling in Santiago de Cuba, Cuba (Mosler, Tamas, Tobias, Caballero Rodríguez, & Guzmán Miranda, 2008; Tobias, Bruegger, & Mosler, 2009). Every household in the sample received a sheet of paper with the following text (in Spanish): “Please classify and separate! Glass—Aluminium—Paper—Cardboard—Plastic.” This paper was intended to remind the residents to separate waste for recycling. For a month, the participants filled out a short, daily questionnaire that asked about their behavior and various psychological constructs (Tobias & Inauen, 2010). The respondents reported their behavior on a scale with six categories, indicating that about 0%, 10%, 25%, 50%, 75%, and 100% of the waste were being separated for recycling. The interviewers collected the questionnaires approximately once a week. These visits turned out to have an additional reminding effect.
To calibrate the simulation model based on these data, one has to solve the problem that the number of parameters is too large for efficiently running general optimization algorithms. Although the agents of the model investigated here had only 20 parameters available to them, this number must be multiplied by the number of agents (or different data sets) considered. The estimation problem could be simplified by setting the parameters to the same values for all the agents (such parameters were denoted as global). For many parameters, this would actually make sense, considering that models would abstract from minor variations to focus on the principal differences. However, the power of microsimulation models would be lost if all attributes would be set to equal values for all agents. Therefore, in the present model, three parameters were estimated differently for each agent (agent-specific parameters): the BF at the beginning of the simulation (BFStart), the CI, and the accessibility of the start behavior (AccStart).
Estimating parameters that would affect all agents and at the same time, parameters that would only affect some agents, would lead to new problems due to differences in parameter sensitivity. Obviously, changes in the parameter values that only affected a single agent would have a minimal impact on the model fit compared to changes in the parameters that affected all agents. Consequently, many agent-specific parameters would be estimated at suboptimal values. We therefore proposed a hierarchical optimization approach (Figure 3). At a lower level, only the agent-specific parameters were estimated for each agent separately, given a set of values for the global parameters. These parameter sets were determined by the higher level optimization algorithm. The value of the error that the higher level algorithm minimized was calculated as the sum of the errors of the best solutions found for each of the agents. In the studied case, the agents did not interact, which simplified the problem. In the case of interacting agents, further iterations at the lower level would be required to reduce inconsistencies between agent inputs and outputs.

Flow diagram of the hierarchical optimization approach used in this study.
In this specific case, the higher level algorithm estimated 17 global parameters. Due to the enormous search space and the nonlinearity of the model, the use of the computation-intensive, differential evolution approach (Storn & Price, 1997) seemed justified. In this approach, populations of candidate solutions are randomly generated by combining solutions of previous rounds according to the errors each solution produces. Solutions with lower errors have higher chances of producing “offspring” solutions for the next generation (i.e., optimization round). Due to the random element and the fact that this approach does not guaranty finding the absolute optimum, for different random seeds, different (local) optima are found.
At the level of the agents, one parameter had discrete values (BF); the values were systematically varied. The remaining two parameters were estimated by using a downhill simplex method (Nelder & Mead, 1965). The basic idea of this approach is to search an n-dimensional space by randomly selecting n + 1 starting points and, in each optimization round, to replace the solution with the largest error by a new set of parameter values. Based on a number of rules, such as reflecting this point of the simplex to the other side and expand or shrink the simplex, solutions with lower errors can be found. Again, this approach does not guarantee to find the absolute optimum and the solutions depend on the random seed.
The Apache Commons Math package version 2.0 was used for the implementation of both algorithms. As the target function, the mean absolute error (MAE) was minimized. 2 The optimizations ran for about 300 iterations of the higher level algorithm. In most cases, the optimization converged after about 100–150 iterations. The total computing period for one parameter estimate was about 2 days. To investigate the identifiability of the parameters (i.e., whether similar solutions could be achieved with different parameter values), we started the optimization 300 times, each with a different random seed.
Results of the Parameter Estimation
The algorithm led to parameter estimates, for which the model output fit the data well (MAE = 0.110, root mean square error = 0.229, R 2 = 75%). Of 1,170 data points, only 44 (3.76%) had an absolute error greater than the measurement resolution (0.25). Only 50 (of 300) solutions had an error larger than 110% of the best-fitting solution, and 101 solutions had an error less than 105% of the best solution. Within the 5% best-fitting solutions (largest error = 102% of the best solution), three patterns of parameter values were identified by visual examination. The main difference in these patterns was the prompt effect. While in one solution, prompts showed no effect, we selected the solution where the prompt effects were strongest. This selection was based on theoretical and practical considerations. Many empirical studies showed that prompts could be effective interventions for changing behaviors (e.g., Fry & Neff, 2009). Therefore, we could exclude the solution where prompts showed no effect. Furthermore, for optimizing the usage of prompts, the solution with stronger prompt effects would lead to clearer results. Nevertheless, the experiments were run for all three solutions, without affecting the conclusion, except that prompts might have been less effective.
Application to Campaign Planning: Promoting Waste Separation Habits Based on Prompts and Home Visits
Experimental Design
With the model calibrated to the empirical data, the following research questions were investigated: (1) What is the optimal period between household visits and setting up prompts? (2) What is the optimal combination of prompts and visits? (3) What effects do improvements of the design and coverage of prompts have on the campaign effects? In investigating these questions, we varied the period between visits, the number of visits, the number of prompts, and the characteristics of the prompts. Such systematic explorations of the model were straightforward. Nevertheless, a number of issues had to be handled.
First, a systematic exploration of a model by simulation experiments would produce an enormous amount of data, and it would be difficult to keep track of the relevant information. To simplify the analysis and presentation, for each simulation outcome, one value was calculated that characterized the success of the campaign. For a better understanding, Figure 4, Diagram A, presents two time-series examples.

Examples of simulated time series (x-axis represents simulation steps or days). Diagram A illustrates the parameters a and b for calculating campaign effect. Diagram B illustrates the effects of varying the period between visits for one follow-up visit (V) and Diagram C of varying the number of follow-up visits every 4 days, both for the case of one prompt (P) installed during the first visit. Diagram D illustrates the effects of varying the number of prompts for one follow-up visit after 4 days. Diagram E illustrates the effects of varying the salience (S) of the prompts and Diagram F of varying the coverage (C) of the prompts.
Typically, visits lead to a sharp increase in the BF, which can then quickly decay. If the BF decays to low values, the action is performed too seldom for habits to develop, and it stabilizes at low values (Person 1 in Figure 4, Diagram A). If the BF decays to levels that are not too low, habits can still develop, and after a while, the BF increases again to stabilize at the maximal BF (1.0; Person 2). Thus, we have two characteristics to express the campaign effect: (1) the behavior change reached at the end of the campaign (a) and (2) the time required to stabilize the maximum behavior (b). Obviously, we prefer campaigns where the value of a is as large as possible. However, if we have several campaign designs reaching the maximum behavior, we prefer the design where this behavior is reached earlier (i.e., where the value of b is smaller). Following this logic, these two characteristics are combined into one indicator, labelled the “campaign effect”:
The campaign duration was set to 100 days. If the BF is below 1.0 after 100 days, the behavior change achieved (a) is returned. If BF = 1.0 after 100 days, the period until this behavior is stable (b) is considered. The smaller the b value is, the more value is added to the behavior change (up to one, in case of an immediate change to the maximum BF). In Figure 4, Diagram A, for the less successful campaign (Person 1), the campaign effect = 0.15 – 0.10 = 0.05. For the more successful campaign (Person 2), the campaign effect = 1.0 – 0.10 + (1–21/100) = 1.69. The interpretation of the campaign effect is straightforward; the greater the value is, the more the behavior changes and the faster the highest possible behavior stabilizes.
Additionally, the composition of the target populations might vary regarding accessibilities, CIs, and behaviors performed before the campaign. To achieve more generalized results, we varied these characteristics over their entire ranges of values and used the average of all outcomes as the result of one parameter combination (“mean campaign effect”). While not presented here, the entire table of campaign effects was investigated before calculating the averages for each of these results. Mostly, the tendencies presented held true for any segment of the population, with one exception—the observed ceiling effects were less pronounced in populations with more difficult starting conditions (i.e., low CI, low BFStart, and low AccStart). This meant that in these cases, increasing campaign efforts (e.g., additional visits) would have more additional effects, making them more efficient. However, in some cases, it was necessary to separate the results according to the CI to obtain the full picture of the intervention effects.
Our approach allows extracting and clearly presenting the main effects found in the experiments we conducted for this study. However, the aggregated indicators are relatively abstract values. To give a better sense of how the experimental manipulations affect the processes, we present some individual time series. Nonetheless, it is important to note that such single results can differ strongly from the core tendencies found when investigating entire populations and experimental series. To allow for comparability, all time series are calculated for AccStart = 0.4, BFStart = 0.15, and CI = 0.4. This represents persons who will be interesting targets of a campaign because their behavior is low but they have a medium receptiveness to interventions. Compared to the results averaged over the entire space of values investigated for the three agent-specific parameters, these time series represent cases with somewhat stronger campaign effects for most settings.
For the experiments, the numbers of prompts (0, 1, and 2) and visits, in addition to the initial visit (0, 1, 2, 3, and 4), were varied. The periods between visits varied, with ranges of 2–20 days (1 additional visit), 2–10 days (2 additional visits), and 2–6 days (3 and 4 additional visits). Therefore, irrespective of the period and number of visits, the intervention activities lasted for a period of up to 20 days.
The salience and coverage of the prompt were varied in the following ways. Four levels of salience were investigated: a rather unimpressive prompt (low salience, as used in the campaign in Santiago de Cuba, AGCRem = 0.00), a medium salience (AGCRem = 0.05), a high salience (AGCRem = 0.10), and a prompt that had a similar daily impact as a household visit (extreme salience, AGCRem = 0.15). The highest level of salience might be similar to the effect of using mobile electronic devices to remind persons of the behavior. Four levels were also considered for the coverage of the prompt. At the lowest level, the prompts were setup in about 10% of the situations in which the behavior could be performed (TP = 0.25). At the next three levels, approximately 25% (TP = 0.50), 40% (TP = 0.75), and 50% (TP = 1.00) of the occasions for performing the behavior were covered, respectively. Obviously, the higher the salience and the coverage were, the stronger the effect of a prompt would be. However, because the relation was not linear, we wanted to find out at which levels further increasing the salience and/or coverage would lose efficiency. To avoid confounding the effects due to prompts with the effects due to follow-up visits, no follow-up visits were considered for these experiments.
Results of the Experiments
Optimal frequency of and periods between home visits
In the first step, we investigated the periods between visits that had the strongest effect on reminding persons to perform an action. Table 2 compiles the mean campaign effect for different periods between visits, different numbers of visits, and different numbers of prompts. The optimal period between visits turned out to be every 4 days. When only one follow-up visit was possible, and no prompt was setup, the highest mean campaign effects were reached by a visit 6 days after the initial one. The time series presented in Figure 4, Diagrams B and C, show rather strong reactions to the interventions (compared to the mean effect). Only in the case without the follow-up visit would the campaign have failed, but the right setup could considerably accelerate the habit development.
Mean Campaign Effects for Various Periods Between Visits, Various Numbers of Visits, and Various Numbers of Prompts.
Note. The simulated campaigns were setup to last about 20 days. Thus, for more frequent visits, only simulations with shorter periods between visits were run (e.g., up to 10 days for two visits, up to 6 days for three and four visits). The periods that reached the highest mean campaign effect are in italics.
Optimal combination of prompts and home visits
Knowing the optimal period between visits, we could then investigate different combinations of reminders (i.e., visits and prompts). These comparisons were done for 4-day periods between visits. Furthermore, since the results differed between cases with low and high commitment, the results that considered only parameter combinations with high commitment are presented separately. This information is compiled in Table 3.
Mean Campaign Effects for Varying Numbers of Prompts, Follow-Up Visits in the Case of Visits Every 4 Days.
Generally, the principle “the more, the better” held true for these techniques. However, due to ceiling effects, less was won by adding more elements to the campaign. For example, prompts were more effective with fewer visits, and additional visits were more effective with fewer prompts. Also, a second prompt had only about 23% of the additional effect as the first prompt. From an applied perspective, it would seem always worthwhile to setup a prompt, since distributing memory aids would be far cheaper than additional visits. However, at least in theory, the number of visits would be unlimited, while refreshing prompts would not be very effective. However, considering the additional effects, more than one additional visit would not seem very efficient, and more than two would not be recommended. With one additional visit, even refreshing the prompt would still be effective, and the mean campaign effect would be 86% of the effect of four visits and two prompts (1.09 compared to 1.27). In the case of a highly committed population, the effect would even be 96% of this effect (1.33 compared to 1.39). With two visits and two prompts, 93% and 99%, respectively, would be reached. The time series presented in Figure 4, Diagram D, show relatively strong reactions to the number of prompts, with no prompt leading to the campaign’s failure and two prompts to almost immediate success.
Effects of prompt coverage and salience
Finally, we investigated the effects of changing the prompt characteristics. We were particularly interested in setting up prompts in a way that more occasions for performing the behavior fell under the influence of the prompts. We were also interested in the effect of increasing the salience of the prompts. Table 4 shows the results of such variations for the cases that did not involve additional visits.
Mean Campaign Effects for Prompt Salience and Coverage (Percentage of Places With Relevant Prompt Effects).
Note. The salience is graded low (L), medium (M), high (H), and extreme (E; i.e., equal to household visits).
Again, we observed the more, the better tendency and a strong interaction with the commitment level of the population. Increasing the salience of the prompt was particularly effective in the case of low-committed populations and in places with more occasions to perform the behavior. For example, by changing a prompt from low to medium salience, the mean campaign effect increased by 0.52 for low-committed persons (from 0.70 to 1.22 for 50% coverage, Table 4), while for high-committed persons, the mean campaign effect increased only by 0.07 (from 1.30 to 1.37, 50% coverage).
Trying to cover more locations to affect behavior was an inefficient strategy. By increasing prompt coverage fivefold (i.e., covering 50% of the places where the actions should be performed, instead of 10%), the campaign impact increased by only between 10% (high-committed population with high salience: change from 1.27 to 1.40) and 28% (low-committed population with low salience: change from 0.54 to 0.70). A higher increase in the campaign effect could already be achieved by increasing the salience from low to medium. However, the exception was in the case of high-committed populations, in which increasing the salience had the same (minimal) effect as increasing the coverage.
Figure 4, Diagrams E and F, presents some examples of time series to illustrate these effects. For this specific group of persons, prompts with low coverage and low salience would be ineffective. However, with either a small increase in salience or in coverage, campaign success would have been reached quickly.
Discussion
We developed an algorithm for estimating the parameters of the model of prospective memory and habit development (Tobias, 2009) based on 48 empirical time series gathered during a campaign promoting recycling habits. The excellent fit of the model to the dynamic data was an indicator of the model’s validity. However, three types of parameter profiles were identified that all led to a similarly good fit but different interpretations regarding the effectiveness of using prompts. We selected one solution based on theoretical and practical considerations. We also systematically varied relevant parameters to investigate how campaigns targeting routine behaviors could be designed in the most efficient way. Using an indicator of campaign success, we could reduce the enormous amount of data produced by the simulation experiments to derive concrete recommendations for how to design such campaigns. We first discuss the practical and then the methodological implications of our results; finally, we present more general conclusions.
Recommendations for Implementing Campaigns for Changing Routine Behaviors
The optimal period of home visits turned out to be every 4 days. This might be a bit surprising, as one might think that in a 3-week campaign, it would be better to do a follow-up visit much later to have a longer lasting effect. However, the most critical period comprises the first days after the initial visit and maybe setting up a prompt. If after a week, the BF is too low, the behavior stabilizes at a low level (BFStart or slightly above), and habits do not develop for a higher BF. In contrast, if the BF does not fall below a critical value, habits start to develop for a higher BF and maybe slowly but certainly increase BF.
Combining many visits with many prompts only makes sense in populations with strong memory issues and low commitment. Often, combining prompts with two visits—an initial one for setting up a prompt and a follow-up visit for refreshing the prompt—is the most efficient strategy. In cases where prompts cannot be used or are ineffective (e.g., because of very low commitment), a third follow-up visit (but not more) makes sense. The reason that more visits are inefficient is the same as explained above; the critical phase is shortly after the start of the campaign. Later visits do not have much effect.
In the case of low-committed populations, the salience of the prompts can be increased. However, highly salient prompts might be expensive or difficult to implement. In highly committed populations, the salience of the prompts is of low importance; therefore, in such cases, simple prompts are the most efficient solution. Trying to set up prompts on many locations where the behavior should be performed turns out to be inefficient. It appears that waste is produced in many different places and situations (see Mosler, Drescher, Zurbrügg, Caballero Rodríguez, & Guzmán Miranda, 2006). For other populations whose behavior has to be performed on few, easily determinable locations, setting up prompts on these sites might be an option.
Discussion of Methodological Implications
Although already introduced by Tobias (2009), the model used is in many ways different from typical social simulation models. The model is completely based on empirical evidence and hypotheses that were tested by fitting the model to empirical time-series data. Furthermore, the parameters all have real-world meaning, and the effects of interventions were explicitly modeled, so that the model could be applied to replicate, forecast, and optimize real-world campaigns. Finally, the model is kept so simple that it allows estimating its parameters based on empirical data and systematically exploring the solution space. Although the model considers interacting agents and processes of preference change, Tobias’ (2009) decision to switch off these elements is particularly remarkable because this dramatically simplifies the model but only minimally reduces its explicative power (i.e., fit to the empirical data) and applicability (i.e., its capability to answer specific research questions). Of course, these exact simplifications are not possible when investigating other phenomena. The point we want to make is that the goal should be to avoid as much complexity as a specific investigation allows, instead of developing models that are more realistic but too complex to be calibrated or systematically explored. This refers not only to the model structure but also to other elements of a simulation, such as the initialization (e.g., in our case, that many parameters are set to the same value for all individuals). A model that effectively replicates empirical data and answers the research questions posed is never too simple, but it might be too complex.
The model parameters were estimated with an algorithm specifically designed for the model used—even though most of the algorithm was based on generalizable algorithms and principles. The lack of generalizability of the algorithm could be criticized, but since no generic algorithms for calibrating any agent-based model exist, the next best solution is to develop algorithms for specific models or jointly design models and the algorithms for calibrating them. Another problem we faced was identifiability issues. Due to the large number of parameters and the model’s dynamic nature, we found many different solutions with similarly good fit. Furthermore, sensitivity analyses (not presented here) indicate that the parameters should be estimated with high accuracy (the tolerance for maintaining an acceptable fit is less than about 5% of the value ranges of the parameters). Such problems with the estimation of parameter values are expected to be even larger in models with more parameters and more complex dynamics. Thus, if we estimate parameter values only once and/or roughly (or not at all), we should be very cautious when comparing different models or outputs or interpreting the parameter values. Further, theoretical and practical considerations are required to select solutions for further analyses or to reduce the parameter space to be searched by the optimization.
Conclusions
We presented an approach that allows using psychological computer models for explaining real-world phenomena and even forecasting future developments or preparing for future interventions. For this, models have to be designed simple enough for calibration and systematic exploration, but the mechanisms over which real-world interventions would influence the output have to be explicitly modeled. Further, the model design must consider the available data and optimization algorithms or, then, dynamic data have to be gathered and algorithms designed particularly for this model. Results derived from models that were not or only roughly calibrated have to be interpreted with caution, particularly if the interpretation refers to parameter values. We expect most complex dynamic models to have various parameter settings that lead to similar outputs and one has to decide, based on theoretical and practical considerations, which solutions to analyze. Finally, models have to be explored systematically to derive generalizable results. This, in turn, requires that key aspects of the output are expressed in simple indicators, so that the relevant information can be extracted from hundreds or thousands of simulation experiments. Such an approach, then, allows deducing concrete recommendations for how to optimally design real-world campaigns.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
