Abstract
When observing an unfamiliar neighborhood, people use indicators of physical disorder to judge the local community (i.e., community perception), associating them with crime and weak relationships between neighbors. The authors argue that these judgments depend on people’s definition of disorder, which is adapted to their local community. This is tested with an experiment. Undergraduate students from across New York State rated the collective efficacy (i.e., social quality) of neighborhoods from a single city using images of physical structures. Participants reported which features they attended to when making these judgments. Participants were categorized as being from New York City (NYC), NYC suburbs, or the less densely populated upstate region. Images were from an upstate city. Those from NYC attended more to pavement than others. Ratings by those from upstate were most accurate and positive. These results supported the initial hypotheses and suggested that community perception combines heuristics and familiarity to make inferences.
When encountering a new environment, people must quickly determine how best to navigate it. They search the scenery for elements that might provide relevant information and guide them to adaptive decisions. In a forest, for example, individuals are attracted to items that inspire curiosity, like open paths, which suggest an opportunity to look for resources (Kaplan, 1992). Likewise, individuals entering an unfamiliar neighborhood will search for clues about how best to behave. Such decisions take on a new complexity, however, because the individual is encountering not only an environment but also a residential community. Adaptive behavior in this case requires social considerations, and people would need to attend to items that are informative about the local community, and the types of interactions that might occur there. Recent work indicates that people do generate judgments of the local community from elements of the scenery, judgments that then inform their subsequent interactions within the neighborhood. We refer to this process as community perception (O’Brien & Wilson, 2011).
Community perception can only be accurate, however, if the individual is focused on those elements that do in fact reflect the local community’s quality. Taken another way, if two individuals differ in how they define disorder, they may arrive at different interpretations of the same neighborhood. In turn, the individual focusing on more informative features will make more accurate inferences. We argue here that these differences in interpretation will be derived from one’s local environment, leading individuals to focus on those cues that are more prevalent in their daily lives. From this it follows that individuals from different backgrounds (e.g., rural vs. urban) might have definitions of physical disorder tailored to the particularities of their own communities, a process we refer to as local adaptation. The present study evaluates this argument for the specific case of using physical disorder and incivilities to judge the social quality and overall safety of a neighborhood. The first section further develops the potential role and implications of local adaptation in community perception. This culminates in a set of hypotheses that are then tested with an experiment that asked naive participants to evaluate the social quality (operationalized as collective efficacy) of a series of neighborhoods from a single city using only images of physical structures. The undergraduate participants span New York State, and we use its geography to categorize individuals by their background. We are then able to compare the way each group defines disorder, and how this impacts the relative accuracy and valence of their ratings.
Community Perception
When meeting a stranger for the first time, individuals make quick, reasonably accurate assessments about the individual’s personality, lifestyle, and culture, using various cues including appearance, facial expressions, and mannerisms (Ambady, Bernieri, & Richeson, 2000; Borkenau, Mauer, Riemann, Spinath, & Angleitner, 2004). This process, known as person perception, is used to determine the stranger’s quality as a social partner, and how best to interact with them. O’Brien and Wilson (2011) posited that a similar calculus, which they referred to as community perception, would be adaptive when entering an unfamiliar neighborhood, as it too portends social interactions that could be either positive or negative. Houses and neighborhoods bear traces of the attitudes and behaviors of the individuals and communities that reside there (Altman & Chemers, 1984; Nasar, 1989; Werner, Peterson-Lewis, & Brown, 1989), and observers might use this information to develop informed judgments. For example, Harris and Brown (1996) demonstrated that naive participants could accurately infer the territorial commitment of residents to their home or neighborhood (i.e., pride and care about quality and upkeep) from pictures of houses. Participants based their assessments on cues of upkeep, levels of decoration, and other signs of personalization, items that do in fact indicate the territoriality of residents.
One of the most important inferences one might make about an unfamiliar neighborhood is the quality of the social environment, particularly whether there is the potential for criminal threat to one’s person or belongings. This gradient in social quality is largely determined by the ability of the local community to establish and enforce a set of shared expectations for behavior in the neighborhood. Known as collective efficacy, this capacity for informal governance can be instrumental in preventing crime (Kawachi, Kennedy, & Wilkinson, 1999; Sampson, Raudenbush, & Earls, 1997), discouraging juvenile delinquency (Sampson, 1997; Sampson, Raudenbush, & Earls, 1999), and promoting prosocial norms (Wilson, O’Brien, & Sesma, 2009). Neighborhoods with higher levels of collective efficacy also take better care of the public space, whereas those with lower collective efficacy tend to have more litter on the streets and sidewalks, unkempt vegetation, and dilapidated housing (Sampson & Raudenbush, 1999). Such physical incivilities are visible, recognizable elements that one might use to estimate a neighborhood’s safety, as well as its overall strength as a community.
Accordingly, research has demonstrated that people utilize physical incivilities when evaluating the quality of a neighborhood. Studies have repeatedly shown that residents who perceive higher levels of physical disorder also report a greater fear of crime than their neighbors (e.g., Gau & Pratt, 2008; Karakus, McGarrell, & Basibuyuk, 2010; Perkins, Meeks, & Taylor, 1992; Pitner, Yu, & Brown, 2012; Sampson & Raudenbush, 2004). This effect is often found to be equivalent to or stronger than the perception of social incivilities (e.g., public drunkenness) or of one’s personal relationships with neighbors, and is consistently independent of the socioeconomic status of both the individual and the neighborhood. This sense of threat communicated by physical incivilities can have far-reaching consequences, elevating stress and depressive symptoms, diminishing trust and self-esteem, and engendering an overall sense of hopelessness (Haney, 2007; Kruger, Reischl, & Gee, 2007; Ross & Jang, 2000; Ross, Mirowsky, & Pribesh, 2001; Wen, Hawkley, & Cacioppo, 2006).
In an attempt to isolate this perceptual phenomenon experimentally, O’Brien and Wilson (2011) showed participants physical sceneries from across a single city. When asked to rate the collective efficacy of a pictured neighborhood, participants in the first experiment relied entirely on the prevalence of physical incivilities (explaining 90% of the variation between neighborhoods). Their evaluations were in strong agreement with measures of collective efficacy from a survey of neighborhood residents (50% of the variation shared). In the second experiment, participants played a game with monetary implications based on the prisoner’s dilemma with a resident of each of the pictured neighborhoods. In this case, participants treated residents from more disorderly neighborhoods as less trustworthy. Even if told that the resident had already chosen to cooperate, participants still tended to eschew cooperation, meaning they were avoiding social exchange with the neighborhood altogether. Overall, these results described an example of community perception with important behavioral implications; not only were informative cues used to make accurate inferences, they were then translated into social decisions appropriate for a threatening environment.
The work linking physical incivilities to fear describes a general adaptation for navigating neighborhood environments, but it is not clear that the indicators of safety are consistent across communities. Two communities might differ in their infrastructure or norms in such a way that they have distinct expectations for physical upkeep. If the residents of these communities then carry these rubrics with them, it is possible that they will develop contrasting opinions when faced with the same neighborhood. This might be referred to as a localized, or local, adaptation. We posit here that these differences exist, shaped in large part by the structure of public spaces, and that individuals will focus more on those elements that are most prevalent in the spaces with which they are familiar.
Implications of Local Adaptation
Local adaptation in this case rests on the assumption that communities with different elements or contexts will have different definitions of physical disorder. This is not to say that people with different backgrounds will disagree on all elements of maintenance. One might expect, for example, that all cultures have at least some aversion to unsanitary conditions, like the improper disposal of waste. However, there will be some differences, and individuals will probably be more attentive to features that are salient in their daily lives. For example, those from urban areas might place a greater emphasis on pavement when observing a new neighborhood in comparison with those from rural areas, reflecting the prominence of pavement in the one ecology versus the other. Such reasoning could be expanded to any major ecological distinction between two communities.
This has two main implications for how community perception will actually operate when people encounter neighborhoods that do not fully resemble their own, because they have either a different infrastructure or a distinct set of social norms. The first of these is a matter of accuracy: In trying to apply one’s own rubric in an unfamiliar context, one might make errors. For example, take two cities, one where hedge care is an important indicator of community strength and another where it is not as relevant. If a resident of the first city found the second city’s hedges to be unkempt, they might view its neighborhoods as less safe. These impressions, though, are inaccurate because they are derived from elements that do not reflect neighborhood quality in the local environment.
Second, when encountering an unfamiliar scenery, individuals might find it difficult to identify cues that are informative about local safety. Alternatively, they might be subconsciously disoriented by unfamiliar patterns and items. In either case, the resultant uncertainty would make it adaptive to bias oneself toward the behavior that will be less costly if an error is made (Error Management Theory [EMT]; Haselton & Buss, 2000). Taking the case of assessing safety, if one mistakes a safe neighborhood for a dangerous one, the cost is limited to passing up social interactions that may in fact be pleasant. If one mistakes a dangerous neighborhood for a safe one, however, an open, trusting attitude could find itself easily victimized. Thus, it would be adaptive to express a negative bias when judging neighborhoods whose elements are unfamiliar.
The Present Study
The present study uses an experiment to test for local adaptation in community perception. Participants were asked to assess the collective efficacy of unfamiliar neighborhoods from across the city of Binghamton, New York, using only images of physical structures. The study was made possible by the Binghamton Neighborhood Project (BNP), a collaboration between multiple academic disciplines and community groups oriented around the study and improvement of urban life. The BNP database includes a library of images from a citywide assessment of physical incivilities and measures of collective efficacy from a resident survey. Together, these resources provided the images shown to participants and the ability to assess the accuracy of their responses.
Our central thesis requires the comparison of individuals from communities with different structures or social norms. Because the latter is much more difficult to specify, we focus on structures. To this end, the study capitalizes on the unique organization of New York State. The majority of the population lives in the southeast corner, in New York City (NYC) or in its surrounding suburbs (NYC suburbs). These two areas have markedly different levels of population density (~40,000 vs. ~5,000 residents per square mile). As one moves further from NYC, the landscape is increasingly rural, though dotted with small cities. These “upstate” and “downstate” regions are typically treated as distinct. The sample is from a public university, providing participants from these three groups: residents of NYC, residents of NYC suburbs, and residents of upstate. The neighborhoods used as stimuli in the experiment were from a single upstate city. For this reason, we argue that the structure and features of these neighborhoods will be most familiar to those from that part of the state. With this experimental design in hand, we propose the following hypotheses:
Hypothesis 1: The three groups will vary in their definition of physical disorder, in accordance with their respective ecologies.
Hypothesis 1a: Residents of NYC will attend more to pavement, given the greater concentration of paved spaces in densely populated urban areas. Those other two groups will place more emphasis on vegetation.
Hypothesis 1b: Examples of poor sanitation, like loose garbage, will be an important signal of collective efficacy to all groups.
Hypothesis 2: The three groups will vary in their accuracy when assessing a neighborhood’s collective efficacy, and those from upstate will be the most accurate.
Hypothesis 3: Following EMT, lack of familiarity will lead individuals to evaluate neighborhoods with a negative bias relative to those more comfortable. In other words, those from downstate will provide overall lower ratings of collective efficacy than those from upstate.
Hypothesis 4: Different definitions of disorder will underlie much of the variation in accuracy across backgrounds, but not necessarily the variation in perceived quality. This follows as accuracy depends on whether an individual is focusing on the correct elements, while perceived quality is a response to familiarity with the elements of the landscape.
Readers might note that we have not provided an explicit definition for the term neighborhood. To this point, we have used it to denote a largely self-contained, place-based community. We do not make claims here as to what the exact geographic scale of a neighborhood should be, in part because urban researchers have long debated and continue to debate this very question (e.g., Coulton, Korbin, Chan, & Su, 2001; Galster, 2001; Grannis, 2009; Suttles, 1972). What is important in the present study is that, when prompted to judge the characteristics of a neighborhood, the participants conceptualize a community that has reliable social patterns and is responsible for the space pictured in the images. Studies that have asked people to describe their own neighborhood suggest that this is how people intuitively define the term (Guest & Lee, 1984; Sastry, Pebley, & Zonta, 2002); thus, we feel confident that the proceeding experiment indeed operationalizes the way individuals judge residential communities.
Method
Participants
Participants included 128 students in an introductory biology/anthropology course at a public university in New York State. Participants reported their hometown. Analysis was restricted to residents of New York State, or a bordering region of an adjacent state (n = 105), who were then categorized as being from NYC, NYC suburbs, or upstate. The first group was defined as those living in NYC (n = 30), the second as those living within two counties of NYC, which approximates the crossover from continuous to patchy settlement (n = 43), and the third as those living beyond this line (n = 32).
Materials
The city of Binghamton contains 63 census block groups (CBGs) that are used to approximate neighborhoods in most BNP studies, of which 62 have substantial residential populations. From these 62, 128 addresses were chosen at random, with each CBG represented at least twice. In September 2009, four digital photos were taken at each of these addresses: one facing the address, one looking across the street from the address, and one looking each way down the street. When placed together, these approximate a 360° view from the street in front of the address (see Figure 1 for an example). No photo included images of people. Of these 128 addresses, 25 were randomly selected, each from a different CBG, which were pared down to 20 that maximized socioeconomic and geographic variation. These 20 acted as the stimuli in the experiment.

Example images from one neighborhood as seen by participants
Procedure
The core of the experiment asked participants to evaluate the 20 addresses specifically chosen for this study. Before this, they were shown 20 other images randomly selected from the image library. These images appeared for 5 s apiece, and participants watched them without responding in any way, giving them the opportunity to cognitively adjust to observing neighborhoods. Subsequently, the 20 target neighborhoods were displayed. Each of the four pictures taken from a neighborhood appeared individually for 5 s, followed by a view of all four images together (as seen in Figure 1) for 30 s. During this 30-s period, the participants rated the collective efficacy of the neighborhood. After viewing and rating all neighborhoods, participants reported the amount of attention they paid to each of a collection of elements in the images (e.g., lawn care). This methodology was approved by the university’s Human Subjects Resource Review Committee.
Measures
Neighborhoods
Objective disorder
Five independent raters provided objective ratings of the physical features listed in Table 1 for each image in the set of 128 addresses. None of these raters was a participant in the experiment. This procedure was identical to that described in O’Brien and Wilson (2011). Ratings were left blank if nonapplicable for a given image (e.g., rating lawn quality in a picture of a street). Cronbach’s alphas indicated that ratings were consistent across raters for all variables (see Table 2), permitting the averaging of all ratings of an image to create an image-specific score for each measure. Scores were then averaged across the four images taken at each address to create an address-specific score for objective disorder. Table 2 reports correlations between these scores for the set of 128 addresses, and the subset of 20 used in this study.
Neighborhood Physical Features, Their Perceived Importance for Assessing Safety, and Correlations Between Them.
Note: Correlations, means, and standard deviations based on responses from complete sample of 128 participants.
On a 5-point Likert-type scale.
p < .1. **p < .05. ***p < .01. ****p < .001.
Correlations Between Physical Features in Images, Calculated at the Address Level.
Note: Above diagonal includes all photographed addresses, below diagonal the subset used in this study. Diagonal contains number of addresses that include a score for each item. See Table 1 for unabbreviated feature descriptions.
p < .1. **p < .05. ***p < .01. ****p < .001.
Resident reports of collective efficacy
In January 2009, 703 students at Binghamton High School responded to a modified version of the Developmental Assets Profile (DAP), a survey developed by Search Institute (http://www.search-institute.org/) to assess the quality of life in adolescents (see O’Brien, Gallup, & Wilson, 2011 for more information). This survey included measures of social cohesion (three items: “People in my neighborhood are willing to help each other”; “I have neighbors who help watch out for me”; “I have good neighbors who help me succeed.”) and social control (three items: “If there were a fight in my neighborhood, neighbors would break it up”; “If children were disrespecting an adult in my neighborhood, other adults would stop them”; “If children were skipping school and hanging out in my neighborhood, adults would tell them to go to school.”) in one’s neighborhood. These two scales reflect the components of collective efficacy and were drawn in large part from Sampson et al. (1997). The scales were combined to form a single measure of collective efficacy, placed on a 0 to 100 scale (α = .83). Using ArcGIS (v. 9.6), each response was mapped to the CBG of residence. A neighborhood measure of collective efficacy was calculated by averaging the responses of all respondents residing in a CBG. Descriptive statistics for this variable in the 20 neighborhood sample were similar to those for the whole city (MTotal = 48.69, SDTotal = 11.59; MSample = 48.07, SDSample = 13.67).
Respondents
Participant judgments of collective efficacy
Participants rated the collective efficacy of the pictured neighborhood using a five-item scale (two reflecting social cohesion—“People around here are willing to help their neighbors” and “There are adults in this neighborhood that children can look up to”—and three reflecting social control—“This is a safe place”; “If there were a fight in this neighborhood, neighbors would break it up”; and “If children in this neighborhood were skipping school and hanging out on a street corner, neighbors would take action”). Four of these items were taken directly from those used in the survey of high school residents, and the fifth (“This is a safe neighborhood.”) was included as a more general question capturing a participant’s overall evaluation of the neighborhood. All responses were on a 5-point Likert-type scale (1 = strongly disagree, 5 = strongly agree). The construct had strong internal reliability (Cronbach’s α = .88). For all respondents, a rating score was calculated for each neighborhood by averaging the responses to the five items and placing them on a 0 to 100 scale. Participants left blank any address they believed they recognized (final N = 2,099 responses).
Attention to objective features
Participants reported the extent to which they used the quality of different physical elements of a neighborhood (e.g., lawn care; see Table 1 for a complete list, descriptive statistics and correlations between them) to make assessments during the experiment. These ratings were on a 5-point Likert-type scale (1 = not at all, 5 = very much).
Accuracy
To create a measure of accuracy, we used Hierarchical Linear Modeling (HLM; see analysis) to nest responses within participants and used resident reports of collective efficacy to predict responses across neighborhoods. This analysis estimates a separate linear relationship between resident reported collective efficacy of a neighborhood and each participant’s assessment across neighborhoods. The residual error for each response indicates deviation from this line. A participant has 19 such residual scores (one for each neighborhood; one neighborhood excluded, see below), and larger absolute values mean greater deviation from a linear relationship that matches the variation in actual neighborhood quality. The average squared residual was then calculated for each individual, creating a measure of accuracy.
Analysis
ArcGIS (v. 9.6) was used to link the participant ratings of each address to the appropriate CBG, creating a design with responses nested within neighborhoods. HLM 6.06 (Raudenbush, Bryk, Cheong, Congdon, & du Toit, 2004) was used to partition the variance associated with descriptors of raters (first-level; e.g., one’s background) from descriptors of neighborhoods (second-level; e.g., collective efficacy as reported by residents). HLM does not calculate standardized Beta’s, so we report unstandardized Beta’s, t statistics, effect sizes, and significance levels for these tests.
Results
Question 1: Defining Disorder
Categorizing Cues
Participants reported the extent to which they used the quality of eight different features when making assessments of a neighborhood’s social dynamics. Many of these features are similar in nature, and their relative importance often correlated across individuals (see Table 1). For example, one highly attentive to the quality of street pavement was similarly interested in driveway pavement (r = .68, p < .001). A principal components analysis was used to reduce these eight variables into a smaller set of measures. This analysis utilized the entire original sample because the question is not related to hypotheses based on the geography of New York State and because a principal components analysis requires a large sample size (Tabachnick & Fidell, 2006; N = 127 because one individual failed to rate the importance of garbage). The results, shown in Table 3, suggest that people categorize these features into three groups: vegetation, pavement, and paint/garbage. This last we will refer to as sanitation.
Principal Components Using Participant Responses to Reduce Image Features to Cognitive Categories.
Note: Only loadings >.4 reported (N = 127). The values are in bold to differentiate the measures from the other parameters in the table.
Based on this result, we calculated three new individual-level variables measuring attention to forms of disorder (i.e., attention to vegetation, pavement, sanitation), averaging the appropriate measures for each participant, for example, attention to pavement = Mean (driveway, street, sidewalk). The same categorization and process was used to calculate three indicators of objective disorder for each address (i.e., upkeep of vegetation, pavement, sanitation). For one address that lacked sidewalks, regression imputation was used to calculate a pavement score. Correlations between these measures and resident reports of collective efficacy are reported in Table 4.
Correlations Between Address-Level Measures of Visible Disorder and Resident Reports of Collective Efficacy.
Note: Above diagonal includes all neighborhoods (i.e., census block groups; n = 62), below diagonal the subset used in this study (n = 20). Collective efficacy measure derived from 2009 administration of Developmental Assets Profile survey to Binghamton residents.
p < .1. **p < .05. ***p < .01. ****p < .001.
Attention to Features Across and Between Groups
The three measures of attention to disorder acted as dependent variables in a within- and between-participants ANOVA that tested two things simultaneously: Are members of all groups more attentive to some features than to others? and Do people’s backgrounds predict which features they focused on? Answering the first question, there was a general tendency for certain features to be favored over others when assessing neighborhood quality, F(2, 202) = 19.21, p < .001. Post hoc contrasts of within-participants measures demonstrated that participants were more attentive to sanitation than pavement or vegetation (p < .01 for both comparisons; see Figure 2).

Attention to image features as reported by participants across groups.
Regarding the second question, a significant interaction between background and feature type indicated that there were differences between groups in how strongly they weighted certain features, F(4, 202) = 3.01, p < .05. We investigated this further by decomposing the initial analysis into three between-participant ANOVAs, each comparing the attention to a particular category of features across groups. Of these, only attention to pavement varied across groups, F(2, 103) = 3.95, p < .05 (see Figure 2). This effect was driven by a tendency of those from NYC to pay more attention to pavement quality than those from either NYC suburbs or upstate (Student Newman-Keuls [SNK] categorization test, p < .05). There was no statistical difference between the NYC suburbs and upstate groups. Attention to vegetation, F(2, 103) =0.46, p = ns, and sanitation, F(2, 103) = 2.15, p = ns, did not vary substantially between groups.
Question 2: Participant Accuracy in Judging Collective Efficacy
Whole Sample
Before analyzing differences in accuracy between groups, we ran an HLM model using resident reports of collective efficacy to predict participants’ responses to the photographed neighborhoods. Participants were only moderately accurate (B = .36, t18 = 2.13, e.s. = .44, p < .05), but post hoc analysis of the residuals showed that this relationship was unduly weakened by a single address that contained views of a side alley, the side of a building, and a river. Kaplan (1992) found that ambiguous urban scenes like this one can arouse ambivalence, thus we removed this address before proceeding with analysis. 1 With this adjustment, the agreement between participant and resident ratings across neighborhoods rose dramatically, with a shared variance of 41% (B = .69, t17 = 3.80, e.s. = .67, p < .01). This replicated the basic result of O’Brien and Wilson (2011) with a similar effect size (.67 vs. .73).
Differences in Accuracy Between Groups
To test for differential accuracy between groups, we ran three separate regressions allowing resident ratings of a neighborhood’s collective efficacy to predict participant judgments within each group. HLM’s parameter for this test would be based on the number of addresses that were compared at the second level, which would yield a very low number of degrees of freedom and limited power to identify differences between groups. For this reason, we use standard regressions with individual responses as the unit of analysis. 2 Individuals from upstate exhibited the greatest accuracy (βdf = 602 = .49, R2 = .24, p < .001), followed by those from NYC suburbs (βdf = 814 = .41, R2 = .17, p < .001) and then those from NYC (βdf = 562 = .35, R2 = .12, p < .001). Comparing these R2s, judgments provided by the upstate group were the most accurate (compared with NYC suburbs, p < .05; compared with NYC, p < .001; using the test recommended by Cohen, Cohen, West, & Aiken, 2003), but there was no significant difference in accuracy between the other two groups. Because HLM controls for individual-level noise when calculating neighborhood-level correlations, switching to an individual-level design artificially lowers R2 as a measure of accuracy across neighborhoods. To permit comparison with the analysis of the entire sample reported above, we provide the R2s for each group if run in HLM: upstate = .55; NYC suburbs = .44; NYC = .32.
Question 3: Biases in Judgment Across Groups
To test whether groups less familiar with Binghamton rated neighborhoods with a negative bias, we compared average ratings across groups. A negative bias would be reflected by lower average ratings. An ANOVA demonstrated that the average judgment of a neighborhood varied across participant background, F(2, 1981) = 4.62, p < .01. Post hoc tests found that judgments made by individuals from upstate were more positive (Mupstate = 66.58, MNYC suburbs = 64.35, MNYC = 63.53; SNK test categorized rural separately from the other two at p < .05).
Question 4: How Objective Features Explain Judgments of Collective Efficacy
Objective Features and Ratings in the Whole Sample
To test whether participant responses to a given address were a direct response to cues of disorder, we ran an HLM in which the quality of an address’ pavement, vegetation, and sanitation predicted participant ratings of collective efficacy (see Model 1 of Table 5). Only vegetation quality was positively, significantly associated with responses (B = 5.38, e.s. = .54, p < .05), but note the large effect sizes for the other parameters (sanitation: e.s. = .30; pavement: e.s. = .26), which are likely nonsignificant given the low degrees of freedom (df = 16). This model explains 71% of the variance in ratings across neighborhoods, as opposed to the 41% explained by resident reports.
Parameter Estimates From Hierarchical Linear Models Using Visible Features, Participant Background, and Participant-Reported Attention to Visible Features to Predict Judgments of Neighborhood Collective Efficacy.
Note: NYC = New York City. N = 1,984 responses nested in 19 neighborhoods.
Effect sizes calculated using
Parameter compares with individuals from upstate.
Measure of attention paid to each set of features, as reported by participant.
Measure of maintenance of features in image, as objectively evaluated.
Interaction between individual report of attention to a set of features, and the quality of those features in images.
p < .1. **p < .05. ***p < .01. ****p < .001.
Validity of People’s Self-Reported Attention to Objective Features
Although it seems apparent that neighborhood ratings are generally in response to the objective upkeep visible in the image, it is not clear that people are actually attending to the specific features they claim to be. To test this, we ran a new regression model introducing six variables: the individual respondent’s interest in each of the three types of disorder, and the cross-level interaction effects between each of these and the corresponding measure at the image level (e.g., individual’s attention to pavement × image’s objective pavement score). A significant positive result for any of these interactions would imply that individuals who reported greater attention to that set of features were indeed more sensitive to variation in its maintenance. As can be seen in Model 2 of Table 5, all three interactions were positive and significant, indicating that participants were aware of which features influenced their ratings of neighborhood quality. This model also found that individuals who reported paying greater attention to vegetation quality had higher assessments on average (B = 1.23, e.s. = .18, p < .001), and those who paid greater attention to pavement made lower ratings (B = −0.43, e.s. = .09, p < .001).
Attention to Image Features and Variation in Accuracy and Average Ratings Between Groups
Pavement is not a good predictor of collective efficacy in Binghamton (r = .12, p = ns), and overemphasizing it could lower one’s accuracy in the current context. Also, those paying greater attention to pavement made more negative assessments of neighborhoods. Because attention to pavement quality is strongest in those from NYC, it may account for some of the variation in accuracy and bias between groups. Evaluating this requires two separate tests, one for accuracy and the other for average ratings. The latter entails a simple modification to the last model, thus we present it first, and follow by testing pavement’s role in influencing accuracy, which is somewhat more complicated.
Variation in average ratings
We added two dummy variables to the HLM, one for those from NYC and the other for those from NYC suburbs (individuals from upstate acted as the reference group) to estimate the baseline effect of participant background on one’s average response (see Model 3 of Table 5). The results suggest that those from NYC suburbs rated a given neighborhood’s collective efficacy as about 2 points lower than did their counterparts from upstate (B = −2.18, e.s. = .06, p < .001). The same parameter was estimated at 3 points for the NYC group (B = −2.99, e.s. = .08, p < .001).
Added to this model were the three variables measuring an individual’s attention to image features, and the corresponding cross-level interactions, as in Model 2 above (see Model 4 of Table 5). These inclusions lowered the parameter for NYC by 20% (B = −2.45, e.s. = .06, p < .001) and the parameter for NYC suburbs by 11% (B = −1.94, e.s. = .06, p < .001). Note that this differential effect was expected as those from NYC, and not those from NYC suburbs, indicated greater attention to pavement upkeep when judging collective efficacy.
Variation in accuracy
To test whether differences in accuracy across backgrounds could be explained by attention to pavement, we first created a measure of total error for each individual. An ANOVA found that accuracy differed between the three groups, F(2, 102) = 5.05, η = .30, p < .01, explaining 9% of the variation across individuals. We then added an individual’s self-reported focus on pavement to the model as a covariate. It had a marginal association with inaccuracy, F(1, 101) = 3.72, η = .18, p < .10, and mediated about one fifth of the effect of background, F(2, 101) = 3.27, η = .25, p < .05. These results are similar to those for overall ratings presented above, implying that a small but substantial amount of the difference between groups in accuracy and valence can be attributed to an elevated attention to pavement in those from NYC. Over and above this effect, however, there is still an unexplained tendency of those from NYC and NYC suburbs to rate neighborhoods more negatively and inaccurately.
Discussion
It is well established that people use physical incivilities as indicators of a neighborhood’s social quality, but, we have argued, this process is subject to what a given individual considers a “physical incivility,” or, more globally, how said individual defines “physical disorder.” If this holds, then it might be that such definitions are derived predominantly from one’s native environment, creating a state of local adaptation. This argument generated four main hypotheses, each of which was borne out by the results of the experiment, at least in part if not in whole. Individuals from different backgrounds paid greater attention to those elements that were more salient in their native environment (Hypothesis 1), and accuracy and average ratings were lower in those from regions more removed from the neighborhoods being observed (Hypotheses 2 and 3). The difference in attention then partially explained the differences in accuracy and average ratings across groups (Hypothesis 4).
In sum, these results provide substantial evidence for the process of local adaptation, but they depict a more complex story than that presented at the outset. When comparing people’s attention to different types of features, those from NYC were distinguished from the other groups by the attention they paid to pavement. A different split emerged from the analysis of accuracy and average ratings, however, classing those from NYC and its suburbs together, but apart from upstate participants. The original hypotheses posited that the emphasis of certain types of physical disorder would be directly responsible for differences in accuracy, but this would have predicted that the division of groups be the same for both of these analyses. Instead, these results suggest that local adaptation may manifest itself at two levels, with heuristics simplifying sceneries, but familiarity playing an important role in understanding them.
Heuristics and Categories of Physical Incivilities
In any given situation, the amount of potentially available information can be overwhelming. For this reason, cognitive efficiency depends on a suite of heuristics that reduce myriad cues to a more manageable set of patterns. An individual can then weight certain patterns as more informative than others (Kahneman, 2011). The first of these processes, reduction, is apparent from the way participants categorized the various attributes of a neighborhood’s physical scenery into three groupings: pavement, vegetation, and sanitation. Individuals then weighted the importance of physical incivilities in each of these domains when evaluating the social quality of a neighborhood. These weights were captured by the attention respondents reported placing on each of the neighborhood’s features. Importantly, these self-reported weights corresponded to the way ratings varied across images, as demonstrated by the interaction effects in Models 2 and 4.
Central to the point of the present study is how these weights varied within and across individuals. For example, participants from all three backgrounds saw sanitation as the most important indicator of community dynamics, an attitude that could be cross-cultural. Being that all residential areas have at least semipermanent structures and occupants who produce waste, these areas of maintenance would typically be indicative of the energy residents have invested in their domiciles and community. This is further borne out by the strong correlation between paint quality, levels of litter, and collective efficacy (see Table 4). Not all categories were viewed as equally important across backgrounds, however. As predicted, pavement quality was a more salient indicator of neighborhood quality for those from NYC than participants from the other two groups. The other hypothesized difference, that those from less densely populated areas would place greater emphasis on vegetation, was not supported. In retrospect, this prediction may have been somewhat naive, as urban areas do feature parks, sidewalk vegetation, and, in some areas, small lawns.
The attention paid to pavement by those from NYC is telling in regard to the nature of local adaptation. One might argue that individuals should focus exclusively on those things that are most informative in the local environment, but pavement, though it accounts for substantial amount of the scenery in urban centers, does not hold a strong relationship with collective efficacy. Streets and sidewalks in most cities, NYC being no exception, are not maintained by residents but by municipal agencies. Although there is evidence that neighborhoods with greater collective efficacy tend to have more reliable pathways to procuring government services (Sampson, 2012), it is not clear that this indirect path creates a reliable correlation between pavement upkeep and collective efficacy. In the data presented here, no such correlation is visible. Heuristics, however, are structured as much for efficiency as they are for accuracy. If the goal is to quickly evaluate the level of physical disorder, a locally adapted heuristic might emphasize some elements merely because they are prevalent. This is not to say that people do not also use informational content to privilege some cues over others—indeed, sanitation, strongly associated with collective efficacy, was seen as the most important indicator of social quality by all groups—simply that both processes lead to the weighting of these various elements.
Familiarity, Accuracy, and Average Ratings
Heuristics played an important role in how groups processed neighborhood sceneries, accounting for 20% of the differences in accuracy and average ratings between those from NYC and other groups. Nonetheless, this left a substantial amount of variation to explain. Furthermore, individuals from NYC’s suburbs exhibited no such bias toward pavement but were just as prone to inaccuracy and negative bias as their NYC counterparts. We propose that these differences can be best explained by an interplay between heuristics and one’s familiarity with the setting.
The communities of NYC, its suburbs, and upstate vary on more than just their relative proportions of vegetation and pavement. They might differ in any number of ways, from the structure of buildings, to the way individuals decorate their space, to the expectations for how a lawn or house “should” look. Such differences can make for limited legibility, or the ability to interpret the landscape, when one enters an unfamiliar neighborhood outside of her own region. When those from NYC suburbs evaluated the social dynamics of neighborhoods, their interpretations were less accurate than those from upstate, despite attending to the same categories of incivilities. Although the overall heuristic was weighted the same way, a lack of familiarity with the specific cues of an upstate city may have hampered their ability to make accurate inferences.
The impacts of familiarity and its consequent legibility are generally limited to accuracy when individuals are merely pursuing information. For example, Nasar, Stamps, and Hanyu (2005) took images of city halls and museums from a variety of cities and asked participants to identify which function each building fulfilled. With little experience with the specific buildings, individuals were just barely more accurate than expected by chance. Evaluating the social quality of a neighborhood, however, plays the more vital role of determining one’s immediate safety. Being unable to interpret the elements of the scenery will impair one’s ability to elect an optimal strategy for navigating the neighborhood. As mentioned above, EMT (Haselton & Buss, 2000) provides a model of how people might operate in such a state of uncertainty. EMT argues that people should be biased toward those decisions that are less likely to be costly in the case of a mistake. This reasoning makes it easy to see why one in a neighborhood with unfamiliar characteristics might be biased to perceive it as threatening: Mistaking danger for safety could have far greater consequences than misreading a safe situation as a dangerous one. This is what we see in the present study, as those from NYC and its suburbs provided lower overall ratings of the neighborhoods of Binghamton, an upstate city.
In making this argument, it is important to distinguish between what EMT does and does not say. It provides a rubric for adaptive outcomes but does not prescribe the cognitive or developmental processes by which they arise. Although the results here demonstrate that those not from upstate negatively biased their judgments of safety when unequipped to make an accurate assessment, they do not indicate a mechanism for generating this pattern. If we were to speculate, mere exposure seems to be the most likely candidate, whereby individuals favor elements with which they are accustomed over those with which they are not (Zajonc, 1968). Although this does not describe a conscious link between accuracy and valence, it would give rise to the crucial correlation: Familiarity creates both legibility and preference, whereby a lack of familiarity hinders accuracy and motivates a negative bias.
An alternative explanation for this pattern could be tensions between the different regions of New York State. The greater NYC region is more affluent, and if its residents have a negative impression of Binghamton, this could depress judgments of the pictured neighborhoods. This would not account, however, for differences in accuracy across groups. Thoroughly distinguishing between these two interpretations would require a more complex research design in which both participants and images are sampled from two or more locales. All participants would then respond to all images, allowing a robust test of the impact of familiarity. The main difficulty with this design would be the need for measures of collective efficacy for the neighborhoods from both regions.
Practical Implications
Evaluations of neighborhood safety are associated with a number of extenuating outcomes. They lead to lower social and financial investment (e.g., Nasar, 1990; O’Brien & Wilson, 2011). They also have an apparent relationship with physiological processes, like stress and depression (e.g., Wen et al., 2006). In addition, these judgments are not limited simply to perceived safety but are also used to develop other, less founded opinions about a neighborhood and its residents (O’Brien & Wilson, 2011). If individuals from different backgrounds define and perceive physical disorder differently, there could be meaningful consequences in one or all of these areas when distinct communities live alongside each other.
At the individual level, those residing in a community with a structure different from the one to which they are accustomed may regularly feel a sense of disorientation and may have a negatively biased view of local safety. Modern immigration presents an opportunity to explore this question. Churchman and Mitrani’s (1997) study of Russian–Israeli immigrants found that differences in physical structure and organization between Russian and Israeli neighborhoods moderated the extent to which the immigrant participants felt connected to their new neighbors. At the societal level, these differences in perception might exacerbate intergroup tensions. For example, norms regarding casual socialization vary across neighborhoods of different socioeconomic statuses; rich neighborhoods socialize in private and poorer neighborhoods make use of the public space to socialize. Ethnographic work indicates that affluent residents often use this as evidence that the poor neighborhoods are dangerous and unregulated (Duneier, 1999; Fagan & Davies, 2000; Stinchcombe, 1963). This is just one example of how differences in norms across adjacent communities can lead to perceptual judgments, though many such examples seem plausible.
Conclusion
A variety of fields concerned with the function and health of urban neighborhoods have come to focus on the way disorder can directly affect the residents of a neighborhood. Little is known, however, about the cognitive and developmental mechanisms that mediate these effects. The present study has demonstrated how one’s local experiences shape community perception, providing the raw information for making interpretations (familiarity) and a rubric for weighting categories of elements (heuristics). We hope that further research explores these areas, helping to clarify the role physical disorder plays in everyday urban life.
Footnotes
Acknowledgements
We would like to thank Andrew Gallup and Yasha Hartberg for comments on earlier versions of this manuscript, and two anonymous reviewers for their suggestions. Thanks also goes to Sarah Fecht, Eliana Frim, Joe Pickerill, Kevin Ralston, and Robert Stark, who coded objective image features; Alexandra Breines, Samantha Calem, Jeremy Cooper, Cory Jankow, Cynthia Rivera, Adam Toporovsky, and Anna Yeo for assistance in administering the experiment; and the Binghamton Neighborhood Project community and its partners, particularly the Binghamton City School District.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
