Abstract
This research uses high-density anonymised mobile phone application (MPA) global-positioning system (GPS) data to describe exposure to racial diversity in different social contexts with an aim to clarify the mechanism linking residential diversity to opportunities for diverse social interactions. In particular, it explores the hypothesis that a diverse residential context does not lead to diverse social contact by comparing three exposure measures – residential, observed and interaction – on the census block group level in Chicago. In doing so, it also explores the contribution of activity spaces to opportunities for diverse social contact. The findings show that the exposure to opportunities for diverse social contact measured by MPA data is generally higher than what is implied by residential census data, especially in areas of high residential segregation in the city. Further, measures using MPA data reveal more spatiotemporal heterogeneity of exposure than that implied by the residential context.
Introduction
The proposal that exposure to neighbourhood diversity improves socioeconomic outcomes through social and physical integration is central to current theory and practice in urban development and housing policy in North America. It is seen as a response to the deleterious consequences of socioeconomic and racial segregation at the neighbourhood level, especially in the context of residential segregation (Freeman, 2009; Yinger, 1986), which can mediate social processes and amplify the concentration of poverty (Crowder et al., 2012; Wilson, 1987). Neighbourhood diversity is also the central aim of many affordable housing measures (Goetz, 2003; Schwartz, 2014). Its efficacy, however, is often challenged on the insufficient empirical grounding to explain the mechanism that links diversity of context to actual diverse social exposure and contact (Galster, 2012; Galster and Friedrichs, 2015).
Previous research on housing mix suggests diverse social contact is less than expected given neighbourhood demographics (Blokland and Van Eijk, 2010; Butler, 2003; Robson and Butler, 2001). While these studies provide meaningful insight into the social processes and networks that compose neighbourhoods, they are often case studies or ethnographic studies that may be contingent on particular social, political and institutional contexts. Quantitative methods used to describe neighbourhood diversity generally rely on census demographics, presenting an epistemological challenge as census data typically only reflect the residential context despite the diversity of activity spaces – the physical locations individuals have direct contact with as a result of day-to-day activities (Golledge and Stimson, 1997; Horton and Reynolds, 1971; Lynch, 1960). This methodological divide potentially abets contradictory conclusions and theoretical divisions.
This analysis uses high-density anonymised MPA GPS (simplified to ‘MPA’ from here) data to detect activity spaces and test whether exposure to racial diversity in a neighbourhood differs across three types of exposure measures, which I consider to delineate three increasingly representative opportunities of social interaction: residential exposure as measured by census data, observed exposure as measured by all observed activity in a neighbourhood and interaction exposure, which measures the diversity of opportunities for interactions as defined by a space-time activity space overlap of two or more people in the MPA data set. Tracing the links and differences between these three measures can clarify the mechanism by which residential and observed social context translate into opportunities for social interaction.
This study seeks to investigate the following research questions: Can MPA data be representative of the urban population? How do the three exposure measures differ in their estimated exposure to racial diversity? If actual diverse social contact occurs less often than presumed in a diverse neighbourhood, as suggested by previous literature, is this evident in how interaction exposure behaves relative to residential and observed exposure? I hypothesise that:
The MPA data can be used to sufficiently replicate Chicago’s residential population.
Observed exposure to racial diversity will be higher than residential exposure as it represents a wider range of activity spaces within the neighbourhood.
Interaction exposure will be lower than residential exposure, as previous literature suggests that actual social interaction is less diverse than residential context might indicate.
If hypothesis (1) can be validated, it can enable us to test how these three different types of social contexts relate to one another. A wider range of activities and opportunities of encounter should be considered for a more accurate characterisation of segregation dynamics beyond the residential context and MPA data may be sufficiently robust in characterising these contexts. Though the analysis remains at the level of measuring opportunities for social interaction rather than actual contact, I suggest that the likelihood of contact is a function of physical proximity. Developing a better understanding of (2) and (3) will shed light on the gap between our theoretical and empirical knowledge of social interaction mechanisms and their relation to neighbourhood diversity.
Background
Exposure to neighbourhood diversity
The neighbourhood context has material and social influences on residents in terms of access to education, health, public and private services, crime and employment opportunities. Previous research has shown that residential location encodes spatially manifested inequities such as exposure to particular environmental conditions (Bowen et al., 1995), differential access to services (Altschuler et al., 2004), amenities (Lee and Lin, 2017), as well as opportunities (Galster and Killen, 1995). In the USA context in particular, the legacy of concentrations of poverty in racially segregated, lower-income neighbourhoods (Cutler et al., 2002; Wilson, 1987) has motivated measures to deconcentrate poverty and promote social integration, typically enacted through housing (Briggs, 2005; Galster, 2002). Studies suggest that improved economic opportunities (Cutler et al., 2008), intergenerational prospects (Chetty et al., 2016), social trust (Abascal and Baldassarri, 2015; Putnam, 2007) and opportunities for professional mobility (Galster et al., 2008) are all associated with neighbourhood diversity.
Understanding the social contact mechanism
The primary mechanism through which neighbourhood diversity is theorised to improve social outcomes is a link between increased exposure to diversity and increased diversity of social interactions (Galster and Friedrichs, 2015). The effects of this link have not always been clear or consistent, or they suggest only small benefits (Manley et al., 2011; Ostendorf et al., 2001). For example, Crowder et al. (2012) find that, despite the growing prevalence of multiethnic ‘global neighbourhoods’ (Logan and Zhang, 2010), social integration and mobility remain stratified along racial lines.
Galster and Friedrichs (2015) propose that the diverse social contact implied by neighbourhood diversity is rarely present. Previous literature suggests diverse neighbourhoods may not be conducive to actual exposure to diversity. True social mix is, at best, mediated by race, class and ethnicity divisions (Blokland and Van Eijk, 2010; Galster et al., 2010; Robson and Butler, 2001), social network distance (Arthurson, 2012), or different time schedules (Fraser et al., 2013). Robson and Butler (2001) describe the ‘tectonic’ nature of the parallel and non-integrated social relations within gentrifying neighbourhoods that are purportedly in favour of diverse racial and ethnicity diversity. The incongruity between the theory and evidence begs clarification, with Galster (2002) suggesting the necessary and sufficient conditions for interaction may be more stringent than expected. Though confirmed interactions are not investigated here, this analysis isolates spatiotemporal overlaps of activity spaces as greater opportunities for interaction.
Measuring exposure to neighbourhood diversity
Beyond the residential context
Residential context as typically defined by census data and administrative boundaries has a sociological meaning; however, Kwan (2013) notes that the focus on residential census data is static and limits our scope in describing daily activities across a range of activity spaces. Residents often do not even perceive or consider administrative neighbourhood boundaries in activity decisions (Coulton et al., 2001, 2004). Revisions of traditional notions of residential diversity have developed a more ego-centric and temporally dynamic characterisation of the activity space (Boterman and Musterd, 2016; Browning and Soller, 2014; Candipan et al., 2021; Kwan, 2009, 2013; Matthews, 2011; Vallée et al., 2010; Wong and Shaw, 2011). In particular, Browning and Soller (2014) study routine social exposure to activity spaces and conclude that shared local exposure in activity spaces is conducive to higher levels of social capital. Similar to this paper’s questions, Jones and Pebley (2014) use granular survey data to ask whether residential neighbourhoods are representative social contexts. They find the activity space landscape more heterogeneous than residential contexts.
Using mobile phone GPS data to understand neighbourhood diversity
Anonymised MPA data represent actual activities across time and space, including but not limited to the residential one. Large geographic coverage of mobile phone data and the possibility of scaling mobility counts can facilitate comparison across different regions. Previous studies using call detailed record (CDR) data have shown the accuracy of population estimates against actual population counts (Deville et al., 2014; Jiang et al., 2017; Liu et al., 2018; Yao et al., 2017), which may be used to improve understanding of dynamic urban flows (Reades et al., 2007; Yuan and Raubal, 2012) and the feasibility of using these data to represent individual activity spaces (Calabrese et al., 2013; Iovan et al., 2013; Jiang et al., 2017). Whereas CDR data have a coarse accuracy of around 200–300 m (Jiang et al., 2013), MPA data have greater accuracy, with a median accuracy of 12 m in this paper.
Other recent studies using GPS data have demonstrated racialised or class-based social-spatial isolation (Palmer et al., 2013; Phillips et al., 2019; Wang et al., 2018), the importance of non-residential contexts in understanding exposure (Jones and Pebley, 2014; York Cornwell and Cagney, 2017), and the spatiotemporal heterogeneity of exposure even for residents of the same neighbourhood (Browning et al., 2017a).
Data and methods
This study analyses three types of exposure to racial diversity on the census block group level: residential exposure as implied by the 2014–2018 American Community Survey (ACS) demographic data, observed exposure representing the empirically measured stays in the neighbourhood using MPA data, and interaction exposure, also using MPA data, where individuals share a space-time overlap in activity spaces, creating opportunities for social interaction.
In this data set, there are approximately 40 million data points daily, with the analysis period between 1 July and 31 October 2019. The final data set includes home locations for 87,631 residents in the city with a total of 1,644,955 stays and 29,888 unique interaction clusters. While this ultimately represents 3.4% of Chicago’s 2016 population (as the representative population for the 2014–2018 ACS data) within the 2103 residential census blocks studied, with some potential selection bias, an analysis of data representativeness against the ACS population data (detailed in the following section) suggests an unbiased estimation of population can be created.
Data pre-processing
The use of anonymised MPA data is provided by Cuebiq, a location intelligence company that collects de-identified data from opted-in users, 1 takes advantage of both high resolution and temporal variation in the data in describing individual activity patterns. The data are derived from applications such as transportation, social media, exercise and weather categories, amongst others, within Chicago city boundaries. The characteristics in the data set are: latitude and longitude, location accuracy (in metres), timestamp and an anonymised and encrypted user ID.
A summary of the data cleaning, filtering and aggregation process is provided in Table 1, along with summary statistics. The initial stage of data pre-processing concerns removing data with poor accuracy and data on transit, because of the difficulty of distinguishing vehicular modalities – using this data. driving alone in a car, which is not conducive to exposure (Wong and Shaw, 2011), appears effectively the same as multiple people on a bus. All data that have higher than a 200 m accuracy – roughly the radius of a circle inscribed in an average block group area of 165 km2 in the city – are removed. All points within a 18.29 m (or 60 ft) diameter buffer around the arterial road centrelines are also removed. 2 This diameter is determined by the typical residential street width in Chicago of 20.12 m (or 66 ft) diameter. 3 All points with velocities greater than 24 km per hour or 14.91 miles per hour – faster than a standard walking pace but slower than the average vehicular speeds of 24 miles per hour in the city (Smith, 2018) – are also removed. This corrects for some of the selection bias as more than half the data appear to be transit data by these metrics.
Summary of data preprocessing, home and stay locations and interaction cluster creation.
Notes: aFor pre-processing, percentiles for raw data based on a representative day. For home and stay location, percentiles for location centroids per user across the entire period. For interaction clusters, percentiles for each interaction cluster.
Stays and demographic profiles
Two fundamental assumptions underpin the exposure measures using the MPA data: the first is the assumption of a probabilistic demographic profile for each individual in the data based on their estimated home locations (Wong and Shaw, 2011), which are then joined to census data. The second is the assumption that an individual’s daily activity space can be extrapolated from the data.
The initial step for the study is an estimation of home locations for each individual. This is estimated using the density-based spatial clustering of applications with noise (DBSCAN) algorithm (Ester et al., 1996). DBSCAN allows the discovery of activity spaces based on a density of raw points for a given radius. The process of estimating home locations is summarised in Table 1. An individual’s home location is defined as the centroid of their most frequented clusters, using the DBSCAN algorithm with a maximum 150 m threshold 4 from the hours of 01:00 CT to 05:00 CT, when individuals are generally likely to be at home. The 150 m threshold is created to be inscribed within an average Chicago block group. These are then filtered by Chicago zoning districts to keep only those locations within feasible areas of residence according to zoning regulations (residential, downtown and planned development districts). Block groups that show a high ratio of home location densities versus the census population density – primarily in areas of high activity across all hours of the day such as downtown and both O’Hare and Midway airports – are also removed from this analysis, because of the higher probability these are not home activities. The details of this process are described in Table 1.
This analysis employs the notion of ‘stay’ locations (Jiang et al., 2013, 2017) as the centroid of a cluster of at least three different raw points that are within 10 m and 5 minutes of one another. 5 Stay centroids are kept only for individuals who have home locations in Chicago city boundaries as a portion of the individuals (around 67%) are visitors to the city. These stay centroids (simplified as ‘stays’ from here on) operationalise the concept of activity spaces using MPA data. The higher precision of the MPA data allows for a smaller clustering radius that could realistically represent close physical proximity. From stays, it is then possible to distinguish spatial-temporal overlaps of activity spaces for multiple people – an opportunity for interaction. Like stays themselves, interaction clusters are determined by stays that are within 10 m and 5 minutes of one another for at least two different residents. Figure 1 illustrates an example of this.

Example of how overlapping stays (shaded circles) create an interaction cluster (outlined in grey). Individuals A and B have stays that are within 10 m range from one another and also overlap in time.
Residential location and demographics
The resulting data represent a sample of 87,631 anonymous residents. Though the number of residents in this analysis is smaller than the estimated 300,000 total number of users in the original data set due to the necessity of filtering the data for residents, the resulting Pearson correlation of the MPA-derived log population density to the ACS log population density on the Census block group is 0.83 and is within the same range of other studies using similar scales (Bachir et al., 2017; Deville et al., 2014).
Derived home locations and other census characteristics are used to create an estimated population based on the 2014–2018 American Community Census 5-Year estimates at the block group level, along with probabilistic demographic profiles for each resident in the final data set. The estimated population, using the ordinary least squares (OLS) model below, is employed across all exposure calculations. For block group
This model produces an unbiased prediction of population density with an adjusted R-squared of 0.793,
7
confirming Hypothesis 1. This model is then used to create a predicted population

Distribution of predicted versus actual census demographics.
To mitigate ecological fallacies, each individual is assigned a probability of membership in a particular group based on their home census block demographics in the 2014–2018 ACS 5-Year estimates. For race and ethnicity, I use the percentage categories for non-Hispanic White, non-Hispanic Black, non-Hispanic Asian and Hispanic. For instance, if an individual’s home location is in a block group with 70% non-Hispanic Black, 20% non-Hispanic White, 5% non-Hispanic Asian and 5% Hispanic, their probability of membership in each racial category will be the same.
Exposure measure
Standard measures of exposure to diverse social contact are based on Lieberson’s (1981) formulation. This analysis uses a measure of local interaction potential based on the Simpson index – similar in nature to the standard Lieberson measure – which reflects the probability that two randomly selected people will belong to members of different groups and is agnostic to boundaries of area units and the number of groups in the comparison. The three exposure measures are calculations of the Simpson index in three different contexts. Here the baseline residential exposure for each block group
where
where
where k indexes on clusters within i. All shared activity spaces within a block group
Results
Table 2 summarises the results of comparisons between three exposure measures, with three different radii for interaction clusters to determine sensitivity of the interaction exposure to distance. The median for the residential exposure is 0.33, indicating that there is a 33% probability that any two random residents in the block group are of different races. The observed exposure probability is generally higher than the residential exposure, with a median of 0.49, which confirms Hypothesis 2. This result may be intuitive as the observed exposure represents a wider range of activities spaces, including the residential space implied by the residential exposure measure. Residential exposure is less than 0.13 for 25% of block groups – unsurprising given the city’s residential segregation along racial lines (Massey, 2016; Massey and Tannen, 2015; Sampson, 2011) – while the other measures are closer to 30% at the 25th percentile. The interaction exposure at the 10 m radius is 0.44 and only slightly less than observed exposure. The statistics for larger radii indicate that the exposure levels do not change much with a looser social interaction clustering radius. There is an 81% and 63% correlation between the residential exposure and observed and interaction exposures, respectively, implying that areas of high residential exposure are also areas with high exposure for other kinds of activity spaces. For the majority of block groups, the observed exposure is significantly higher than the residential exposure, as shown in Figure 3. In terms of the distribution of interaction exposures these are all as expectedly lower than the observed exposures, though surprisingly, most interaction exposure levels are higher than the residential exposure. In other words, when people have an opportunity to interact, the degree of exposure to racial diversity generally tends to be higher than what is implied by the residential exposure measures. Thus, Hypothesis 3 cannot be firmed.
Exposure score summary.

Residential exposure (left), difference between observed and residential exposure (middle) difference between interaction and residential exposure at 10 m radius (right).
Figure 3 shows the geographic variation of residential exposure and how the observed and interaction exposure measures deviate from residential. All three exposure measures have generally similar regions of high and low exposure levels (in the middle and right maps in Figure 3 these are grey and light-coloured areas near 0). Regions that are most exposed to diversity are along the city’s major roadways where the city has seen growth and gentrification in the last few decades and where there are significant Hispanic populations. The middle map in Figure 3 shows that observed exposure levels are higher than residential exposure, especially in the city’s South and West sides, which suggests an underestimation of exposure when relying on census data. The rightmost map showing the difference between interaction and residential exposure loosely corroborates this proposition but also highlights the spatial heterogeneity of the interaction measure. Of note is the total number of census blocks covered by the interaction exposure measure, which is a smaller subset of the residential and observed measures. This is due to the fact that not all blocks contained interaction clusters; interaction exposure is conditional on individuals having the opportunity to come into close proximity.
Figure 4 shows the density of social interaction clusters and their average interaction exposure measure by the hour across the days of the week. The density of clusters are typically higher during the day and begin to decline around 16:00 h and also generally declining starting on Thursdays (with August being an exception, where density is higher starting Thursdays). The mean interaction exposures tend to be highest between the hours of 04:00 h and 12:00 on weekdays and gradually decline over the course the afternoon and evening, with exposures generally lower on the weekends. 9

Density distribution of social interaction clusters across the week by hour (top). Average interaction exposure throughout the week by hour, with shading representing the 95% confidence interval (bottom).
Discussion and conclusion
This analysis demonstrates the possibility of creating an unbiased and representative predicted population and also describes possible links between neighbourhood diversity and opportunities for diverse social interaction through comparing residential exposure to actual and more specific contexts of interaction opportunities. The results confirm Hypothesis 2: observed exposure to racial diversity is higher than residential exposure, with West side neighbourhoods such as North Lawndale and Garfield Park and South side neighbourhoods such Washington Park and Englewood showing much higher exposures than the residential data in these ‘hyper-segregated’ (Massey and Denton, 1989) neighbourhoods would indicate. Nevertheless, observed exposure is still fairly correlated with residential measures. Although here I consider the full scope of an individual’s daily activity space, this finding aligns with previous literature comparing workspaces and transit activities exposures to residential exposures.
The findings regarding Hypothesis 3 are more surprising. The interaction exposure generally shows higher values than residential exposure, conditional on opportunities for interactions occurring in space and time. This is likely also due to the wider range of activity spaces accounted for when using MPA data. The findings from interaction exposures also reveal a higher degree of spatial heterogeneity relative to residential exposure, even on racially segregated West and South sides of the city, which confirms previous research (Jones and Pebley, 2014). Of note are the pockets of high interaction exposure in these parts of the city next to ones of low exposure. Ultimately, these results do not necessarily refute previous research. The interaction exposure describes a set of opportunities for interaction and not actual interactions; further, in this paper, these opportunities are defined by space-time proximity. This may engender one type of opportunity for interaction but exclude other types of proximity that are also conducive to interaction, such as social network proximity. The weekly temporal variation of the degree of interaction suggests a possible source of exposure: exposure is generally higher during regular weekday activities, perhaps mostly consistently articulated in the morning, which are likely workplace or school activities. The monthly variation of exposure in Figure 4, especially the reduced interaction clusters in October, also suggests seasonality in these patterns and sensitivity of exposure due to the weather. Lastly, the degree of overlap across all three exposure measures suggests that residential context remains an important factor in determining exposures across a broader range of activities and contexts. This may result from residents extending their activity spaces most frequently to places in and near their home locations.
There are several limitations and future research questions that this analysis introduces: first, there are questions of representativeness of the home and stay locations. The estimation of the home locations assumes the MPA users represent a random sample of residents in the block group, which may not be the case if there is selection bias for MPA users. Furthermore, while the residential population can be estimated using the MPA data, the lack of reliable temporally dynamic population counts means temporal calibration of MPA stay data is difficult to validate. This presents a potential area of future research. Second, as previously mentioned, opportunities for interaction do not necessarily result in actual interaction. The limit in precision in data – 12 m at the median – also constrains more detailed analysis; thus, future research may explore mixed methods approaches to validating actual interactions. Third, future research could study whether exposure differs for different racial and ethnic groups while also considering its intersection with socioeconomic status (Browning et al., 2017b). Lastly, though use of these data for research purposes required compliance with privacy restrictions set by the provider, the increasing availability of MPA data raises broader questions of privacy protection. While the USA lacks legislation similar to the European Union’s General Data Protection Regulation, research on geographic differential privacy to obscure traceability is increasingly prevalent (Andrés et al., 2013; Xiao and Xiong, 2015) and should be considered in future analysis.
From a policy perspective, most affordable housing policies in the USA favour mixed-income neighbourhood diversity and ‘policy-driven gentrification’ approaches, despite the unresolved link between neighbourhood diversity and opportunity through, amongst other influences, the potential for social contact. This article suggests there is more potential exposure to diverse social contact than previous literature finds and opens up new methods to explore the social contact mechanism, which also demonstrates an empirical and scalable approach to elucidating these policy questions. Future investigations that incorporate specific policies, particular groups of interest and an intersectional approach to diversity, such as between race and class, may improve our understanding of the overall effects of social mixing-driven housing and urban development policy.
Footnotes
Acknowledgements
This article was originally presented at the ‘Predicting neighborhood change using big data and machine learning: Implications for theory, methods, and practice’ symposium on January 9, 2020 at the University of California Berkeley. Thank you to the organizers and attendees for their feedback on an earlier draft. I would also like to thank Lance Freeman, Jamie Saxon, Shin Bin Tan, the presenters and attendees of the Association of Collegiate Schools of Planning Annual Conference on November 5, 2020, David Weisberger, the anonymous reviewers, and the editor for their feedback.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
