Abstract
Analyzing social change requires detecting patterns of continuity and difference over time. While time-series clustering offers a valuable approach, existing techniques are often limited by assuming fixed cluster definitions and static assignments of entities to clusters. To address these limitations, we introduce a unified framework of temporal clustering methods that allows for both dynamic cluster definitions and the transition of entities between clusters, generalizing and extending previous work. We also provide new algorithms for this dynamic clustering that optimize global objectives, with optional constraints on the transitions of entities across clusters. This framework expands the methodological toolkit for analyzing social change, and we provide guidelines for its application. We illustrate our approach with three case studies: polarization of social and political attitudes across U.S. states; cross-national cultural change; and the evolution of neighborhood business patterns. We conclude with directions for further research.
Introduction
Understanding social change—the aggregate shift in characteristics of units like neighborhoods, countries, or states over time—is a central task in sociology and related fields (Fosse 2023a, 2023b; Fosse and Winship 2023; Silver et al. 2022a; Silver and Silva 2021, 2023; Silver et al. 2022b). 1 Recent studies have examined such change in various areas, such as political party identification (Fosse and Winship 2023), gender attitudes (Meagher and Shu 2019), and verbal ability (Fosse 2023c), across a range of units, such as societies, organizations, and neighborhoods (Dias and Silver 2021). However, analyzing social change is challenging because it requires identifying patterns of continuity and difference across many units, each measured on multiple features, while the broader social context itself may be shifting (Fosse 2023a).
Clustering methods are well suited for analyzing complex, multidimensional patterns of social change. By grouping units on many features at once, they can reveal important underlying patterns, even when relationships are nonlinear and heterogeneous.
While conventional clustering methods are useful, they typically group entities based on their similarities without considering their temporal sequencing. The analysis of social change, however, requires incorporating a temporal dimension. This extension leads to temporal clustering: grouping complex entities based on their trajectories over time, such as cultural values of countries, voting patterns in U.S. states, or business compositions of neighborhoods. 3 This paper presents a unified framework for temporal clustering that expands the methods for analyzing social change. Our framework organizes existing approaches, while introducing new ones.
Our framework is structured around two key dimensions: (a) cluster definitions and (b) cluster labels. Cluster definitions (the characteristics that define a group) can be either static or dynamic. For example, the definition of a “high” fertility country or a “red” state may evolve. Second, cluster labels (the assignment of entities to groups) can also be static or dynamic. A country might move from a “high” to a “low” fertility cluster, or a U.S. state might move from a “red” cluster to a “blue” cluster. Combining these two dimensions creates a typology that organizes existing temporal clustering methods (e.g., Aghabozorgi et al. 2015; Delmelle 2016; Liao 2005), and identifies new ones, including those where both labels and definitions change simultaneously.
The remainder of this paper proceeds as follows. We first present our unified framework and the four classes of temporal clustering methods it subsumes. We then detail the methodology, from the underlying clustering objectives and optimization algorithms 4 to practical considerations for implementation. Next, we illustrate our framework with three case studies: state-level political polarization in the United States, cross-national cultural change, and the evolution of business establishments in Chicago neighborhoods. We conclude by summarizing key contributions and suggesting directions for further research.
Unified Framework for Temporal Clustering
In this section, we present the proposed unified framework for temporal clustering. The foundation of the proposed framework is built on top of center-based clustering, whose intuition was discussed in the previous section. In this context, a cluster definition refers to the set of characteristics that describes the typical profile of a cluster that is shared among all entities in this cluster (e.g., a cluster with countries that share the “increasing fertility over years” trend), which can be directly inferred from the center of each cluster. The cluster label denotes the assignment of each entity to a specific cluster at a given time (e.g., a state being labeled as “red” or “blue”).
However, existing clustering approaches are often limited by two critical assumptions. First, many methods assume fixed cluster definitions, such that the characteristics defining each group remain constant over time. This assumption is problematic because the substantive meaning of categories can shift significantly: for example, the threshold for “high fertility” declined substantially during the late twentieth century, and the characteristics distinguishing “red” and “blue” U.S. states have evolved alongside changes in political polarization. This can result either in incorrect cluster definitions, where evolving temporal trends are obscured by a static definition, or in misaligned cluster assignments, such that a high-fertility country is incorrectly grouped into a low-fertility cluster as its fertility rate declines over time, making it more similar to the static definition of low fertility.
Second, conventional methods frequently impose static cluster membership, meaning that once an entity is assigned a cluster label, it cannot transition even if its characteristics change. This assumption fails to capture important dynamics, such as states like Texas and Oklahoma shifting their political alignments across decades, or countries moving toward greater secularization in global value surveys. As we later demonstrate in case studies, applying methods that assume fixed definitions and cluster labels would obscure meaningful shifts or exaggerate stability where none exists. Without accounting for the possibility that both the definitions of clusters and the membership of entities may change, traditional clustering methods risk producing misleading or incomplete representations of social change.
To address these limitations, we introduce a unified framework for temporal clustering that systematically incorporates dynamic cluster definitions and dynamic entity memberships. The framework expands the methodological toolkit available for analyzing how entities evolve across time and space. We present this unified framework based on two key dimensions. This framework is shown in Table 1. The rows of the table represent the dynamism of cluster definitions, while the columns of the table indicate the dynamism of the cluster labels or the composition of the clusters. For the cluster definitions, the key question is: Is the definition of the cluster allowed to change over time? For example, with respect to political polarization, does what it means to be a “blue” versus “red” state change over time? This might occur, for example, if being a “red” state initially depends on mass social attitudes but then becomes more of a function of party identification. For the cluster labels, the key question is: Are the cluster labels of the entities allowed to change over time? For example, is a state such as Ohio allowed to move from a “blue” state cluster to a “red” state cluster, or must Ohio always be classified as either “red” or “blue”? This framework yields four classes of temporal clustering methods, as illustrated in Table 1.
Temporal clustering methods: Definitions and use cases The table presents four main temporal clustering methods categorized by two dimensions: the dynamism of cluster definitions (rows) and the dynamism of cluster labels (columns). Examples are based on different assumptions of the definition and composition of “red” and “blue” states in the U.S.
When considering different models of temporal clustering, it is important to specify the degree of dynamism of the cluster labels (i.e., how often and how many entities are allowed to transition from their original cluster assignment.) We can represent the number of allowable cluster label changes using a hyperparameter
Cluster Definitions and Labels are Both Static
The first approach of temporal clustering assumes that the definition and composition of clusters are static. This approach, which we call Static Clustering (SC), is shown in the upper left quadrant of Table 1. With SC, the hyperparameter
Dynamic Cluster Definitions but Static Cluster Labels
The other possibility is to assume that the cluster definitions change, but that the cluster labels are fixed, so that the composition of the clusters remains the same. This approach is shown in the lower left quadrant of Table 1 and is referred to as Time-Series Clustering or TSC for short (Aghabozorgi et al. 2015). Although uncommon in sociology and demography, TSC is a powerful method for identifying patterns and structures in evolving sequential data and has been applied in many fields, including finance, medicine, biology, and geography. TSC extends traditional center-based clustering by incorporating similarity measures that take into account entire sequences of time-ordered data. With TSC, for example, the definition of a “red” and “blue” state evolves to reflect the fact that a “red” state may increasingly be a function of mass social attitudes rather than political partisanship. However, with TSC the hyperparameter
Static Cluster Definitions but Dynamic Cluster Labels
Another possibility is to assume constant cluster definitions, but allow dynamic cluster labels so that entities can move from one cluster to another. This is shown in the upper right cell of Table 1. We call this approach Sequence Label Analysis (SLA). In the context of political polarization, SLA assumes that the definition of a “red” and “blue” state does not evolve, so that any change is limited to the composition of the clusters rather than their definitions.
SLA can be understood as an adaptation of a two-step technique recently developed in studies of neighborhood evolution based on sequence analysis (Delmelle 2015, 2016, 2017; Dias and Silver 2021; Kang et al. 2020; Patias et al. 2020; Silver and Silva 2021). In the first stage, the temporal relationship is ignored by clustering each time point in each entity independently, rather than clustering the entire temporal sequence as in TSC. Each entity has an aggregated sequence constructed from the cluster labels of its time points, preserving the original temporal order. Thus, all entities are allowed to have different cluster labels at different time points. In the second step, these aggregated sequences are analyzed, for example, using similarity measures such as the optimal matching (OM) distance (Delmelle 2016; Kang et al. 2020) or Markov matrices (Silver and Silva 2021). The entire temporal data of each entity is then grouped based on the similarity between cluster label sequences or transition probabilities between clusters. The first stage of this two-stage approach is a form of SLA, as defined in Table 1.
Specifically, SLA in existing studies essentially assumes a value of
Cluster Definitions and Labels Both Dynamic
The last logical possibility, shown in the lower right cell of Table 1, is that both the cluster definitions and the labels are dynamic. In other words, not only do the cluster definitions change, but the compositions of the clusters change as well, reflecting the fact that entities can move from one cluster to another. Temporal clustering in this category is called Dynamic Clustering (DC). In the context of political polarization, DC assumes not only that the definition of a “red” and “blue” state evolves, but so do the compositions of the clusters.
As with SLA, there are bounded and unbounded variants of DC. The unbounded version of DC corresponds to a value of
For this reason, in practice the bounded version of DC is generally more desirable. This version of DC corresponds to a value of the hyperparameter
Methodology of Temporal Clustering
In this section, we discuss the methodology of temporal clustering, addressing the technical challenges and presenting our solutions. We formalize the criteria for optimal temporal clusters as mathematical objectives that account for intra-cluster similarity and accommodate various clustering methods, extending the objective of traditional center-based clustering to incorporate temporal relationships within the dataset. We also present the objectives for each of the four main temporal clustering methods within our unified framework in Table 2. To optimize these objectives, we briefly outline two approaches: MILP and an iterative approach inspired by the Majorize-Minimization (MM) algorithm. 7 We conclude this section by discussing practical considerations crucial for the efficiency and interpretability of clustering results, including similarity measures, feature selection, determining the number of clusters, and setting the maximum number of cluster label changes allowed.
Mathematical objectives of temporal clustering methods the temporal extension of center-based clustering’s objective aims to minimize the sum of distances between each entity and its closest centers, considering the status of cluster definitions (centers)
Clustering Objectives
Temporal clustering presents several technical challenges. First and foremost, it is crucial to rigorously define the criteria for optimal temporal clusters. We have formalized these criteria as mathematical objectives that account for intra-cluster similarity while accommodating the various clustering methods shown in Table 1. These objectives extend the traditional center-based clustering approach, which minimizes the sum of distances between each entity and its nearest center, to incorporate temporal relationships. Our objectives involve changes in both cluster labels and definitions, thereby accommodating the full range of clustering methods in our framework, from static clustering to fully dynamic approaches. Due to the technical nature of the following section, applied researchers may wish to skip to the subsequent section (“Practical Considerations”).
To clarify the discussion that follows, we introduce some notation. Let
Using the notation above, Table 2 presents the mathematical objectives for each class of temporal clustering methods in our framework (cf. Table 1). As previously discussed, the primary objective of all four methods is to minimize the sum of squared distances between each entity,
The primary distinction lies in whether cluster centers
It is important to note that the clustering objectives remain unchanged for the methods in the rightmost column of Table 2, specifically SLA and DC, regardless of whether the maximum allowable label changes
Another crucial consideration is selecting an appropriate optimization method for solving the objectives in Table 2. Traditional clustering algorithms, such as Lloyd’s algorithm for K-means clustering (Lloyd 1982), use an iterative approach inspired by the MM algorithm (Kenneth et al. 2000). We extend this method to optimize the objectives of methods for temporal clustering. The MM algorithm approximates the optimal solution by alternately updating cluster centers and reassigning entities to the updated clusters, gradually improving the clustering result until convergence. We also incorporate a post hoc analysis during each iteration to identify the top
Practical Considerations
Besides selecting a suitable temporal clustering method and optimization approach for a given use case and dataset, researchers must also weigh several additional factors to maximize the effectiveness and interpretability of their results. These practical considerations can significantly influence the analysis. In this section, we outline important aspects to evaluate when implementing temporal clustering, including the choice of distance measures, feature selection and normalization, determining the optimal number of clusters, and setting restrictions on the number of cluster label changes.
Choosing a distance metric
Distance measures are widely used to assess similarity between entities and cluster centers in clustering and social sequence analysis (Aisenbrey and Anette 2010). They are effective for evaluating both feature values and temporal relationships, and are efficient for large datasets (Shieh and Keogh 2008). However, choosing the right distance measure is critical in temporal clustering, because an entity may appear close to a cluster center according to one measure, but far away according to another. Appendix B in the online supplement provides an example illustrating how different distance functions yield varying results for the same set of entities. Thus, the choice of distance measure can greatly influence the results of one’s analysis, not only for conventional clustering techniques and sequence analysis but also for temporal clustering.
Manhattan distance (also referred to as the
In general, the choice of distance measures depends highly on the temporal dataset and specific use case. We recommend starting with metric distance functions like Manhattan Distance or Euclidean Distance when applying temporal clustering. These measures are simple, effective, and scalable for large temporal datasets. In some applications, specialized similarity measures may be necessary, such as dynamic time warping (DTW) for aligning sequences with different durations or phase shifts, especially when traditional distance metrics fail to capture relevant similarities (Müller 2007). Alternatively, cosine distance, which captures the cosine of the angle between two vectors, may be useful for handling temporal sequences of different lengths after a dataset has been normalized. Regardless of the chosen measure, the selection of an appropriate distance metric is critical as it forms the basis for assessing closeness between data points, which the clustering algorithms then use to group similar entities together. For this reason, researchers should generally consider different distance measures and their impact on how the data is clustered to determine the robustness of their results.
Selecting and normalizing the features
Another important consideration is the number of features to consider when conducting a cluster analysis. An excessive number of features can increase the runtime of the optimization process as well as complicate the interpretation of temporal patterns. In general, we recommend considering features that are most relevant to the research question and that are expected to show meaningful variation across entities and time. As with any analysis, domain knowledge can guide the selection of key features.
Another related consideration is normalization. As discussed above, the distance function is sensitive to how features are scaled. Differences in feature scales can lead to distorted results, such that features with larger scales appear more substantial than those with smaller scales. Therefore, techniques such as min–max normalization and z-score normalization are essential. Normalizing each characteristic and time point separately can also eliminate common trends across entities, allowing for more interpretable comparisons between entities for certain features (e.g., mitigating the effects of inflation on median income). Because of the relative ease of implementation and interpretation, we generally recommend applying z-score normalization to each feature separately at each time point.
Determining the number of clusters
In traditional clustering, the results are sensitive to the number of clusters specified

Visualization of methods for determining the optimal number of clusters
A drawback of both methods is the need to run the clustering algorithm multiple times with different values of
Setting the maximum number of cluster label changes
The maximum number of label changes across clusters among all entities, given by the hyperparameter
Taken together, the methods presented in this section provide a comprehensive framework for temporal clustering, encompassing both mathematical foundations and practical guidelines. Researchers can effectively apply these methods by formalizing clustering objectives, exploring optimization techniques through MILP and iterative approaches, and addressing key considerations in data preparation and parameter selection to analyze complex patterns of social change. The flexibility of this framework allows for adaptation to various types of entities and temporal datasets, while also ensuring robust and interpretable results. As demonstrated in the subsequent case studies, careful application of our framework enables deeper insights into diverse sociological contexts, from political polarization to cultural value shifts to patterns of urban development.
Temporal Clustering: Sociological Examples
To evaluate and demonstrate our proposed methods, we introduce three case studies on social change: state-level social and political attitudes in the United States, cross-national cultural change in the World Values Survey (WVS), and neighborhood business development in Chicago. These cases were chosen not only because they are important topics in sociology and related fields, but also because they reveal the dynamic nature of social change, showing how entities can move between clusters as well as how new clusters emerge, reflecting evolving social, political, and cultural contexts.
The temporal clustering methods used were selected based on factors such as the number of time points and features within the dataset as well as the natural characteristics of social change of the entities (e.g., whether the definitions of clusters should evolve or remain static). In addition, we present several approaches to illustrate some of the practical considerations outlined in the previous section, in particular determining the number of clusters
State-Level Temporal Polarization
The issue of political polarization in the United States has received considerable attention across a wide range of studies in various social science fields (Brown and Enos 2021; Iyengar et al. 2019). Researchers have observed that the U.S. population has increasingly shifted towards a strong affiliation with one of the two major political parties, the Democrats and the Republicans, over time. This “polarization” is evident in voting patterns, the proportion of seats in the Senate and House between the two parties in different states (Poole and Rosenthal 1984), and social attitudes as revealed in survey data. McCarty et al. (2016) also highlighted the growing disagreement on policy issues between political elites in the two parties. In our first case study, we employed temporal clustering on various cultural and political features derived from mass survey responses to evaluate prevailing accounts of population-level political polarization between states.
Existing temporal clustering methods, such as TSC and SLA, fall short as they do not account for changes in both cluster definitions and labels over time. By contrast, DC allows U.S. states to transit between Democrat-affiliated and Republican-affiliated clusters and modifies the definitions of two clusters of states. Accordingly, DC offers an arguably more realistic description of temporal polarization than existing methods, as it does not assume that a state is always Democrat-affiliated (as with TSC) or that a Democrat-affiliated cluster has an identical proportion of votes for Democrats across years (as with SLA). Furthermore, by adjusting the maximum allowable changes in labels using the hyperparameter
Dataset
Our analysis examines changes in policy views, attitudes, and partisanship across different states over the past seventy years (1946–2014). The temporal data was collected by the Correlates of State Policy Project (CSPP) (Grossmann et al. 2021), a comprehensive and publicly accessible database covering over 3,000 state policies and politics-related variables from various sources spanning seven decades. We selected five features from the CSPP for analysis. One key feature is the ratio between the proportion of Democratic and Republican identifiers in each state, indicating residents’ political affiliations. A ratio greater than one indicates a larger proportion of Democratic identifiers than Republican identifiers, while a ratio less than one suggests the opposite. 9
We also included four CSPP measures of estimated policy and mass resident liberalism in each state. These features consist of measures of social and economic policy liberalism at the state level (“policysociallib_est” and “policyeconlib_est” in the CSPP codebook) as well as estimates of mass social and economic liberalism among state residents (“masssociallib_est” and “masseconlib_est”). The estimates of mass liberalism are constructed using public opinion surveys weighted by the proportion of identifiers in different parties (Caughey and Warshaw 2018).
Pre-proprocessing and tuning the hyperparameters
As noted above, the temporal dataset spans yearly data for the five features from 1946 to 2014 across all 50 states. For Alaska and Hawaii, which became states in 1959, data from 1949 to 1958 were imputed using the values from their first available year (1959). We observed common trends among all states in certain features, such as a general increasing trend in the mass social liberalism over the duration of the dataset. We applied z-score normalization separately per feature per year to eliminate these common trends and facilitate easier identification of notable temporal trends within entities. While we chose to apply z-score normalization in this analysis to highlight relative temporal trends within entities, we note that researchers may choose to normalize or not normalize variables based on their specific substantive goals. Importantly, the proposed framework operates seamlessly with both normalized and unnormalized inputs, without requiring any structural modifications to the clustering objectives or optimization procedures. This flexibility allows researchers to tailor the analysis to emphasize either relative patterns of change (using normalized variables) or absolute levels and shifts (using raw variables).
To determine the appropriate number of clusters
Using a similar two-stage approach, we determined the maximum allowance for changes in cluster labels as
For this particular application, we first we used the elbow method to evaluate the reduction in the DC objective as the number of allowable changes increased, similar to determining the number of clusters
Temporal clustering results
The DC method (see Figure 2) reveals a number of hidden temporal clusters and trends in the state-level temporal data. 11 Since all features are z-score normalized, the cluster definition for the “Ratio of Proportion of Identifiers between two Parties” feature loses its original interpretability (i.e., a value below one does not necessarily indicate a higher proportion of Republican identifiers compared to Democratic identifiers). To address this, we z-score normalized the value of one alongside other data points during pre-processing. The normalized value of one is depicted as the black dash-dot line in Figure 2. This line represents an equal proportion of identifiers between the two parties for each of the time points. Trends diverging significantly above this line suggest a stronger association with the Democratic Party, whereas values below indicate a stronger association with the Republican Party.

Changing definitions of “red” and “blue”: line plots of cluster definitions across U.S. States. The definition of each cluster is illustrated as dashed lines in the line plots below, with a distinct color for each cluster. Five features are used in the analysis. The black dash-dot line represents the normalized value of one, such that values above this line indicate a higher proportion of identifiers in Democratic parties and values below it indicate the opposite. For the “Relative Proportion Democrat versus Republican” feature, the temporal trends of each cluster exhibit increasing polarization over time, as illustrated by the divergence of the dashed lines from the black dash-dot line. Analyses based on DC with
Figure 2 shows increasing polarization between states, as reflected in the changing cluster definitions. This is indicated by the widening disparity in the proportions of identifiers affiliated with the two political parties over time. Initially, most clusters exhibit a nearly equal ratio between the identifiers of the two parties, suggesting a balanced distribution between parties across states. However, this ratio diverges significantly as most clusters shift towards polarized extremes. For example, the definition of cluster 1 (represented by the deep blue dashed line) increasingly aligns with the Democratic party, whereas the definitions of clusters 5 and 6 (depicted by the orange and deep red dashed lines) show a decreasing ratio, indicating a stronger association with the Republican party. These temporal trends underscore patterns of polarization in resident political orientation observed across multiple clusters of states over the years. A similar pattern is observed in the various measures of liberalism. Before 1960, there was a lack of distinction in liberalism across most clusters, such as the volatile and overlapping pattern observed in mass resident economic liberalism (see Figure 2) within cluster definitions. Post-1960, however, a significant divergence emerged in the cluster definitions, such that state policies and resident opinions on economic and social matters began to exhibit more distinct patterns across states.
Based on the geographic distribution of each cluster (illustrated in Figure 3) and a detailed analysis of the entities’ temporal data within each cluster (included in the Appendix C in the online supplement), the following sections offer an in-depth description of the six types of states. These descriptions of how clusters have changed with respect to political affiliation, liberalism in economic and social policies, and residents’ opinions on economic and social issues.

Geographical distribution of sociopolitical clusters in 2014. This choropleth map depicts the distribution of U.S. states across six clusters in the final year of observation, 2014, based on five features related to political affiliation, mass liberalism, and policy liberalism. The legend provides a brief description of each cluster’s temporal pattern. Notably, Texas was initially part of Cluster 6 but shifted to Cluster 4 in 1984. Alaska and Oklahoma also experienced shifts in their cluster labels: Alaska moved from Cluster 2 to Cluster 4 in 2000, and Oklahoma transitioned from Cluster 4 to Cluster 6 in 1983. The geographical distribution of clusters before any changes in cluster labels, as determined by DC’s results, is shown in Figure 4. Analyses based on DC with
Cluster 1: Transition from balanced to strong democratic affiliation with high and stable liberalism
The states in Cluster 1 initially had a similar proportion of identifiers between the two parties, indicated by a ratio near one before 1960. Subsequently, the ratio shifted significantly toward the Democratic party, with a pronounced increasing trend, making this cluster the highest average ratio among all states after 2000. This trend suggests a continuously increasing and strong affiliation toward the Democratic party among the residents of these states.The four types of liberalism measures in these states consistently ranked among the highest of all states throughout the studied period. Resident economic liberalism increased from the 75th percentile to the highest rank between 1946 and 1960. This temporal pattern reveals a high degree of liberalism in both states’ policies and residents’ opinions. 12
Cluster 2: Minor differences between two parties with moderately increased democratic affiliation and moderate to high liberalism
The time series of the five features for the states within Cluster 2 maintained a small difference in the proportion of residents identifying with the Democratic and Republican parties, with a slightly higher proportion of Democratic identifiers. This stable ratio indicates a persistent partisan lean favoring the Democratic party over the Republican party.
The temporal trends of estimated liberalism for both policies and residents exhibit distinct patterns in two epochs, before and after 1970. Before 1970 there was higher volatility in all four measures of liberalism, particularly in mass resident economic liberalism, which fluctuated between high and low levels without clear trends shared among states. 13 After 1970, volatility decreased significantly within all four types of liberalism, indicating more stable state-level policies and residents’ social and economic attitudes. Cluster 2’s temporal pattern stabilized around the 75th percentile among all states, revealing a moderate and high degree of liberalism in states within this cluster after 1970.
Cluster 3: Increasing democratic party affiliation until 1990, followed by decline, with high and stable median liberalism
Cluster 3 includes only two states, West Virginia and Kentucky, but demonstrates distinct temporal trends in the features, making it a unique cluster. The ratio between the proportions of identifiers in the two political parties exhibits three stages over the years. First, before 1960 the ratio maintained a consistent value, indicating a stronger affiliation with the Democratic Party, though the difference was relatively insignificant. Next, from 1960 to 1990, there was a significant increase in the ratio, indicating a larger proportion of residents affiliating with the Democratic Party. Finally, from 1990 to 2012, this high ratio toward the Democratic Party declined, shifting back to a range similar to that before 1960. Note, however, that the Democratic Party has maintained a substantial proportion of residents in these two states.
The measures of liberalism also reflect similar temporal stages as the party affiliation ratio. Before 1960, the estimated liberalism of economic-related policies and resident opinion on economic matters exhibited upward trends from the 25th percentile to the median among all states, while the two measures of social liberalism remained within the 25th percentile. Then, from 1960 to 2000, social policy first showed an increasing trend into the median and maintained within the median rank alongside economic policy liberalism. Mass resident liberalism exhibited two polarized trends: while resident economic liberalism continued to rise into the 75th percentile from the median, resident social liberalism remained within the lowest range of all states. After 2000, there was a significant declining trend in resident economic liberalism back toward the median to the 25th percentile, revealing low liberalism in mass resident opinions on economic and social matters, while the two types of policy liberalism maintained the median rank.
Cluster 4: Transition from moderately democratic to balanced and moderately republican affiliation with medium but highly volatile liberalism
In terms of partisanship, Cluster 4 indicated a tendency toward Democratic affiliation before 1960, as evidenced by the ratio of identifiers greater than one. After 1960, this ratio decreased, suggesting an increasing proportion of Republican identifiers, and after 1990 both parties had a similar proportion of identifiers in these states, with some states having a higher proportion of Republican identifiers. This trend continued through 2014, the last year in our data.
Measures of liberalism in some states within this cluster show high volatility. For example, Colorado experienced a significant drop from a high ranking in economic policy liberalism to the 25th percentile. Nevertheless, the cluster definition, which represents the average of all states in this cluster, shows relatively stable trends. This stability can be attributed to two factors. First, most states with high volatility shifted within a small range of values. Second, states with high volatility in trends make up a small proportion of this cluster. The cluster definition shows that social liberalism in state policies and residents’ opinions has remained within the median of all states. By contrast, economic liberalism in these states is in a lower range, between the median and the 25th percentile.
Cluster 5: Transition from balance toward strong republican affiliation with low and declining liberalism estimates in both residents and policies
The states in Cluster 5 show an increasing tendency to become more Republican over time. This shift is evident in the two-step transition of the ratio of party identifiers. Prior to 1975, both parties had comparable proportions of identifiers among residents, resulting in a ratio of nearly one to one. After 1975, however, this ratio declined significantly, falling below one, indicating that the proportion of Republican identifiers exceeded that of Democratic identifiers in these states. Along with this shift in partisanship, measures of economic liberalism among both residents and state policies show a significant downward trend. These measures of economic liberalism initially ranged between the median and 75th percentile, but eventually fell to around the 25th percentile or lower, suggesting that economic policies and opinions in these states are becoming less liberal. A similar, though less pronounced, downward trend is observed for both policies and attitudes on social issues.
Cluster 6: Decline from strong democratic affiliation to similar proportion (Significant increase in republican) with low liberalism estimates in state policies
States in Cluster 6 had the highest ratio of two-party identifiers before 1960, with some states showing increasing Democratic Party affiliation, indicating a strong Democratic alignment among residents. However, this pattern shifted after 1960, with most states in the cluster experiencing a declining trend in this ratio. The timing of this decline varied; for example, Alabama maintained its increasing trend toward the Democratic Party until 1976, while Louisiana’s decline began in 1964. This temporal trend reveals a substantial increase in the proportion of Republican identifiers and a corresponding decrease in Democratic identifiers, transforming these states from strongly Democratic to more Republican-leaning.
Measures of liberalism in states’ social policies, economic policies, and residents’ opinions on social issues consistently remained between the lowest rank and the 25th percentile, suggesting relatively low liberalism in state policies on economic and social issues, along with a declining trend in mass social liberalism. By contrast, residents’ opinions on economic issues underwent a two-stage evolution. Before 1960, it showed an increasing trend from the 25th percentile to between the 75th and the highest ranks, indicating a relatively high degree of economic liberalism among residents. After 1960, it declined and stabilized around the median, showing a decline in economic liberalism but still at a higher level than the other measures of liberalism in this cluster.
Before turning to our analysis of notable movements across clusters, it is important to clarify the relationship between continuous measurement and categorical clustering in our approach. Although we occasionally invoke familiar political terms such as “blue” and “red” states to aid interpretation, these labels were not produced by our clustering methods. Rather than imposing binary thresholds (e.g., strict Democratic or Republican cutoffs), our framework identifies six distinct clusters based on shared trajectories within a continuous, multidimensional feature space. The underlying cluster definitions remain fully embedded in continuous variables over time in Figure 2. The number of clusters (
Analysis of notable movements of states across clusters
Using values of

Geographical distribution of sociopolitical clusters in 1982. This choropleth map illustrates the distribution of states across six clusters in 1982 using the same five features. As mentioned in Figure 3, Texas, Alaska, and Oklahoma were in a different cluster in 1982 than their final cluster label, indicating notable changes in their cluster labels. Analyses based on DC with

Line plots of features for the U.S. states with notable movements across clusters. The time series features of the three states (illustrated with a solid line with text annotation on the side) that exhibited changes in their cluster labels (illustrated as changes in the color of the line segment within a single state). Dashed lines indicate the definitions of the clusters. Analyses based on DC with
In 1983, Oklahoma moved from Cluster 4 to Cluster 6. Oklahoma’s time series in the relative proportion of Democrats (vs. Republicans) shows that prior to the transition to Cluster 4 in 1983, it had a stable trend with a higher proportion of Democratic identifiers, which closely matches the definition of Cluster 4 (pink dashed line in Figure 5). However, the slope of its declining trend differs from the Cluster 4 definition and the other states in Cluster 4. This slope indicates a substantial shift in Oklahoma, more in line with the temporal pattern of Cluster 6 (red dashed line). If Oklahoma were statically assigned to Cluster 6 such that time were ignored, its trend before 1983 would show little similarity to the definition of Cluster 6, as Oklahoma has a significantly smaller share of Democrats compared to the other states in Cluster 6.
Further analysis revealed similar characteristics in other features contributing to this change. For example, Cluster 4 has maintained median ranges in estimated liberalism and state policies over the years (pink dashed line in the social and economic policy measures in Figure 5). By contrast, the states in Cluster 6 (red dashed line) show a continuous decline, with lower scores on these liberalism measures than in Cluster 4. Prior to 1983, Oklahoma showed stable liberalism with median values that fit only with Cluster 4. Conversely, Oklahoma’s declining trend after 1983 fits only with Cluster 6, indicating that Oklahoma embodies two different temporal patterns, marking it as a state with a notable shift in its cluster label. 14
Similar scenarios are evident in two other states that experienced changes in their cluster labels. For example, Alaska moved from Cluster 2 to Cluster 4 in 2000 due to a decline in the liberalism of social and economic policies as well as in residents’ opinions on social issues. This decline distinguishes it from the other states in Cluster 4, which maintain measures of liberalism at the 75th percentile throughout the years studied. In addition, Alaska lacks the increasing trend in the proportion of Democratic versus Republican identifiers that is evident in the definition of Cluster 2 (blue dashed line in Figure 5). These unique patterns observed in Alaska led to a change in its cluster label in 2000 to Cluster 4, which has a more similar temporal pattern to Alaska after 2000.
Texas shifted from Cluster 6 to Cluster 4 because of an increase in the liberalism of social policies and residents’ opinions on social issues between 1970 and 1980. The trend thereafter sets Texas apart from the low and declining pattern of Cluster 6 and aligns it more with the stable trend within the median range in Cluster 4. Texas also had an earlier and faster decline in the proportion of identifiers from the two major political parties around 1970 than other states in Cluster 6. After 2000, unlike the temporal pattern observed in the Cluster 6 definition, which shows a continuous declining trend through the end of the years studied, Texas maintained a stable and balanced proportion between the two parties. Consequently, all of the observed differences in Texas support an earlier transition from a state with a high Democratic affiliation to one with a balance between the two political parties compared to other states with similar temporal patterns. Furthermore, Texas exhibits a higher degree of social liberalism, both in its policies and among its residents. These unique temporal patterns are reflected in a change in its cluster label in 1984, as shown in Figure 5.
This case study underscores the power of DC to reveal complex patterns of social change. In describing the dynamic nature of polarization in the United States, DC accounts for both the fluidity of state affiliations as well as the evolving meaning of “red” and “blue” clusters. Our analysis has highlighted six distinct clusters as well as three states, Oklahoma, Texas, and Alaska, that deviate from the broader shifts in cluster definitions, moving from one cluster to another.
Global Cultural Change: World Values Survey
For our second case study, we investigate cross-national changes in cultural values using the WVS (Haerpfer et al. 2022). First, we provide empirical definitions of four types of countries based on ten indicators across three time points aggregated from seven waves of the WVS conducted from 1981 to 2022. We then analyze how countries transition between these types over time to examine cultural changes.
Overview
The WVS is an ongoing research program that aims to measure the cultural values and beliefs of citizens in different countries using standardized questionnaires. A notable research product of the WVS is the “Inglehart–Welzel Cultural Map” (Inglehart 2006), which categorizes countries in each survey wave into distinct cultural clusters. The core methodology involves subsetting to each wave of the WVS, applying factor analysis to select ten indicators, and clustering countries based on their positions in the resulting two-dimensional space. The dimensions represent “Traditional” versus “Secular-Rational’ values and “Survival” versus “Self-expression” values (Inglehart 2006). Next, the countries are grouped manually into nine cultural zones (cf. Huntington 2020) taking into account factors such as language, geography, and religion. Patterns of social change are then identified by visually examining shifts in this two-dimensional space across survey waves.
Despite the immense contributions of the work by Inglehart and colleagues to ongoing debates about the nature and extent of cultural change around the world, their main clustering approach has several limitations. First, reducing the original ten features to just two dimensions through factor analysis can lead to a considerable loss of information. Second, the manual clustering approach lacks a clear mathematical objective, rendering the results subjective and challenging to reproduce. Additionally, since the analysis is performed independently for each wave of survey responses with different subsets of countries, there is a risk of inconsistency in the definition of reduced dimensional space across years. Finally, the cultural zone definitions are based on external, ad hoc decisions that can significantly affect the conclusions drawn from the analysis. As we show below, our proposed temporal clustering framework can address these issues and provide a more rigorous, data-driven analysis of global cultural change using the WVS data.
Dataset
We aggregated the seven WVS waves from 1981 to 2022 into three time points to maximize the number of countries with complete data for the ten selected indicators (see Table 3). The scores of the features were calculated as the mean of the survey responses per country, taking into account sampling weights and eliminating missing responses. Z-Score normalization was applied to each feature to ensure all features are under the same scale. The resulting dataset covers 30 countries across three time points, with ten features capturing various aspects of cultural change.
World values survey indicators this table lists the names, questionnaire codes, descriptions, and possible responses for the ten survey indicators used in this case study, as obtained from the “WVS time series 1981-2022 variables report, version 5.0.” note that the original numerical scores are scaled differently for each question. For some indicators, such as the happiness variable (A008), we have adjusted the numerical scores to facilitate a more straightforward interpretation of the results.
Methods and hyperparameters
Recall that SLA assumes that the cluster definitions remain constant over time, making it most suitable for datasets with few time points.
15
Moreover, when dealing with high-dimensional datasets containing many features, SLA can provide easily interpretable results. Moreover, as we show below, by using the bounded version of SLA, where we set the maximum number of allowed cluster label changes (
Detailed cluster definitions
To begin our analysis, we applied the unbounded version of SLA, where

Cluster definitions in the world values surveys. The radar plots depict the cluster definitions as the scores on the ten survey indicators used. Analyses based on SLA with
Note that SLA directly incorporates the average survey responses into the cluster definitions, providing specific scores for each indicator. This contrasts with the manual clustering and reduction to two principal components used in the original “Inglehart–Welzel Cultural Map.” However, to compare our results to those from Inglehart and colleagues (2006), it is informative to visualize the distributions of the countries within each cluster across the first two principal components derived from the ten survey items.
In Figure 7, we show the position of each country at each of the three aggregated time points, along with the centers of the four clusters. Cluster labels for each time point are represented by different colors, with the center of each cluster indicated by crosses. The colored regions are formed by connecting four archetypal points (Stone and Cutler 1996) of entities within each cluster. Similar to Inglehart (2006), the first two principal components in Figure 7 reveal two key dimensions of cross-cultural variation: (a) traditional versus secular-rational values and (b) survival versus self-expression values. For instance, a lower score on the first principal component (y-axis) indicates that a country leans more towards traditional values, while a higher score indicates a stronger alignment with secular-rational values. The centers of the four clusters, represented by colored crosses, are positioned at the extremes of these dimensions, defining four distinct value types.

Cultural maps of 30 countries across three time points. To compare our results with those from Inglehart and colleagues (2006), the distribution of each country at each aggregated time point is plotted using the first two principal components derived from ten survey indicators. Cluster labels for each time point are represented by different colors, with the center of each cluster indicated by crosses. The colored regions are formed by connecting four archetypal points of entities within each cluster. Analyses based on SLA with
Figure 7 shows, with remarkable detail, how countries have evolved with respect to Inglehart’s fundamental value orientations. To better understand this figure, recall that SLA assumes the cluster definitions remain the same over time (see Table 1). Accordingly, in Figure 7, the position of cluster centers (represented by colored crosses) are fixed across all three time points. However, in contrast to the fixed cluster centers, SLA allows for dynamic cluster labels, meaning that countries can shift between clusters as their positions on the plot change, moving closer to the center of a different cluster at different time points. These transitions are illustrated by the fact that a single country’s label can appear in different colors across the three plots in Figure 7. For instance, in 1995 and 2012, the United States was grouped into Cluster 2 (Traditional and Self-Expression Values), as it was located closest to the center of this cluster (marked by a red cross). However, as the United States’ position shifted toward secular-rational values, it became closer to the center of Cluster 1 (Secular-Rational and Self-expression Values), thereby resulting in a change in its cluster label.
A similar situation occurred with Russia and Ukraine. In 1995, these countries were part of Cluster 3 (Secular-Rational and Survival Values). However, they moved from secular-rational to traditional values in 2012, causing a shift in their labels to Cluster 4 (Traditional and Survival Values). Then, in 2020, Russia and Ukraine shifted back to Cluster 3 (Secular-Rational and Survival Values). It is important to note that, compared to the United States, the positional shifts of Russia and Ukraine were less notable in Figure 7. When entities can move between clusters without restriction (i.e.,
Identifying notable movements between clusters
The results so far are based on the unbounded version of SLA, where
Pakistan: From survival to self-expression
Pakistan experienced a notable shift from Cluster 4 (Traditional and Survival Values) to Cluster 2 (Traditional and Self-expression Values), reflecting a fundamental shift change from survival to self-expression values. As shown in Figure 8, Pakistan shifted towards higher levels of happiness, as well as a shift from materialist to post-materialist values, both of which are consistent with a transition from Cluster 4 to Cluster 2.

Pakistan: temporal changes on selected indicators. These line plots depict the mean survey responses for happiness and post-materialist values in Pakistan across three aggregated time points. The color-coded line segments indicate the cluster labels at each time point, while the dashed lines denote the cluster definitions, which are static. These results are consistent with the Pakistan’s movement from Cluster 4 to Cluster 2, reflecting a shift from survival to self-expression values. Analyses are based on SLA with
Japan: From survival to self-expression
Japan shifted from Cluster 3 (Secular-Rational and Survival Values) to Cluster 1 (Secular-Rational and Self-Expression Values). Japan is notable in that it exhibited a substantial increase in the belief that homosexuality is justified (see Figure 9), setting it apart from other East Asian countries. Alongside this trend was a shift towards higher levels of happiness, such that country became increasingly aligned with the center of Cluster 1 rather than Cluster 3. Note that Japan’s transition from survival to self-expression values is consistent with the trends presented by Inglehart (2006).

Japan: temporal changes on selected indicators. These lines plots depict the mean survey responses for happiness, views on homosexuality, and pride in nationality in Japan across three time points. The color-coded line segments indicate the cluster labels at each time point, while the dashed lines denote the cluster definitions, which are static. These results are consistent with the Japan’s movement from Cluster 3 to Cluster 1, reflecting a shift from survival to self-expression values. Analyses are based on SLA with
United States: From traditional to secular-rational
The United States shifted from Cluster 2 (Traditional and Self-Expression Values) to Cluster 1 (Secular-Rational and Self-Expression Values), reflecting an underlying shift from traditional to secular-rational values. As shown in Figure 10, the data indicates a decrease in national pride, less respect for authority, and a decline in the importance of God, all of which are consistent with this shift in values.

United States: temporal changes in selected indicators. These line plots illustrate the three selected features across three time points based on the mean survey responses in the United States. The color-coded line segments indicate the cluster labels at each time point, while the dashed lines denote the cluster definitions, which are static. These results are consistent with the United States’ movement from Cluster 2 to Cluster 1, reflecting a shift from traditional to secular-rational values. Analyses are based on SLA with
This case study demonstrates the utility of our temporal clustering framework for analyzing global cultural change using the WVS data. By applying SLA, we obtained detailed definitions of four cultural clusters and directly showed how these clusters relate to the ten major survey items commonly used to examine cross-national cultural change in the WVS. In addition, we illustrate a method for identifying important shifts in cluster labels by setting the value of the hyperparameter
Business Development Patterns in Chicago
In our final application, we examined the temporal patterns and shifts within Chicago neighborhoods based on their historical records of business establishments across several categories. Neighborhood change research has deep roots in urban sociology, dating back to the Chicago and Atlanta schools. Key concepts from this era, such as the concentric zone and neighborhood succession models, continue to inform contemporary research despite undergoing modifications over time. Mid-century advances introduced factorial ecology, which used factor analysis to classify neighborhoods based on economic, family, and ethnic dimensions. Although later research often shifted focus to single-variable approaches like segregation indices, geodemographic techniques reintroduced multivariate analysis (Silver and Silva 2021), emphasizing the importance of detecting complex neighborhood types and changes through cluster analysis.
The field has recently advanced through the incorporation of sequence analysis techniques, as exemplified by Delmelle’s (2016) work, which has expanded the geographic and temporal scope of multidimensional neighborhood change studies. This research confirms many earlier findings, while offering new insights into the spatial patterns of poverty, wealth, and ethnic diversity across urban landscapes. Adapted from the DNA and protein sequence alignment problem in biology (Carrillo and Lipman 1988), this approach often considers neighborhood change as a sequence analysis problem and employs a two-stage analysis of neighborhood changes (Delmelle 2016; Dias and Silver 2021; Kang et al. 2020; Patias et al. 2020). In this method, as noted previously, the first stage disregards the temporal dimension by treating each time point as an individual entity in K-means or other standard clustering methods, equivalent to SLA in the proposed framework. As mentioned in the prior case study using the WVS, SLA allows each entity to have a distinct cluster label at each time point, reflecting neighborhood changes within entities. However, the cluster definition remains fixed because temporal relationships are not considered during clustering. The second stage examines how neighborhoods change their type over time (e.g., shifting from “residential” to “commercial”). OM distance has often been used to re-group neighborhoods based on the similarity in their temporal sequences through cluster definitions. Evolving cluster definitions can then be obtained by aggregating temporal data from entities within the same group.
However, there are several limitations in this approach. Although OM is the dominant approach for conducting a sequence analysis in the social sciences (Aisenbrey and Anette 2010), the validity of measuring the similarity between sequences using OM is questioned (Levine 2000). Accordingly, there has been much debate and experimentation around how to conduct the second stage of the analysis (Ling and Delmelle 2016; Silver and Silva 2021). Also, the temporal dimension is only analyzed using the entities’ cluster label rather than the raw temporal data, which can exhibit information loss. The two stages do not use the same similarity measures or mathematical clustering objectives, raising concerns about potential biases in the clustering results. Moreover, the dynamic cluster definitions derived from aggregating neighborhoods’ temporal data after re-grouping differ from the initial static cluster definitions. However, the entities’ labels at each time point and their transitions between clusters still rely on the original static definitions. Consequently, the social changes within neighborhoods (changes in label) and the social changes across neighborhoods (changes in cluster definitions) are not identified under a consistent set of cluster definitions.
Due to these deficiencies, we propose that DC would be more suitable for neighborhood change studies. With DC, cluster definitions can be directly interpreted as temporal trends rather than static values, eliminating the need for a second-stage sequence analysis to summarize the temporal relationships using the OM distance, as has been done in previous studies. This reduces external factors that can introduce biases and eliminates possible information loss. As illustrated in prior case studies, adjusting the number of maximum cluster label changes,
The existing two-stage approach of neighborhood change analysis has been employed in studies of many cities. For instance, Delmelle (2016) investigated neighborhood changes in Chicago using various socioeconomic indicators, Kang et al. (2020) applied this approach to the New York metropolitan area across different census tracts, Silver and Silva (2021) studied Toronto, while Delmelle (2017) and Dias and Silver (2021) have expanded to larger sets of cities. We selected the temporal data of the business establishments in Chicago in this case study. Utilizing the unique features of DC, our objective extends beyond merely identifying various types of neighborhoods in Chicago based on temporal trends in business establishments and geographical distributions. We also aim to identify the neighborhood with the most atypical movement across the clusters. This approach can serve as an exploratory tool for further study to discern the causes behind these changes.
The temporal data for this study was sourced from the ZIP Codes Business Patterns (ZBP) (Bureau 2023), covering annual statistics from 2007 to 2016 on the number of establishments across various business sectors. These sectors are categorized according to the North American Industry Classification System (NAICS). These data have been featured in sociological research on the organizational resources in poor urban neighborhoods (Small and McDermott 2006) and the cultural profile of U.S. neighborhoods (Silver and Clark 2016), among others. Our analysis focuses on six specific types of business sectors, detailed in Table 4. This selection was made to balance computational feasibility with the diversity of business types analyzed. Out of the 61 ZIPs in Chicago, the study included data from 50 ZIPs. The selection criteria were based on the number of establishments in the six identified business types exceeding a predefined threshold. This approach ensures the elimination of ZIPs with near-zero values, which could potentially skew the analysis and results. The final temporal dataset for this case study contains 50 entities with six features across 10 time points.
NAICS codes and descriptions of six key features.
*Except Convenience Retailers.
Feature selection and eliminating abnormal entities also reduce the data points in this temporal dataset, making it feasible to utilize DC under MILP optimization, which guarantees globally optimal solutions and fully reproducibility results. The number of clusters
To demonstrate the unique capabilities of DC, we applied the existing approach, which corresponds to SLA without limiting cluster label changes in our framework, under the same setting (i.e.,
Cluster definitions using SLA
Recall that SLA disregards temporal relationships by clustering each time point independently. Applying SLA with a 3-cluster setting results in three static cluster definitions over time; each definition was obtained by aggregating all time points within this cluster using the mean (see Table 5).
Static cluster definitions for three neighborhood types using SLA the table below presents the values of the static definition for the three clusters (rows) across the six studied features (columns). These values represent the mean number of business establishments of different types across all time points within each cluster. Detailed descriptions of each neighborhood type are provided in the main manuscript. Analyses are based on SLA with
By analyzing the scores of the cluster definitions across different features, we identified the unique characteristics of each cluster based on their pattern of business establishments. Integrating with the spatial distribution of neighborhoods in each cluster (see Figure 11), the complete description of each cluster is stated below.

Geographical distribution of different neighborhood types using SLA. Given that neighborhoods’ cluster labels are dynamic, each neighborhood can belong to multiple clusters over time. The opacity of the choropleth map represents the frequency with which each neighborhood is assigned to a particular cluster. High opacity indicates that a neighborhood frequently belongs to this cluster, while low opacity suggests it is part of this cluster for a limited duration. The number of time points each neighborhood (minimum of one year) spends within this cluster is also annotated on the choropleth map. Hotels and Restaurants are predominantly concentrated in downtown Chicago, whereas Daily Needs establishments are primarily located in the outer regions. Although there is no clear concentration area for neighborhoods classified as Low Business Activity, some of these neighborhoods are located on the boundary of Chicago. Analyses are based on SLA with
As shown above, SLA can highlight important clusters in the data. However, the significant drawback of using SLA is its restriction that cluster definitions are temporally invariant, such that the characteristics of each cluster type consist of static comparisons rather than evolving patterns. For example, while SLA can identify that neighborhoods within the Hotels and Restaurants cluster tend to have more hotel-related business establishments than others, it does not account for whether this number increases over the years or if other business establishments remain constant or decline. Similarly, it cannot determine if neighborhoods in the Low Business Activity cluster consistently have lower numbers of all types of business establishments or if there is a shift from high to low starting from a particular time point.
Cluster definitions using DC
In contrast to SLA, DC utilizes dynamic cluster definitions, such that the cluster centers themselves are temporal trends (see Figure 12). The clustering results from DC, with

Definition of each neighborhood type using DC. Each of the three plots below illustrates the cluster definition of one of the neighborhood types when employing DC. In each plot, solid lines represent the values of business types (z-axis) after min–max normalization on each feature separately over various years (x-axis).
18
Each line corresponds to a different type of business establishment. Analyses are based on DC with

Geographical distribution of different neighborhood types using DC. The choropleth maps below depict the geographical distribution of neighborhoods within the three clusters identified by the DC clustering results. Each entity’s time points within the corresponding cluster are annotated on the choropleth maps. Analyses are based on DC with
Neighborhood volatility analysis
One of the key questions raised in neighborhood change research concerns the degree of volatility or continuity exhibited in different parts of a city (Silver and Silva 2021). A precondition to pursuing this question is the ability to reliably discern when a neighborhood is changing types. With SLA, a seemingly straightforward way to measure a neighborhood’s volatility is by counting the frequency of its label changes across the studied time points. More frequent changes indicate transitions between different neighborhood types, suggesting more volatility. In our 3-cluster SLA setting, eight entities experienced cluster label changes. ZIP 60643 stood out with four changes, alternating between Low Business Activity and Daily Needs clusters, making it the most frequently changing entity. However, a close examination of this case shows the limitations of existing SLA methods for identifying neighborhood volatility and the improved capacities offered by DC.
In fact, a thorough analysis of ZIP 60643’s temporal data (Figure 14) reveals no significant changes across the six considered features over time. The number of business establishments in “Child day care services” and “Supermarkets and grocery stores” shows contrasting values in the definitions of the two clusters to which ZIP 60643 was assigned. However, there are no substantial atypical trends that would justify these cluster label changes. This discrepancy suggests a potential misalignment between cluster labels and actual trends in the underlying features. One possible explanation for this misalignment is that ZIP 60643’s scores on these contrasting features lie on the boundaries between Low Business Activity and Daily Needs clusters, far from the centers of both clusters (Figure 15). This relative position makes ZIP 60643 vulnerable to statistical noise when clustering using conventional techniques, such as those commonly used in the literature on urban social change. Consequently, neighborhoods like ZIP 60643 may be incorrectly identified as an entity with atypical movements when using the conventional version of SLA, where

Neighborhood cluster label change using SLA (a) and (b) DC. These panels show line plots for the six features examined for two key neighborhoods. Specifically, the top panel (a) shows changes in the features for the most atypical cluster label change when using SLA with

Scatter plot of two key features in two selected clusters. The scatter plot below illustrates the distribution of time points within the clusters Low Business Activity and Daily Needs across two features (labeled on each axis) that show contrasting values at the centers of these clusters (marked with a star). ZIP 60643, marked with a red cross, is located on the boundary between the two clusters, distant from the cluster centers. This position is susceptible to statistical noise in center-based clustering methods, potentially leading to erroneous identification of entities as frequently changing clusters.
By contrast, when we restricted the maximum allowable number of changes in entities’ label to one (i.e.,
The case studies in this section have demonstrated the utility of our unified framework for temporal clustering in providing new insights into various dimensions of social change. In the first case study, we analyzed a unique dataset of state-level opinions, policies, and political affiliations in the United States. Our method identified distinct temporal patterns that reveal increasing polarization in both political affiliation and the degree of liberalism in state policies and residents’ opinions. Moreover, by restricting the number of entities allowed to move across clusters, we were able to highlight those states with notable transitions that deviate from the broader shifts in cluster definitions. The second case study examined global cultural change using data from the WVS. Applying SLA, we developed a typology of cultural clusters that more closely aligns with the underlying survey indicators, in contrast to previous work ((e.g., see Inglehart 2006)). We also identified several countries, such as the United States and Pakistan, that have undergone significant cultural shifts and have moved from one cluster to another. Finally, in the third case study, we used historical data on business establishments in Chicago neighborhoods to identify distinct patterns of local economic development. By comparing our DC method to more conventional approaches, we demonstrated that the latter is prone to misidentifying atypical shifts arising from statistical noise. Taken together, these diverse applications illustrate the flexibility and insights gained from our unified framework for analyzing complex patterns of social change across a range of sociological contexts and scales.
Conclusion
The study of social change has long been hindered by overly simple theories and methods for joining concepts to data (Blumer 1990; Boudon 1986; Haferkamp and Smelser 1992; Lenski 2005; Tilly 1984). As Hallinan (1997)noted in her presidential address to the American Sociological Association, the field has often been constrained by unrealistic assumptions of “continuity, linearity, and stable equilibrium (2).” She urged sociologists to develop “new models that better portray and explain complex, contemporary social events (2).” This call for “new approaches to theories” to a classic sociological topic of social change Šubrt (2017) has been echoed more recently, highlighting a persistent need for better methodological tools (52).
Our unified framework for temporal clustering directly addresses these calls. By allowing for dynamic cluster definitions and for cases to transition between clusters, our approach moves beyond the limiting assumptions of linearity and continuity critiqued by Hallinan and others. Our paper makes five main contributions to the study of social change. First is the unified framework itself, which expands the methodological toolkit for analyzing complex units such as societies, neighborhoods, and states. By incorporating both evolving cluster definitions and assignments, our framework generalizes existing methods while offering new techniques such as DC, allowing researchers to select an approach suitable to their specific research questions and data.
Second, we formalize the objectives for temporal clustering, extending traditional center-based methods to incorporate an explicit temporal dimension. We provide mathematical objectives for each of the four main types in our framework, accommodating various combinations of static and dynamic cluster definitions and labels. This formalization provides a clear basis for comparing different clustering approaches and ensuring consistent and reproducible results.
Third, we develop new algorithmic methods, particularly for DC and SLA. These methods optimize clustering objectives while allowing researchers to constrain how many times a unit can change clusters. We also offer two complementary optimization techniques: Mixed Integer Linear Programming (MILP) methods, which provide globally optimal and reproducible solutions, and a more scalable iterative technique based on the MM algorithm, useful for larger datasets.
Fourth, we provide practical guidelines for applying temporal clustering in sociological research. These guidelines address key decisions in the research process, including feature selection, normalization, and the choice of an appropriate distance metric, number of clusters, and number of allowable label changes.
Finally, we illustrate our framework with three case studies that highlight the various ways in which our methods can be used in practice: first, an analysis of U.S. states identifies patterns of increasing political polarization and highlights several states that shift clusters over time; second, using the WVS, we develop a typology of cultural clusters and identify countries that transition between them; and, lastly, a study of business patterns in Chicago shows how our approach avoids the misidentification of atypical shifts that can arise from statistical noise in conventional methods.
Our framework has limitations that point to directions for future research. A key issue is the quantification of uncertainty due to random sampling. Our iterative MM-based algorithm, for instance, can produce locally optimal results that are not always reproducible, due to random initialization. While the MILP approach is globally optimal, it does not fully resolve the issue. Future research could address this by incorporating statistical inference methods such as bootstrapping into the clustering process to evaluate the statistical significance of the clustering results.
Another limitation is a reliance on
Scalability for very large datasets is another area for future work. Our framework presents a tradeoff between a globally optimal (MILP) approach that can be computationally intensive and an efficient (MM-based) method that may be more scalable but less optimal. Future research could focus on developing hybrid methods to better balance this trade-off. Possibilities include machine learning techniques for data reduction or feature extraction, parallel computing strategies for MILP optimization, or advanced heuristics that approximate global optima more closely than current MM-based methods.
Finally, our framework’s applications extend beyond the geographical cases presented here. Our approach can be generalized to any feature space with a meaningful order, such as occupational hierarchies, educational levels, or socioeconomic strata. This allows for studying social mobility where the units represent positions in a social hierarchy. The framework is also well-suited for cohort analysis, where units can represent different birth cohorts, making it possible to track generational change (Fosse 2023a; Waite et al. 2021). While such applications require careful thought about defining similarity in abstract spaces, they highlight the broader potential of temporal clustering in sociological research.
Supplemental Material
sj-pdf-1-smr-10.1177_00491241251396789 - Supplemental material for Mapping Social Change: A Unified Framework for Temporal Clustering
Supplemental material, sj-pdf-1-smr-10.1177_00491241251396789 for Mapping Social Change: A Unified Framework for Temporal Clustering by Jiazhou Liang, Jolomi Tosanwumi, Ethan Fosse, Daniel Silver and Scott Sanner in Sociological Methods & Research
Footnotes
Declaration of Conflicting Interest
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Preregistration Statement
This study was not preregistered, as all datasets used are either publicly available or accessible with restrictions from third parties, and existed before the beginning of this research.
