Abstract
Background
Emerging evidence suggests that the built environment (BE) and social vulnerability of communities play a critical role in the Alzheimer's disease (AD) dementia burden. However, to what extent they are predictive of dementia burden nationally and regionally remains understudied.
Objective
This study investigates the cross-sectional associations between county-level BE and social vulnerability index (SVI) variables and AD dementia prevalence across the US and identifies the key variables within each region.
Methods
We applied eXtreme Gradient Boosting (XGBoost) and Random Forest (RF) models nationally and regionally using a combined set of BE and SVI variables. Variable importance was assessed using SHapley Additive Explanations values.
Results
Significant regional differences in AD dementia prevalence were observed (F = 126.56, p < 0.05). XGBoost consistently outperformed RF when using SVI and BE variables, with the highest predictive accuracy observed in the Northeast (R² = 0.92), followed by the South (R² = 0.73), West (R² = 0.72), and Midwest (R² = 0.72). In the Northeast, BE variables alone explained most of the variance, with grocery store density and walkability emerging as key variables. In contrast, the BE-only model for the South performed poorly but improved substantially with the inclusion of SVI variables. In the Midwest and West regions, both BE and SVI variables contributed more evenly to the model predictions.
Conclusions
While BE variables are more influential in urbanized regions with increased infrastructure (Northeast), SVI variables play a stronger role in more socioeconomically disadvantaged regions (South), underscoring the need for region-specific interventions.
Introduction
While individual-level factors such as genetics and lifestyle behaviors influence the development of Alzheimer's disease (AD) dementia, growing evidence suggests that neighborhood environments also play a critical role. Variables such as walkability, transportation access, and proximity to healthcare may shape dementia risk directly or indirectly by influencing cognitive health. For example, Clarke et al. (2015) found that neighborhoods with more community resources, accessible public transit, and well-maintained public spaces were associated with slower rates of cognitive decline in older US adults. 1 Similarly, Yu et al. (2022) reported that neighborhood street connectivity was positively associated with global cognition among older adults in the US, suggesting that built environment (BE) variables may promote cognitive resilience. 2
Recognizing these influences, researchers have increasingly focused on social determinants of health (SDOH) to better understand factors influencing cognitive outcomes. 3 A recent study by Ciciora et al. (2024) demonstrated that lower educational attainment, social isolation, and living in disadvantaged areas are significantly associated with increased risk for AD dementia in a nationally representative US sample. 4 To quantify neighborhood-level SDOH, the Centers for Disease Control and Prevention (CDC) developed the social vulnerability index (SVI). This composite index includes 15 census-based variables reflecting socioeconomic status, household composition, minority status and language, and housing and transportation. 5 Using this index, Desai et al. (2024) found that those living in counties with higher SVI scores had more than twice the risk of developing AD compared to those in less socially vulnerable areas, even after adjusting for age, sex, and APOE ε4 status. 6 While these findings underscore the impact of community-level social disadvantage on AD dementia risk, they represent only one dimension of the neighborhood context.
Beyond social vulnerability, increasing attention has turned toward understanding how the variables of the environment impact cognitive health. While the role of air pollution in AD dementia has been extensively documented—linking fine particulate matter and nitrogen dioxide to increased dementia risk,7,8—other aspects of the environment have received comparatively less attention. The 5D framework—which includes density, diversity, design, destination accessibility, and distance to transit—provides a structured approach for understanding the influence of BE characteristics on health outcomes.9,10 A scoping review conducted by Sturge et al. (2021) found that connections to society and interactions with natural environments and public spaces contribute to the well-being of those living with dementia. 11 Similarly, Ward et al. (2018) found that walkable environments, accessible public spaces, and familiar community resources foster social engagement, independence, and overall well-being in people living with dementia. 12
Machine learning (ML) models offer robust tools for exploring AD dementia by capturing complex, nonlinear relationships between neighborhood environment and disease risk. 13 ML models can process large and multidimensional datasets and provide accurate predictions of AD dementia problems. 14 For instance, Nori et al. (2019) developed an ML model that accurately predicted AD dementia onset using a combination of electronic health records and claims data. 15 Dadu et al. (2024) and Thompson et al. (2023) used ML models to predict prognosis and identify novel therapy targets for AD dementia.16,17 However, despite their predictive power, ML models are often criticized for limited interpretability. 18 Explainable artificial intelligence methods, such as SHapley Additive Explanations (SHAP), can help address this challenge by clarifying the contribution of individual variables to the model predictions. 18
Recognizing the importance of neighborhood characteristics in cognitive health, the National Institute on Aging has launched targeted funding initiatives to advance research in this area.19,20 Nevertheless, nationwide AD dementia studies examining the role of neighborhood variables, especially BE and SVI, remain limited. Much of the existing research in this area has been conducted at the patient level, with relatively little exploration of broader spatial patterns.11,21 Given the considerable heterogeneity in both neighborhood characteristics and AD dementia prevalence across the US,22,23 there is a pressing need for geographically nuanced investigations. Thus, this study explores the role of these variables nationally and regionally (i.e., Midwest, South, Northeast, and West) using two ML models. Specifically, we seek to answer the following questions:
How well do the ML models perform when predicting AD dementia prevalence using only BE variables at both national and regional levels? To what extent does BE-only model performance improve with the addition of SVI, across the US and by region? Which BE and SVI variables are most influential in predicting AD dementia prevalence nationally and regionally?
Findings from this study aim to provide useful insights about geographically tailored public health interventions and support AD dementia prevention efforts at the population level.
Methods
Data collection and preparation
Data on the estimated prevalence of AD dementia (as the response variable) at the county level (n = 3142) was obtained from the article published by Dhana et al. (2020). 23 Dhana et al. (2020) used data from the Chicago Health and Aging Project (CHAP), a large, population-based cohort of over 10,000 older adults aged 65 and above residing in Chicago, which included extensive neuropsychological testing and demographic data. Moreover, to estimate the probability of AD dementia, they employed a generalized additive quasibinomial regression model, adjusting for key demographic covariates including age, sex, race/ethnicity, and education. They applied this model to bridged-race postcensal population estimates from the National Center for Health Statistics, stratified by demographic group. This allowed them to generate county-level prevalence estimates of AD dementia for all US counties. Their approach provided a demographically adjusted, spatially comprehensive dataset uniquely suited for geographically explicit epidemiological modeling and served as the foundation for our study.
The SVI, which reflects demographic and socioeconomic factors that influence how communities respond to stressors, 5 was obtained at the county level from the Centers for Disease Control and Prevention (CDC) and the Agency for Toxic Substances and Disease Registry (ATSDR) (https://www.atsdr.cdc.gov/place-health/php/svi/index.html). Supplemental Table 1 lists and describes the SVI variables used in this study, with further details available in CDC/ATSDR SVI Documentation. 24
To assess the contribution of BE to AD dementia prevalence, the 5D framework was used, which includes ‘density’, ‘diversity’, ‘design’, ‘destination accessibility’, and ‘distance to transit’ categories. 10 The variables in these categories were obtained from the EPA Smart Location Database, a data product and service provided by the US EPA Smart Growth Program (Supplemental Table 1). 25 Additional destination variables included ‘access to grocery and convenience stores’ sourced from the USDA Food Environment Atlas, 26 and ‘access to hospitals’ from the Homeland Infrastructure Foundation-Level Data (HIFLD) database. 27 A more detailed description of these variables is available in the EPA Smart Location Database Technical Documentation and User Guide. 25
To account for rurality and regional variation, Rural-Urban Continuum Codes (RUCC) were used. The RUCC, provided by the US Department of Agriculture, distinguishes US metropolitan counties by the population size of their area and non-metropolitan counties by their degree of urbanization and proximity to a metropolitan area. 28 Following Pruitt et al., RUCC categories were grouped: 1 = large metropolitan areas; 2 = medium metropolitan areas; 3 = small metropolitan areas; 4–7 = urban non-metropolitan areas; and 8–9 = rural non-metropolitan areas. 29 Supplemental Table 1 provides a complete list of all variables, along with their descriptions and corresponding data sources.
All data were collected or prepared at the county level and joined to the US Census TIGER/Line shapefiles. Spatial join operations were performed in ArcGIS Pro (ESRI, Redlands, CA) to calculate the number of hospitals, grocery stores, and convenience stores in each county.27,30 Census tract-level variables from the EPA Smart Location Database were aggregated to the county level using median values (Table 1).
Descriptive statistics for BE and SVI variables classified by US region.
Exploratory analysis
Analysis of variance (ANOVA) tests were conducted to assess differences in mean AD dementia prevalence across US regions (i.e., Midwest, South, Northeast, and West). Where significant differences were detected, Tukey's HSD tests were used to identify which regional pairs differed significantly. A similar procedure was applied across RUCC categories to examine rural-urban differences for each US region.
Multicollinearity among BE variables was evaluated using exploratory regression in ArcGIS Pro, and predictors with variance inflation factors (VIF) ≥ 5 were excluded. After incorporating SVI variables, we reassessed multicollinearity for the combined model and removed any remaining variables with VIF ≥ 5 or high pairwise correlations (|r| > 0.75).
Machine learning models
Following data cleaning and preprocessing, two tree-based ML models—Random Forest (RF) and eXtreme Gradient Boosting (XGBoost)—were implemented to predict AD dementia prevalence. These models were selected due to their ability to handle nonlinear relationships, variable interactions, and structured tabular data.31–34 They were also chosen specifically for their strong predictive performance in health research and their ability to provide interpretable measures of variable importance.33,34 Additionally, both models are relatively easy to implement and tune, and highly flexible with respect to handling missing data and different data types.33,34 Modeling began with BE variables alone to assess their independent predictive performance. SVI variables were subsequently added to assess improvements in model performance.
Random forest
RF is an ensemble learning model that makes predictions by combining the results of many individual decision trees. Each tree in the model makes its own prediction, and the final output is the average of those predictions. 31 To ensure that the trees are not all the same, the model creates each one using a different random sample of the data and a random selection of input variables. 35 This approach helps reduce bias and avoid overfitting. RF is widely used in health research because it can handle large, complex datasets and capture nonlinear relationships that may not be obvious. 31 The model was trained using the Scikit-Learn library in Python. A detailed description of the RF is provided elsewhere. 36
XGBoost
To further enhance predictive accuracy, we employed XGBoost, 37 an optimized gradient-boosting framework. Previous studies have found better performance of XGBoost compared to RF.38,39 Unlike RF, which builds trees independently, XGBoost builds them sequentially. Each new tree is trained to correct the errors made by the preceding trees, allowing the model to progressively improve its performance. 37 Additionally, XGBoost includes built-in safeguards to prevent overfitting, helping the model generalize well to new data. 37 Its speed and high accuracy have made it a popular choice in AD dementia studies.32–34 In this study, we trained the XGBoost model using the ‘xgboost’ library in Python. A detailed description of the XGBoost is provided elsewhere. 39
Feature importance
Model interpretability was assessed using SHAP values that quantify the contribution of individual variables to the model's predictions. 18 SHAP values offer a consistent, model-agnostic approach based on cooperative game theory, by estimating the marginal contribution of each feature across all possible feature combinations to the model. 18 To enhance interpretability, we visualized these scores using SHAP beeswarm plots, which illustrate the distribution and magnitude of each variable's contribution across counties.
Model settings
The dataset was randomly partitioned into training (70%), validation (15%), and test (15%) sets. The training set was used to fit the models, the validation set was used for hyperparameter tuning and overfitting prevention, and the test data was used for final model evaluation.
Hyperparameter tuning was conducted using grid search and optimized via five-fold cross-validation with root mean squared error (RMSE) as the evaluation metric. 40 For RF models, tuned parameters included the number of trees, maximum depth, minimum samples per split, minimum samples per leaf, and the number of variables considered for splitting. 41 For XGBoost models, tuned parameters included the learning rate, maximum tree depth, subsample ratio, and column sampling ratio. 42 Moreover, early stopping was applied during XGBoost training to terminate the process if validation performance did not improve after 10 consecutive rounds. 43 To enhance transparency and reproducibility, a GitHub repository has been provided containing the code used for data processing, model development, and analysis.
Model evaluation
Model performance was evaluated on the test set using RMSE, mean absolute error (MAE), R², and Akaike Information Criterion (AIC). AIC was calculated based on the residual sum of squares and the number of variables to assess model complexity, where lower AIC values indicate a better balance between model fit and complexity. 44 To ensure consistency and comparability across all model configurations, we applied the same set of performance metrics for evaluating models using BE-only, and combined BE + SVI predictors at both national and regional levels. Mapping and visualization of key variables for each region were conducted in the ArcGIS Pro environment. Results were considered statistically significant if p < 0.05.
Results
Preliminary descriptive statistics
Preliminary descriptive statistics of AD dementia prevalence reveal regional differences across the US. Figure 1 depicts the boxplot of prevalence by US Census regions. The South exhibits the highest mean (11.62%) and median (11.30%) prevalence, along with the largest range (12.80%) and standard deviation (1.66%), indicating both higher and more variable rates. The Midwest had a mean prevalence of 10.93% and a median of 10.90%, with relatively low variability (standard deviation = 0.92%) and a more concentrated distribution (range = 7.30%). The Northeast has almost the same mean prevalence (10.87%) but a slightly lower median (10.60%), with moderate variability (standard deviation = 1.11%) and a range of 8.30%. The West had the lowest mean (10.30%) and median (10.10%) prevalence, though it exhibits the second-highest variability (standard deviation = 1.35%) and the second-largest range (10.60%), suggesting more dispersion in prevalence rates across counties.

Boxplot of AD dementia prevalence across U.S. regions.
An ANOVA test showed significant regional differences in AD dementia prevalence across the US (F = 126.56, p < 0.05), rejecting the null hypothesis of equal mean AD dementia prevalence by regions. Post-hoc Tukey's HSD test further found that all regional pairs were significantly different (p < 0.05), except the Midwest and Northeast (p > 0.05). Table 2 shows the results of Tukey's HSD test across regions.
Tukey's HSD test results.
Classified by RUCC categories, ANOVA tests showed no significant differences in AD dementia prevalence nationally (F = 1.25 and p > 0.05). However, regional analyses showed notable differences. In the Northeast, the ANOVA test was highly significant (F = 23.04, p < 0.05), with Tukey's HSD test showing large metro areas differed significantly from all other categories, and medium metro areas were also significantly different from rural non-metro and urban non-metro areas. In the South, the ANOVA test was significant (F = 5.62, p < 0.05), with Tukey's HSD test indicating urban non-metro areas were significantly different from small metro, large metro, and medium metro areas. In the Midwest, the ANOVA test produced a highly significant result (F = 15.93, p < 0.05), with Tukey's HSD test showing large metro areas differed significantly from rural non-metro and urban non-metro areas and rural non-metro areas were significantly different from all other categories. In contrast, the West showed no significant differences in AD dementia prevalence by RUCC category (F = 1.95, p > 0.05), suggesting more homogeneity in prevalence rates between rural and urban areas. Supplemental Tables 2–4 show complete Tukey's HSD test results across RUCC categories.
Variable selection
Final multicollinearity checks confirmed that all retained variables had VIF values below 5 and all the inter-variable correlation coefficients were less than 0.75. The final model included 15 variables representing both BE and SVI characteristics. Table 3 presents the VIF values and maximum inter-variable correlations for each selected variable. See Supplemental Figure 1 for the correlation matrix for all included variables.
VIF and maximum correlations across included predictors.
Hyperparameters
For nationwide models using both BE and SVI variables, the best-performing XGBoost model was tuned using a maximum tree depth of 6, a learning rate of 0.01, a training subsample ratio of 0.8, and a column sampling ratio of 0.6. The corresponding RF model performed best with 300 estimators, maximum depth constraint of 20, square root feature selection at each split, a minimum of two samples required to split a node, and a minimum of one sample per leaf. Moreover, hyperparameter tuning was conducted separately for each region to optimize model performance based on regional characteristics. The selected hyperparameters for all XGBoost and RF models—using either BE variables alone or both BE and SVI variables—are summarized in Supplemental Tables 5 and 6, respectively.
Regional and national model performance
Across both nationwide and regional models, XGBoost consistently outperformed RF in predictive accuracy when using both BE and SVI variables. The nationwide XGBoost model achieved moderate predictive performance, with an R² of 0.51 and RMSE of 1.05, while the RF model showed a slightly weaker fit (R²=0.50, RMSE = 1.06). Regionally, XGBoost models demonstrated stronger performance, particularly in the Northeast (R²=0.92), followed by the South (R²=0.73), the West (R²=0.72), and the Midwest (R²=0.72). In contrast, regional RF models showed lower accuracy overall, with the strongest performance in the South (R²=0.62) and weakest in the Midwest (R²=0.34). Table 4 summarizes the performance metrics for all nationwide and regional models using both SVI and BE variables. Performance metrics for models trained solely on BE variables are included in Supplemental Table 7.
Model performance metrics for XGBoost and RF Using BE and SVI variables across US regions.
Nationwide feature importance ranking
Figure 2 presents the SHAP summary results, illustrating the relative importance and directional impact of BE and SVI variables on AD dementia prevalence across the US as identified by the XGBoost (Figure 2A) and RF (Figure 2B). In both models, SVI variables, such as the “percentages of single-parent households”, “households without a vehicle”, and “people uninsured” emerged as dominant variables. The BE variable “grocery store counts” remained among the top five in terms of variable importance for both models. For each of the top five variables identified, higher values were associated with higher AD dementia prevalence.

SHAP summary plot showing the relative importance and impact of BE and SVI variables on AD dementia prevalence as identified by (A) XGBoost and (B) RF models. Positive and negative SHAP values indicate the direction of each variable's influence.
Regional feature importance ranking
Figures 3 and 4 display the top variables of AD dementia prevalence for each US region for both XGBoost and RF models, respectively, using BE and SVI variables. In the South, SVI variables emerged as primary variables of AD dementia prevalence. Notably, the “percentage of households without a vehicle”, “percentage of single-parent households”, and the “percentage of uninsured individuals” ranked among the top variables in this region. A higher proportion of each of these predictors was associated with increased AD dementia prevalence across both RF and XGBoost models.

Top variables associated with AD dementia prevalence across US regions identified by XGBoost.

Top variables associated with AD dementia prevalence across US regions identified by RF
In contrast, models for the Northeast showed greater reliance on BE variables. The “grocery store counts” and “walkability index” consistently ranked among the most important variables in both XGBoost and RF models for this region. “Employment entropy” also emerged as an influential BE variable in the Northeast for the XGBoost model. In the Midwest and West, both BE and SVI variables contribute to model predictions, reflecting a more balanced pattern compared to other regions. In the Midwest, “grocery store count” and “walkability index” remained important BE variables, while housing-related SVI variables—such as the “percentages of mobile homes” and “group quarter living”—were also among the most influential variables. In the West, important BE variables include “grocery store counts” and “employment entropy”, while key SVI variables include “per capita income” and “percentage of people uninsured”.
Supplemental Table 8 presents the regional variable importance for BE-only variables.
Regional variations versus RUCC
Regional differences in key variables corresponded with variation in rurality, as indicated by median RUCC values. The Northeast had a median RUCC of 3, indicating a more urban composition, which aligns with the greater importance of BE variables such as “grocery store counts”, “walkability index”, and “employment entropy” in that region. In contrast, the Midwest, South, and West all had a median RUCC of 6, reflecting more rural characteristics. In these regions, SVI variables —particularly those related to housing, income, and access to resources—played a more prominent role in predicting AD dementia prevalence.
Discussion
To our knowledge, this is the first study to assess national and regional associations between social and built environment characteristics and AD dementia prevalence in the US. 45 Using XGBoost and RF models with SHAP values, we identified key predictors and quantified their contributions. Large differences in performance were found between XGBoost and RF regionally, but not at the national level. We believe this is due to XGBoost's sequential boosting framework which is generally more effective in identifying and leveraging these subtle patterns than RF, which relies on averaging across many randomized trees. 46 In contrast, at the national level, the model is trained on a much larger and more heterogeneous dataset, where the complexity of local interactions is diluted, leading to smaller relative performance gains between the two algorithms.
Nationally, XGBoost slightly outperformed RF, but BE variables alone showed poor predictive power. Regional analysis of XGBoost revealed strong performance in the Northeast, moderate in the Midwest and West, and poor in the South when using BE variables alone. Adding SVI variables substantially improved model fit, especially in the South, emphasizing the importance of socioeconomic factors (e.g., employment, single-parent households) in these regions. In contrast, BE characteristics, such as grocery store access and walkability, were more influential in the Northeast. This may be partly due to differences in infrastructure, access to resources, and underlying social determinants of health across the US regions. Prior studies have similarly found that variables such as walkability, transportation access, and socioeconomic disadvantage are linked to cognitive health outcomes and dementia risk, particularly in underserved communities.4,11,47
Surprisingly, our findings suggest that the counties with higher median walkability scores experience higher AD dementia prevalence. While walkability is typically associated with health-promoting behaviors, this association may reflect that in more densely populated areas, environmental stressors such as air pollution, noise pollution, and limited green space contribute to cognitive decline. 48 Urban environments are characterized by overcrowding and social disparities, which are linked to chronic stress and may adversely affect cognitive function. 49 The positive association between higher grocery store density and AD dementia prevalence could reflect the broader urban environment. Higher concentrations of grocery stores tend to occur in more urbanized areas, which may also have higher levels of socioeconomic inequality and limited access to healthy food options, even when grocery stores are more numerous. 50 This “food swamp” paradox—where high grocery store density coexists with limited availability of fresh and healthy food—can lead to increased reliance on processed and unhealthy foods. Such dietary patterns have been linked to greater cognitive decline and increased dementia risk.51,52 Therefore, while the count of grocery stores may be higher, the quality of a person's diet in these communities may be poorer, contributing to the higher AD dementia prevalence rates found in this study, which requires further investigation.
This study found that lack of vehicle access is a variable of higher AD dementia prevalence, particularly in the Southern region of the US. These findings align with previous research suggesting that limited vehicle access can impede timely access to healthcare, potentially delaying diagnosis and treatment. 53 Limited vehicle access may be especially consequential for accessing specialist care in rural or underserved areas. Additionally, transportation barriers may increase social isolation, which has been linked to poorer cognitive outcomes. 54 Findings for the other three common SVI variables—mobile homes, unemployment, and lack of insurance—varied by region.
Associations with mobile home prevalence were mixed, both aligning with and contradicting prior research. 55 In some regions, strong social networks and access to local health resources may mitigate risks, while in others, inadequate infrastructure and limited healthcare access likely heighten AD dementia vulnerability among mobile home residents. Similarly, regional differences in the association between uninsured (%) and AD dementia prevalence reflect variations in healthcare infrastructure. While prior studies have found that being uninsured is linked to delayed diagnosis and lower disease awareness, 56 access to safety-net services—such as free clinics or state-funded programs—may buffer these effects in some areas. 57 Furthermore, findings on the role of unemployment (%) were also mixed, consistent with prior work by Toth et al. showing insignificant or inconsistent associations with AD dementia mortality. 58 The health impact of unemployment may depend on the availability of social protections; where benefits and public healthcare are limited, unemployment may amplify stress and reduce access to care, increasing dementia risk. 59
Variable importance and model performance varied across US regions, likely reflecting differences in urbanization and the relative influence of BE and SVI variables. In the more urbanized Northeast (median RUCC = 3), models performed well using BE variables alone, with grocery store counts, walkability index, and transit frequency emerging as top variables—highlighting the role of infrastructure and access in shaping cognitive health outcomes. 53 The limited improvement after adding SVI variables suggests a smaller role for social vulnerability, potentially due to stronger safety nets to support vulnerable populations in this region. 60 In both the Midwest and West (median RUCC = 6), both BE and SVI variables contribute meaningfully to model performance, reflecting a more balanced influence of physical and social factors. In the Midwest, variables like mobile homes (%), walkability, and uninsured (%) showed varied associations, possibly due to local variations in infrastructure, food access, and environmental stressors linked to cognitive decline.50,51,61 In the West, variables like per capita income and unemployment further emphasized the role of economic disparities, where financial barriers may limit access to preventive care and cognitive health services.51,52
In the South (median RUCC = 6), models performed poorly using BE variables alone but improved significantly with the inclusion of SVI variables. Key variables in this region included no-vehicle (%), uninsured (%), and single-parent households—highlighting the dominant role of socioeconomic vulnerability in AD dementia prevalence. In the South, limited transportation infrastructure and weaker social safety nets may amplify the impact of vehicle access and insurance coverage on healthcare access and early diagnosis.53,54 These findings suggest that in more rural and socioeconomically disadvantaged regions, addressing structural barriers such as transportation, insurance, and income inequality might be essential for reducing AD dementia risk.
This study has several limitations. First, the use of county-level data may obscure sub-county variation and introduce an ecological fallacy, where associations observed at the county level do not reflect individual-level or sub-county level burden. However, this geographic scale was dictated by the data source, which was the first-ever nationwide dataset of AD dementia prevalence at the most resolved geographic scale, i.e., county level. Aggregation at this level may mask important neighborhood-level differences in socioeconomic status or BE characteristics that influence cognitive health outcomes. Moreover, while SHAP values provide valuable insights into the relative importance of predictors, these results should be interpreted as exploratory and hypothesis-generating rather than as evidence of causality or as a basis for policy or intervention decisions. The cross-sectional nature of this study limits the ability to draw causal conclusions. Longitudinal studies are needed to assess temporal relationships and identify potential causal pathways. Third, as noted by Dhana et al., their approach did not account for regional differences in other important risk factors—such as cardiovascular and lifestyle factors—which are unavailable at the county level and may modify the relationship between demographics and dementia risk. 23
Future studies should investigate the mediating effects of built and social environment characteristics on rural-urban disparities in AD dementia. This study did not explicitly account for spatial autocorrelation or uncertainty; therefore, future research may benefit from developing Bayesian methods to address geographic dependencies and improve predictive accuracy. Incorporating finer-scale geographic data, such as ZIP code or census tract-level measures, may enable a more accurate understanding of neighborhood-levels effects, supporting the development of targeted interventions and enhancing model precision. Disentangling the distinct contributions of BE and SVI variables to AD incidence and survival would provide deeper insight into the mechanisms driving spatial disparities. However, county-level data on AD incidence and survival are not currently available at a national scale, limiting our ability to explore these dimensions directly. Assessing the persistence of BE and SVI associations with AD prevalence over time would offer valuable insight into the temporal stability of these relationships. However, historical BE datasets at the county level are currently limited in availability and consistency, restricting our ability to implement formal temporal validation. As more county-level data become available, investigating the longitudinal associations of BE and SVI characteristics with both AD dementia prevalence and incidence would substantially strengthen the evidence base.
These findings have important public health implications. Identifying region-specific variables of AD dementia can help guide targeted prevention strategies and resource allocation. For example, in more urban regions like the Northeast, interventions may focus on enhancing access to walkable infrastructure and neighborhood services. In contrast, in more rural regions such as the South and West, efforts may be better directed toward addressing socioeconomic vulnerabilities, including lack of insurance, limited vehicle access, and unstable housing. By aligning local policies and community programs with regionally relevant risk factors, public health efforts can be more effectively tailored to reduce disparities in cognitive health outcomes across the US.
Conclusion
In summary, this study highlights the significant role of neighborhood characteristics in explaining regional differences in AD dementia prevalence across the US. BE variables were the strongest variables in more urbanized regions like the Northeast, while SVI variables were more influential in the South, where healthcare disparities and economic vulnerability are more pronounced. Across both nationwide and regional models, XGBoost consistently outperformed RF in predictive accuracy. Additionally, regional models demonstrated better performance than national models, underscoring the value of region-specific modeling approaches. These findings suggest that developing region-specific public health interventions, addressing both built and social environments, is essential for reducing AD dementia disparities in the US.
Supplemental Material
sj-docx-1-alz-10.1177_13872877251372144 - Supplemental material for The role of built and social environments in Alzheimer's disease dementia prevalence in the United States: A machine learning approach
Supplemental material, sj-docx-1-alz-10.1177_13872877251372144 for The role of built and social environments in Alzheimer's disease dementia prevalence in the United States: A machine learning approach by Mackenzie Kramer, Alexander V Alekseyenko and Abe Mollalo in Journal of Alzheimer's Disease
Footnotes
Ethical considerations
This study did not involve research on individual human participants or animals conducted by the author, and thus, ethical approval was not required.
Author contributions
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: AM and AVA are supported by the South Carolina SmartState Endowed Center for Environmental and Biomedical Panomics (CEABP).
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
