Abstract
Background
Dementia is a common complication of type 2 diabetes mellitus (T2DM), influenced by both genetic susceptibility and social disadvantages. While polygenic risk scores (PRS) have been widely applied to assess genetic vulnerability, the contribution of social determinants and their interaction with genetic risk are less understood.
Objective
This study aimed to investigate the independent and joint effects of PRS and polysocial risk scores (PsRS) on post-T2DM dementia risk.
Methods
A prospective cohort study was conducted using UK Biobank data. PsRS and PRS were derived from multidimensional social and genetic indicators, respectively. Cox proportional hazards models were used to examine their associations with dementia outcomes. In genetically susceptible individuals, seven machine learning models were applied to predict dementia risk. SHAP and ALE were used to interpret feature importance.
Results
Among 5,624 participants with T2DM, those in the highest PsRS group had a markedly elevated risk of all-cause dementia (HR = 2.738; 95% CI: 1.556–4.818, p < 0.001). This association persisted among genetically susceptible individuals. Machine learning analyses in the medium-to-high PRS group showed that eXtreme Gradient Boosting (XGBoost) achieved the best predictive performance (F1 = 0.735, AUC = 0.726). SHAP interpretation highlighted employment status (mean |SHAP value| = 0.45) and educational level (mean |SHAP value| = 0.19) as the strongest social contributors to dementia risk.
Conclusions
Both social disadvantages and genetic susceptibility contribute to dementia risk in individuals with T2DM. These findings underscore the importance of addressing modifiable social factors in targeted dementia prevention strategies for high-genetic-risk populations.
Introduction
Dementia, particularly among individuals with type 2 diabetes mellitus (T2DM), poses a growing global public health challenge. 1 According to the Global Burden of Disease 2019 report, dementia accounted for over 28 million disability-adjusted life years worldwide, ranking as the seventh leading cause of death. 2 Individuals with T2DM face a 50%-100% higher risk of developing dementia, compared to those without diabetes. 3 Increasing evidence suggests that T2DM accelerates neurodegenerative processes through multifaceted mechanisms, including insulin resistance, vascular dysfunction, and chronic inflammation.4,5 Although clinical and genetic risk factors for dementia have been extensively investigated, the impact of social disadvantage remains underappreciated.
Social disadvantage is recognized as a major upstream determinant of health, often outweighing the influence of medical care itself.6,7 However, most previous epidemiological studies only quantified the contribution of a single social determinant of health, overlooking the complex interconnection. To capture the complex nature of social conditions, a novel metric termed the polysocial risk score (PsRS) was recently developed. 8 Unlike traditional approaches that examine isolated social determinants such as education or income, the PsRS provides a multidimensional index that aggregates exposures across various domains including education, employment, housing, and social support.9–11 Thus, it offers a more ecologically valid estimate of social vulnerability. This composite score has shown promise in predicting T2DM and common dementia outcomes, yet its application in the context of post-diabetic dementia remains unexplored.12,13
Genetic predisposition is another key contributor to dementia. 14 Genome-wide association studies (GWAS) have identified multiple susceptibility loci for Alzheimer's disease (AD) and vascular dementia (VaD), leading to the development of polygenic risk scores (PRS), which improve the early identification of high-risk individuals. 15 Some evidence suggests that favorable socioeconomic factors, such as education and income, may attenuate the impact of high PRS. 16 However, in T2DM populations, the combined influence of genetic risk and social disadvantage on dementia remains unclear. While PRS has been widely used to stratify genetic risk, whether broader social conditions captured by PsRS can modify this risk is unknown. To date, no study has examined the joint effects of PRS and PsRS on dementia outcomes in individuals with T2DM, nor has any identified which social domains contribute most significantly to dementia risk in this group.
Therefore, this study leveraged data from the UK Biobank to investigate the independent and joint associations of PRS and PsRS with dementia risk in individuals with T2DM. Specifically, we focused on the medium and high PRS subgroups and constructed a predictive model for all-cause dementia (ACD) using PsRS components alongside clinical factors. This study aimed to delineate socially modifiable risk factors that may inform precision prevention strategies for dementia in genetically susceptible populations.
Methods
Study design and population
This study was conducted using data from the UK Biobank (UKB), a large-scale prospective cohort study that recruited over half a million participants from the general population of England, Wales, and Scotland between 2006 and 2010. 17 Participants provided comprehensive information on sociodemographic characteristics, lifestyle factors, medical history, and physical and cognitive function through baseline questionnaires, physical examinations, and biological sample collections. 18 Longitudinal follow-up has been maintained through linkage with national electronic health records, including hospital admissions, death registries, and primary care records. The UK Biobank study was approved by the North West Multi-Center Research Ethics Committee (REC reference: 11/NW/0382). Detailed information on UK Biobank's ethical oversight is available at https://www.ukbiobank.ac.uk/learn-more-about-uk-biobank/about-us/ethics.
The study population was restricted to participants with T2DM. Participants were excluded if they failed genetic quality control procedures, had withdrawn consent, or had a diagnosis of any form of dementia prior to baseline assessment. Participants with missing data on T2DM status, dementia outcomes, or relevant covariates were also excluded from the final analysis (Supplemental Figure 1).
Construction of polysocial risk score
The Healthy People 2030 Initiative (https://health.gov/healthypeople) defines five core domains of social determinants of health (SDOH): social and community context, education access and quality, economic stability, healthcare access and quality, and neighborhood and built environment. 19 In addition, several frameworks have highlighted socially stratifying factors that may contribute to disparities in health outcomes. 20 In this study, we selected 15 SDOH-related variables based on previous literature and the data availability within the UK Biobank cohort.21–24 All the variables were classified into four domains: social and community context, education access and quality, economic stability, and neighborhood and built environment.
Each variable was recoded into a binary format. Specifically, as detailed in Supplemental Table 1, the “Reference answer” for each variable represents the favorable condition (assigned a score of 0), whereas any status reflecting a worse or more deprived condition compared to this reference was coded as 1 (indicating a socially disadvantageous condition). Responses such as “Do not know” or “Prefer not to answer” were treated as missing data (Supplemental Table 1). All 15 variables were retained in the final model to comprehensively capture the multifaceted nature of social vulnerability. A PsRS was then calculated for each participant by summing the scores of these 15 individual variables, resulting in a total score ranging from 0 to 15. Participants were subsequently categorized into three social risk groups based on their PsRS: low (0–5), medium (6–10), and high (11–15).
Construction of polygenic risk score
To estimate participants’ genetic predisposition to dementia, we constructed a weighted PRS based on summary statistics from a published GWAS meta-analysis focused on AD. 25 This meta-analysis included 71,880 clinically diagnosed cases (including proxy cases) and 383,378 controls of European ancestry (p < 5 × 10−8). A total of 94 single nucleotide polymorphisms (SNPs) were selected according to genome-wide significance thresholds and linkage disequilibrium pruning (r2 < 0.1; distance threshold, 250 kb), ensuring minimal correlation between loci (Supplemental Table 2). Genotype dosage data for these SNPs were extracted from the UK Biobank imputed dataset (BGEN format), and polygenic scores were computed using PLINK 2.0 for data processing and scoring. 26
The PRS was computed as the weighted sum of risk alleles (coded 0, 1, or 2) at each locus, with weights corresponding to β coefficients (the log of odds ratio) derived from the GWAS and normalized by the total sum of β coefficients across all included SNPs. A higher PRS indicated a greater genetic susceptibility to dementia. Based on the distribution of PRS in the study sample, participants were categorized into three genetic risk groups: low (lowest quintile), intermediate (quintiles 2–4), and high (highest quintile).
Study endpoints
All participants in the UK Biobank were registered with the National Health Service (NHS), allowing for prospective follow-up via linkage to national electronic health records, including hospital admissions, death registries, and primary care records. 18 The primary study endpoints were incident cases of ACD, AD, and VaD, whose International Classification of Diseases codes could be found in Supplemental Table 3. For this analysis, follow-up was censored on August 20, 2022, ensuring consistency across all regions (England, Scotland, and Wales) in accordance with the latest available update of linked health records provided by UK Biobank.
Measurements of covariates
Baseline covariates were collected through touchscreen questionnaires, nurse-led interviews, and physical examinations. All multivariable models were adjusted for demographic characteristics including age (Field ID 21022), sex (Field ID 31), and body mass index (BMI, Field ID 21001). For lifestyle and clinical risk factors, adjustments included smoking status (Field ID 20116) and alcohol consumption (Field ID 20117), both self-reported and categorized into three groups: never, previous, and current. Medical history covariates included diagnoses of cancer (ICD-10 codes C00–C97), stroke (I60–I64), hypertension (I10–I15), and coronary heart disease (I21–I25), coded as 0 for no and 1 for yes. Medication use was also included (Field IDs 6153 for females and 6177 for males), covering current treatment with insulin, cholesterol-lowering agents, and antihypertensive medications, all coded as binary variables: 0 = no, 1 = yes.27,28 We additionally adjusted for APOE (apolipoprotein E gene; OMIM:107741) ε4 carrying status (not applied in analyses involving PRS), self-reported family history of dementia, and medical history of 3 common diseases, including cancer, diabetes, and cardiovascular disease (CVD). Participants were labeled as APOE ε4 carriers if they carried 1 or 2 copies of the APOE ε4 allele. Detailed covariate definitions and coding schemes are provided in Supplemental Table 4.
Statistical analysis
Baseline characteristics were summarized across the 3 PsRS level groups and compared using the chi-squared test for categorical variables and analysis of variance (ANOVA) for continuous variables. Follow-up time was calculated from the date when participants first visited the assessment centers until the date of first diagnosis, death, loss to follow-up, or the updating date of the database. We first applied Cox proportional hazards regression models to evaluate the associations between PsRS and the incidence of ACD, AD, and VaD. The proportional hazards assumption was assessed using Schoenfeld residuals. Analyses were performed both in the entire cohort and in a subgroup of participants with intermediate and high genetic risk (defined as the upper 80% of the PRS distribution). Hazard ratios (HRs) with 95% confidence intervals (CIs) were reported. Model performance was additionally assessed using prediction-focused performance (PFP) metrics to interrogate the theoretical percentage of total dementia burden due to disadvantageous social conditions. To explore potential nonlinear associations between PRS, PsRS, and dementia risk, we fitted restricted cubic spline models with four knots placed at the 5th, 35th, 65th, and 95th percentiles of each risk score. These models were applied separately for each dementia subtype to explore dose-response relationships.
We then conducted stratified analyses to investigate whether the association between PsRS and ACD risk varied across different PRS strata. For participants with intermediate or high genetic risk, we further developed machine learning models to predict incident ACD based on the 15 individual PsRS components and selected clinical and demographic covariates. Seven machine learning algorithms, including Logistic Regression (LR), Support Vector Machine (SVM), Gradient Boosting Decision Tree (GBDT), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Categorical Boosting (CatBoost), and Adaptive Boosting (AdaBoost), were employed to develop the predictive models for ACD. To address class imbalance, a mixed sampling strategy combining oversampling and undersampling was employed. 29 In the prediction task, 80% of the participants were randomly selected to construct the training set, while the remaining 20% were integrated into the independent test set. To optimize model performance, we applied Optuna based on Bayesian optimization for hyperparameter tuning with a predefined search space for each algorithm. 30 After tuning, each candidate model's performance was evaluated on the independent test set using multiple metrics, including area under the receiver operating characteristic curve (AUC), area under the precision-recall curve (AUPRC), F1 score, accuracy, and Brier score. Model training was conducted using five-fold cross-validation (Supplemental Table 5). The final model was selected based on its overall performance across these metrics in the independent testing set (Supplemental Table 6).
To explain the model decision process, we applied SHapley Additive Explanations (SHAP) to quantify feature contributions and Accumulated Local Effects (ALE) plots to visualize marginal effects. 31 SHAP provided both individual- and population-level insights into feature importance, while ALE captured the average effect of each feature, accounting for interactions with other variables. 32 For comparison, we also constructed a traditional logistic regression model and nomogram to serve as a benchmark to assess the robustness and interpretability of the machine learning approach. All analyses were conducted using R 4.4.2 and Python 3.9. The aplot R package was used for figure combination and visualization refinement. 33 All statistical tests were two-sided, and p values were adjusted for multiple comparisons using the Benjamini-Hochberg (BH) method, with a significance level of 0.05 unless otherwise stated.
Results
Baseline characteristics stratified by PsRS categories
Baseline characteristics of the study population stratified by PsRS categories are summarized in Table 1. A total of 5624 participants with T2DM history were included, categorized into low-risk (PsRS 0–5, n = 742), intermediate-risk (PsRS 6–10, n = 4612), and high-risk groups (PsRS 11–15, n = 270). The median baseline age was 62 years (interquartile range [IQR]: 57–66), with small but statistically significant differences among the groups (p = 0.002). Compared with the low-risk group, participants in the intermediate- and high-risk groups had higher proportions of males, current smokers, individuals with a history of stroke or CVD, insulin users, and cholesterol-lowering medication users. Significant differences were observed for smoking status (p = 0.009), alcohol consumption (p < 0.001), BMI (p < 0.001), stroke history (p = 0.018), CVD (p = 0.0001), insulin usage (p = 0.027), and cholesterol medication usage (p = 0.025). In contrast, no significant differences were found in cancer history, hypertension status, APOE ε4 carrier status, or family history of dementia across PsRS categories.
Baseline characteristics of participants stratified by PsRS categories in this study.
Continuous variables were displayed as median [IQR] and categorical variables were displayed as n (%). *p values were obtained from the analysis of variance for continuous variables and the χ2 test for categorical variables. CVD: cardiovascular disease; APOE: apolipoprotein E; IQR: interquartile range.
To further characterize the distinct risk profiles associated with dementia subtypes, baseline characteristics were stratified by clinical outcomes (Control, AD-only, VaD-only, and mixed dementia; Supplemental Table 7). Significant heterogeneity was observed among the subtypes. Notably, the VaD-only group exhibited a pronounced burden of vascular and metabolic risk factors compared to the AD-only group. Specifically, participants who developed VaD had a significantly higher baseline prevalence of stroke history (10.2% versus 2.4%), insulin-treated diabetes (44.3% versus 27.7%), and a higher median BMI (31.23 versus 30.02 kg/m2) compared to those who developed AD. Although differences in smoking status and CVD history did not reach statistical significance across all groups, the VaD-only group consistently showed the highest proportions of previous smokers (58.0%) and individuals with CVD history (37.5%). In contrast, the AD-only group was characterized by the highest proportion of APOE ε4 carriers (62.7% versus 58.0% in VaD), reflecting a stronger genetic predisposition distinct from the vascular-driven profile of VaD.
Independent associations of PRS and PsRS with the risk of three dementia subtypes
The newly constructed PsRS followed an approximately normal distribution and was positively associated with the risk of ACD (Figure 1A), AD (Figure 1B), and VaD (Figure 1C). The relationship between PsRS and ACD appeared linear, with no evidence of nonlinearity (p for nonlinearity = 0.354; p overall = 0.010). Similar linear associations were observed for AD (p for nonlinearity = 0.280; p overall = 0.364) and VaD (p for nonlinearity = 0.607; p overall = 0.001).

Relationship between polysocial risk score and incident dementia. The histograms show the distribution of PsRS, and the restricted cubic spline curves show the linear relationship between PsRS and three types of dementia risk (A: all-cause dementia; B: Alzheimer's disease; C: vascular dementia). HRs (95% CI) were estimated with adjustment for age, sex, BMI status, smoking status, alcohol consumption status, baseline history of cancer, hypertension, and CVD, cholesterol drug usage, family history of dementia, and APOE ε4 carrier status. APOE: apolipoprotein E; CI: confidence interval; CVD: cardiovascular disease; HR: hazard ratio; PsRS: polysocial risk score.
In the Cox proportional hazards models, higher PsRS was consistently associated with increased risks of ACD, AD, and VaD. In Model 1, participants with PsRS scores of 11–15 had a 2.55-fold higher risk of ACD (HR = 2.551; 95% CI, 1.451–4.485, p < 0.001) and a 3.66-fold higher risk of VaD (HR = 3.658; 95% CI, 1.225–10.928, p = 0.0202), while the association with AD was not statistically significant (HR = 2.055; 95% CI, 0.895–4.719, p = 0.090). These associations remained robust in the fully adjusted Model 2. Compared with participants with PsRS scores of 0–5, those with scores of 11–15 had a 2.74-fold higher risk of ACD (HR = 2.738; 95% CI, 1.556–4.818, p < 0.001), a 2.25-fold higher risk of AD (HR = 2.253; 95% CI, 0.979–5.186, p = 0.056), and a 3.98-fold higher risk of VaD (HR = 3.975; 95% CI, 1.328–11.898, p = 0.014). The hazard ratio per 1-point increment in PsRS was 1.125 (95% CI, 1.053–1.201, p < 0.001) for ACD and 1.285 (95% CI, 1.152–1.433, p < 0.001) for VaD, but not significant for AD (HR = 1.042; 95% CI, 0.933–1.163, p = 0.464). The population preventable fraction (PFP%) of PsRS in Model 2 was estimated at 32.5% (95% CI: 32.1%–32.9%, p < 0.001) for ACD, 2.0% (95% CI: 1.4%–2.7%, p < 0.001) for AD, and 60.7% (95% CI: 60.4%–60.9%, p < 0.001) for VaD (Table 2).
Risk of incident dementia according to polysocial risk score.
Cox proportional hazards models were used to examine the associations between the polysocial risk score (PsRS) and the risk of all-cause dementia, Alzheimer's disease, and vascular dementia. Unadjusted: included only PsRS; Model 1: adjusted for age, sex, and BMI; Model 2: adjusted for age, sex, BMI, APOE ε4 carrying status, cancer history, hypertension history, cholesterol medicine usage history, and family history of dementia. HR: hazard ratio; CI: confidence interval; PFP: preventable fraction for the population. *p < 0.05, **p < 0.01, ***p < 0.001.
Similarly, the PRS derived from 94 significant SNPs of AD also exhibited a normal distribution and showed a positive association with dementia outcomes (Figure 2A-C). A significant nonlinear association was observed between PRS and ACD (p for nonlinearity = 0.023; p overall < 0.001), where the hazard ratio increased more steeply at higher PRS levels. In contrast, the associations between PRS and AD (p for nonlinearity = 0.725; p overall < 0.001) and between PRS and VaD (p for nonlinearity = 0.056; p overall = 0.004) were linear.

Relationship between polygenic risk score and incident dementia. The histograms showed the distribution of PRS in 5,624 participants, and the restricted cubic spline curves showed the non-linear relationship between PRS and all-cause dementia risk as well as the linear relationship between PRS and Alzheimer's disease risk and vascular dementia risk. HRs (95% CI) were estimated with adjustment for age, sex, genotyping array, and the first 10 principal components of ancestry. HR: hazard ratio; CI: confidence interval; PRS: polygenic risk score.
Associations of PsRS with dementia risk across PRS strata
To assess whether the association between PsRS and ACD differed across levels of genetic susceptibility, we conducted stratified analyses by PRS tertiles (low, medium, and high; Figure 3). In the low genetic risk group, PsRS showed no significant association with ACD risk. In the medium genetic risk group, participants with a PsRS of 11–15 had a 2.026-fold higher risk of ACD (HR = 2.026; 95% CI: 0.948–4.326; p = 0.068) compared to those with scores of 0–5. The strongest association was observed in the high genetic risk group, where participants with PsRS scores of 11–15 had a significantly elevated risk of ACD (HR = 5.325; 95% CI: 1.625–17.452; p = 0.006). Even participants with scores of 6–10 showed an increased risk (HR = 2.288; 95% CI: 0.915–5.722; p = 0.077), compared to the low PsRS reference group. The interaction test between PsRS and PRS was not statistically significant (p for interaction = 0.491).

Association between PsRS categories and all-cause dementia across PRS strata. Forest plot of HRs and 95% CIs from Cox proportional hazards models assessing the association between PsRS categories (0–5, 6–10, 11–15) and incident all-cause dementia, stratified by low, medium, and high PRS groups. The model was adjusted for age, sex, BMI status, baseline history of cancer, hypertension, and CVD, cholesterol drug and insulin usage, family history of dementia, genotyping array, and the first 10 principal components of ancestry. HR: hazard ratio; CI: confidence interval; PRS: polygenic risk score; PsRS: polysocial risk score; CVD: cardiovascular disease.

Receiver operating characteristic (ROC) curves of seven machine learning models in the testing set after hyperparameter tuning. The models evaluated include Logistic Regression (LR), Support Vector Machine (SVM), Gradient Boosting Decision Tree (GBDT), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Categorical Boosting (CatBoost), and Adaptive Boosting (AdaBoost). The Area Under the Curve (AUC) values for each model are displayed.
To further examine this pattern, we combined individuals with medium and high PRS and re-evaluated the association between PsRS and dementia risk (Table 3). In the fully adjusted model, participants with PsRS scores of 11–15 had a 3.19-fold higher risk of ACD (HR = 3.193; 95% CI: 1.713–5.952; p < 0.001) compared to those with scores of 0–5. Each 1-point increase in PsRS corresponded to a 12.3% increase in ACD risk (HR = 1.123; 95% CI: 1.045–1.207; p = 0.002), and the PFP was estimated at 38.5%. Similar associations were observed for VaD (HR = 7.415; 95% CI: 1.911–28.779; p = 0.004), and for AD (HR = 2.951; 95% CI: 1.268–6.871, p = 0.012), albeit with slightly weaker effects for AD.
Risk of incident dementia according to polysocial risk score in medium/high PRS population.
Cox proportional hazards models were used to examine the associations between the polysocial risk score (PsRS) and the risk of all-cause dementia, Alzheimer's disease, and vascular dementia in the medium and high PRS population. Unadjusted: included only PsRS; Model 1: adjusted for age, sex, and BMI; Model 2: adjusted for age, sex, BMI, APOE ε4 carrying status, cancer history, hypertension history, cholesterol medicine usage history, and family history of dementia. HR: hazard ratio; CI: confidence interval; PFP: preventable fraction for the population. *p < 0.05, **p < 0.01, ***p < 0.001.
The association between employment status and all-cause dementia risk in the medium and high PRS risk populations.
Logistic Regression models were used to examine the associations between the employment status and the risk of All-cause dementia in the medium and high PRS population. Unadjusted: included only employment status; Model 1: adjusted for age, sex, and BMI; Model 2: adjusted for age, sex, BMI, APOE ε4 carrier status, cancer history, hypertension history, cholesterol medicine usage history, and family history of dementia. Participants categorized as “Volunteer” (n = 19) and “Student” (n = 5) were excluded due to small sample sizes. OR: odds ratio; CI: confidence interval. *p < 0.05, **p < 0.01, ***p < 0.001.
Predictive modelling of ACD risk in medium and high PRS population
To predict the risk of ACD among individuals with medium-to-high PRS, a variety of supervised machine learning algorithms were applied and evaluated. 5-fold cross-validation results and detailed performance comparisons across all candidate models are provided in Supplemental Tables 5 and 6. While models such as GBDT and CatBoost demonstrated high precision during the training process, and Logistic Regression yielded optimal calibration (Brier score: 0.213) on the independent test set, XGBoost exhibited the most clinically relevant performance profile. Specifically, XGBoost achieved the highest recall (0.930) and a robust F1 score (0.735) on the test set, indicating superior sensitivity in identifying true dementia cases compared to other candidates. Consequently, given its consistent discriminative power (AUC: 0.726) and ability to minimize missed diagnoses, XGBoost was selected as the final model for interpretability analysis (Figure 4). The detailed training process and model selection section are provided in the Supplemental Material.
Feature contributions and model interpretation using SHAP and ALE
SHAP and ALE analyses were subsequently applied to interpret the XGBoost model and identify key predictors of ACD risk among individuals with elevated genetic susceptibility. The most influential feature was employment status (mean |SHAP value| = 0.45), followed by APOE ε4 carrier status (0.34), history of hypertension (0.24), and educational level (0.19). In contrast, the remaining PsRS components showed minimal impact (Figure 5(a)). The SHAP dependence plots (Figure 5(b)) revealed that high-risk categories of employment status and APOE ε4 carrier status were consistently associated with positive SHAP values, indicating a stronger contribution to increased dementia risk. In contrast, low-risk categories showed negative contributions. Hypertension history and educational level demonstrated a similar trend, with SHAP values rising alongside risk severity, suggesting a dose-response relationship.

Global-level most important predictors under SHAP analysis based on the XGBoost model. (a) The bar plot displays the global feature importance for predicting ACD risk among individuals with a medium and high PRS. (b) The beeswarm plot illustrates the distribution of SHAP values across individual observations, with each dot representing one sample. Positive SHAP values suggest a higher likelihood of ACD, while negative values indicate a lower contribution. The farther a point is from the SHAP = 0 line, the greater its influence on the model prediction.
To further elucidate how the model utilized individual-level features for predictions, SHAP waterfall plots and force plots were employed to visualize case-specific contributions (Figure 6). In a representative true positive (TP) case, the predicted probability of ACD (f(x) = 0.643) was primarily driven by APOE ε4 carrier status (+0.27) and high-risk employment status (+0.17) (Figure 6A). Conversely, in a true negative (TN) case, the low predicted risk (f(x) = –1.386) was largely explained by protective factors, including the absence of hypertension (–0.65), low-risk employment status (–0.39), and non-carrier APOE status (–0.25) (Figure 6B). The corresponding SHAP force plots further illustrated how risk factors in the TP case shifted predictions toward higher risk (Figure 6C), while protective factors in the TN case pushed predictions toward lower risk (Figure 6D).

SHAP-based instance-level interpretation for two randomly selected samples. The X-axis denotes SHAP values, where E[f(x)] represents the expected model output and f(x) indicates the specific prediction for each sample. The length of each bar reflects the magnitude of the feature's contribution. (A) Waterfall plot of a random True Positive (TP) sample. (B) Waterfall plot of a random True Negative (TN) sample. (C) Force plot of a random True Positive (TP) sample. (D) Force plot of a random True Negative (TN) sample.
To complement these findings, a SHAP summary heatmap was generated to visualize feature contributions across individuals (Supplemental Figure 3). Features such as employment status, APOE ε4 carrier status, hypertension history, and educational level showed consistent, high-magnitude SHAP values across many individuals, aligning with their key role in overall risk predictions. Based on the SHAP analyses, employment status and educational level were identified as the two most critical social predictors of ACD risk. To further characterize their effects, ALE plots were generated for each variable (Supplemental Figure 4). When employment status shifted from low-risk (value = 0) to high-risk (value = 1), the ALE value increased from approximately −0.11 to +0.10. Similarly, for educational level, the ALE value rose from approximately −0.06 to +0.06 as the variable shifted from low-risk to high-risk. These linear and symmetric patterns suggest that both social disadvantage in employment and low educational attainment independently and consistently increased the risk of ACD, reinforcing their cumulative impact observed in the SHAP results.
Association between employment status and all-cause dementia risk
In the formal analysis, employment status emerged as the most influential social factor associated with ACD, as identified by SHAP. Accordingly, we further examined its independent association with ACD using multivariable logistic regression models. Among individuals with medium and high PRS, being retired was associated with a 1.52-fold increased risk of ACD (OR = 1.516; 95% CI: 1.015–2.318; p < 0.05), while those unable to work faced an even greater risk (OR = 2.065; 95% CI: 1.183–3.540; p < 0.01), compared with currently employed individuals (Table 4). These associations remained robust after adjusting for demographic, clinical, and genetic covariates in the fully adjusted Model 2.
Comparison of predictive performance between Logistic Regression and XGBoost
To evaluate the robustness and interpretability of the XGBoost model, a conventional logistic regression model was constructed as a benchmark. In the medium and high PRS population, logistic regression achieved an AUC of 0.73 on the test set, which was slightly lower than that of the XGBoost model, indicating the latter's superior discrimination ability. A nomogram was subsequently developed based on the logistic model, providing a visual tool of individual-level risk prediction (Figure 7). Among all predictors, hypertension history and employment status demonstrated the highest regression coefficients (1.119 and 0.941, respectively), further confirming their pivotal role in ACD risk estimation. These findings support the reliability of employment status as a key social risk predictor and underscore the added value of machine learning approaches in capturing complex and multifactorial risk patterns.

Nomogram based on logistic regression for model robustness assessment. The chart displays the top 10 predictors of all-cause dementia identified from the logistic regression model. Each predictor contributes to an overall risk score, which corresponds to the estimated probability shown at the bottom. A classification threshold of 0.5 was used to differentiate high- and low-risk individuals.
Sensitivity analyses
To mitigate the potential bias of reverse causality, a sensitivity analysis excluding participants diagnosed with dementia within the first two years of follow-up was conducted (Supplemental Table 8). The association between PsRS and ACD remained statistically significant and robust in the fully adjusted model (Model 2: HR per score increment = 1.125, 95% CI: 1.053–1.203; HR = 2.731, 95% CI: 1.534–4.865). Notably, the association with VaD persisted and appeared particularly strong, with the high-risk group exhibiting a four-fold increase in risk compared to the low-risk group (Model 2: HR = 4.098, 95% CI: 1.243–13.508; HR per score increment = 1.303, 95% CI: 1.165–1.458). While the association with AD maintained a positive trend, the magnitude was attenuated and did not reach statistical significance in the fully adjusted model (Model 2: HR per score increment = 1.048, 95% CI: 0.938–1.171).
Discussion
In this study, we developed a multidimensional PsRS based on 15 SDOH and a PRS constructed from 94 genome-wide significant SNPs in 5624 UK Biobank participants with T2DM. We comprehensively assessed the independent and joint associations of PsRS and PRS with dementia risk and revealed that higher social risk was consistently associated with increased incidence of ACD and VaD, even among individuals with elevated genetic susceptibility. Furthermore, we built an XGBoost prediction model for ACD in the medium and high PRS populations. SHAP and ALE analyses highlighted employment status and educational level as key social factors contributing to dementia risk in genetically vulnerable individuals.
Dementia in individuals with T2DM represents a critical public health challenge, driven by shared pathological mechanisms such as vascular dysfunction and chronic inflammation. Despite this concern, the cumulative impact of social determinants on dementia risk within this high-risk population remains poorly quantified. Prior studies have mostly examined isolated social factors, such as social isolation 34 and education attainment, 35 and findings from the Health and Retirement Study, 36 lacking a comprehensive framework to capture the complex interplay of socioeconomic disadvantages. Crucially, no previous study has constructed a multi-domain PsRS specifically for T2DM, nor has one integrated such a metric with PRS using advanced machine learning to disentangle modifiable from non-modifiable risks. 37 Our study fills this gap by validating a rigorously developed PsRS and combining it with genetic profiling. This integrated approach not only provides the first comprehensive quantification of combined polysocial and genetic risk in a prospective T2DM cohort but also offers a cost-effective, scalable tool for early risk stratification. Given the escalating global burden of dementia, as evidenced by the 7.75 million early-onset cases reported in 2021 and a persistent upward trend in the burden attributable to high fasting plasma glucose, such a framework is vital for enabling targeted preventive strategies without overburdening clinical resources. 38
Our findings are consistent with prior research demonstrating that each unit increase in PsRS is associated with an elevated risk of dementia, further validating its utility as a cross-population indicator. 13 However, our analysis revealed a critical divergence concerning the nature of this risk in T2DM patients: a robust linear relationship between PsRS and ACD risk was observed (p for nonlinearity = 0.354; p overall = 0.01), a pattern that appears to be predominantly driven by VaD. As highlighted in the baseline analysis stratified by outcomes (Supplemental Table 7), the risk of incident ACD was disproportionately concentrated in individuals with a heavy burden of vascular and metabolic comorbidities, such as higher BMI, history of stroke, and insulin dependence. These phenotypes were significantly more prevalent in the VaD subgroup than in the AD subgroup, which suggests that the heightened sensitivity to incremental PsRS exposure in T2DM likely reflects the synergistic effects of social vulnerability and accelerated vascular pathophysiology. The PsRS essentially captures the cumulative “allostatic load” of adverse social determinants, 39 which directly exacerbates metabolic instability40,41 and endothelial dysfunction,42,43 thereby fueling the vascular pathways to dementia. 44 This distinct linear relationship highlights the clinical relevance of PsRS as a sensitive, gradable tool for identifying vascular-driven cognitive decline in diabetic populations.
In the SHAP analysis of the XGBoost model, the predictive contributions of individual PsRS components varied significantly, offering novel insights into the hierarchy of risk factors. Notably, employment status (mean |SHAP| = 0.45) and educational attainment (mean |SHAP| = 0.19) emerged as dominant predictors, with employment status surprisingly surpassing the contribution of the established genetic risk factor, APOE ε4 carrier status (mean |SHAP| = 0.34). The prominence of employment status likely reflects a dual pathway of risk in the T2DM population. First, regarding the “Retired” status, workforce exit often precipitates a withdrawal from social networks and cognitive stimulation, exacerbating social isolation, which is an independent risk factor for dementia. 45 Second, and perhaps more critical in this cohort, the “Unable to work” status likely serves as a strong proxy for severe physical frailty or advanced diabetic complications. This suggests that these individuals face a “double burden” of poor baseline health and socioeconomic deprivation, driving the exceptionally high risk observed. Regarding education, its protective effect is twofold: it fosters cognitive reserve accumulated in early life, which acts as a buffer against neurodegeneration,46,47 and enhances health literacy, which is crucial for optimal diabetes self-management. 35 While our binary scoring captures the broad impact of workforce participation, future studies should further explore how occupational complexity influences this relationship.
Encouragingly, our validation of PsRS as a reliable risk indicator suggests that community-level interventions must complement standard clinical care for high-risk individuals. To effectively mitigate social risk, a multi-sectoral collaborative approach involving healthcare authorities, urban planners, and community organizations is essential. Specifically, increasing investment in social infrastructure, such as recreational centers and communal spaces, can reduce social isolation and help maintain physical and cognitive activity levels for these patients. Educational initiatives should also be stratified by age to address the specific deficits identified in our model: for younger populations, resources should focus on enhancing health literacy and early disease awareness; for older adults, efforts should prioritize disease understanding and adherence to treatment strategies. Furthermore, given the prominent protective role of employment identified in our SHAP analysis, public health policies should consider incorporating “social prescribing” programs. These programs could offer volunteering opportunities or light-load work specifically for early retirees, helping to sustain social connectedness and cognitive reserve, thereby delaying the onset of dementia in the vulnerable T2DM population.
However, this study has several limitations. First, despite the prospective design, the limited number of incident cases in high-risk subgroups may have reduced statistical power. Second, limitations regarding PsRS construction exist: the binary coding and self-reported nature of factors like employment status lack granularity. Additionally, certain SDOH domains derived from US frameworks (e.g., healthcare access) may have limited applicability in the UK due to the universal coverage of the NHS. Third, the observed sex imbalance likely reflects the epidemiological characteristics of T2DM in older adults, where susceptibility increases in women due to postmenopausal hormonal changes and differential survival rates, rather than strictly indicating selection bias. 48 Finally, the restriction to participants of European ancestry and the “healthy volunteer” bias of the UK Biobank limit the generalizability of our findings to diverse populations.
Conclusion
Our study demonstrates that a higher PsRS is associated with an increased risk of dementia in individuals with T2DM, and a notable proportion of dementia cases may be attributed to social disadvantage. We further identified employment status and educational level as key modifiable social factors influencing dementia risk among medium and high genetic susceptibility groups. These findings highlight the importance of addressing social risk in dementia prevention strategies, particularly for genetically vulnerable populations.
Supplemental Material
sj-docx-1-alz-10.1177_13872877261435200 - Supplemental material for Polygenic and polysocial risk scores in post-type 2 diabetes dementia: Risk stratification and predictive modeling in the UK Biobank cohort
Supplemental material, sj-docx-1-alz-10.1177_13872877261435200 for Polygenic and polysocial risk scores in post-type 2 diabetes dementia: Risk stratification and predictive modeling in the UK Biobank cohort by Nan He, Hongguang Wu, Nan An, Jin Bu, Yiran Li, Yaqi Dai, Yanqiu Ou and Feiying He in Journal of Alzheimer's Disease
Footnotes
Acknowledgements
This research was conducted using data from the UK Biobank Resource under Application Number 93913. The authors gratefully acknowledge the UK Biobank participants and coordinators for their valuable contributions.
Ethical considerations
All procedures were conducted in accordance with the Declaration of Helsinki. The UK Biobank study received ethical approval from the North West Multi-Center Research Ethics Committee (REC reference: 11/NW/0382) and obtained informed consent from all participants. Access to the UK Biobank resource was granted under application number 93913 for the present analysis.
Consent to participate
Informed consent was obtained from all individual participants included in the UK Biobank study.
Consent for publication
Not applicable
Author contribution(s)
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by the Guangdong special funds for science and technology innovation strategy, China (Stability support for scientific research institutions affiliated to Guangdong Province-GDCI 2024)), National Natural Science Foundation of China (Grant No.82373529), and Guangdong Provincial Medical Science and Technology Research Fund Project, China (A2024085, A2025075).
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
