Abstract
Background
Deep learning (DL) has been increasingly applied to grade knee osteoarthritis (KOA) on radiographs, but reported diagnostic performance varies across Kellgren–Lawrence (K–L) grades.
Purpose
To systematically evaluate the diagnostic performance of DL models for radiographic KOA grading.
Material and Methods
PubMed, Embase, and Web of Science were searched through November 2024 for studies using DL algorithms to grade KOA on X-ray images. Sensitivity and precision were synthesized. Heterogeneity was assessed using the I2 statistic. Subgroup analyses and meta-regression were conducted according to transfer learning, external validation, multi-task learning, joint training strategy, and data splitting. Publication bias was assessed using funnel plots and Egger's test. Study quality was evaluated using the revised QUADAS-2 tool.
Results
Of 1004 records screened, 32 studies were included. Pooled sensitivity for K–L grades 0–4 was 0.90, 0.66, 0.80, 0.87, and 0.88, respectively, and pooled precision was 0.87, 0.71, 0.81, 0.86, and 0.91, respectively. Diagnostic performance was poorest for K–L grade 1, particularly in sensitivity, indicating limited reliability for early-stage KOA detection. Heterogeneity was high across outcomes and grades, particularly for sensitivity in K–L grades 1 and 2 and precision in K–L grades 0 and 1. Meta-regression identified transfer learning and data splitting as potential sources of heterogeneity. Egger's tests suggested no statistically significant small-study effects.
Conclusion
DL models showed better diagnostic performance for moderate-to-severe radiographic KOA than for early-stage disease. However, the poor sensitivity for K–L grade 1, substantial heterogeneity, and limited external validation suggest that current DL models are not yet reliable for early KOA detection or ready for routine clinical implementation. Further standardized reporting, robust validation, and multicenter external evaluation are required.
Introduction
Knee osteoarthritis (KOA) is a progressive degenerative joint disorder that significantly impairs mobility and diminishes quality of life, particularly among older adults (1). The condition is principally marked by the formation of osteophytes, narrowing of joint spaces, and deterioration of the articular cartilage within the knee (2). In its initial stages, KOA may present with few or no clinical symptoms; however, as the disease advances, patients often experience knee pain, tenderness, restricted range of motion, and swelling (3).
Radiographic imaging, particularly plain X-rays, remains the most commonly used modality for evaluating KOA severity due to its wide availability and cost-efficiency (4,5). Among the various grading systems, the Kellgren–Lawrence (K–L) scale is the most widely adopted in both clinical and research settings (6). The K–L scale provides a clinically meaningful ordinal framework for KOA severity assessment, ranging from grade 0, indicating no radiographic evidence of OA, to grade 4, indicating severe disease with marked joint space narrowing, osteophyte formation, sclerosis, and bony deformity. Distinguishing early grades, particularly K–L grades 1 and 2, is clinically important because these stages may represent subtle or early structural changes that could influence monitoring, lifestyle intervention, and early management decisions. However, these early grades are also more difficult to classify reliably because radiographic findings are often subtle and overlap with normal or borderline appearances. Nonetheless, this system has notable limitations, including substantial inter-observer variability and a strong reliance on subjective interpretation (7,8). These inconsistencies can be attributed to differences in image quality, acquisition protocols, and evaluator expertise, all of which can affect diagnostic performance and therapeutic decision-making (4).
In recent years, artificial intelligence (AI)—and deep learning (DL) in particular—has emerged as a powerful tool in medical imaging (9). As an advanced subset of machine learning, DL has demonstrated remarkable performance in numerous imaging tasks, including disease classification, detection, and segmentation (10–14). By learning hierarchical data representations from large datasets, DL models have the potential to improve diagnostic performance and reproducibility. Despite growing interest in AI-assisted KOA diagnosis, the application of DL models based on X-rays to K–L grading remains an evolving research area. Although numerous studies have demonstrated encouraging results (7,15–45), most reports have emphasized overall model accuracy or aggregate performance, whereas grade-specific diagnostic performance remains insufficiently characterized. This is particularly important because DL models may perform differently across K–L grades, with early-stage grades being more difficult to identify than moderate-to-severe disease. Therefore, the consistency of grade-specific diagnostic performance and the overall reliability of DL models warrant further investigation. Furthermore, key methodological variables, including transfer learning, external validation, multi-task learning, joint training strategies, and data splitting methods, have not yet been comprehensively examined.
Therefore, this systematic review and meta-analysis aimed to provide a comprehensive evaluation of the diagnostic performance of DL models in radiographic grading of KOA.
Material and Methods
This meta-analysis was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies (PRISMA-DTA) guidelines. In addition, the study protocol was prospectively registered before data extraction (PROSPERO registration number: CRD42025630047).
Search strategy
A comprehensive search of PubMed, Embase, and Web of Science was conducted through November 2024, utilizing core and synonymous terms related to knee osteoarthritis and DL (e.g. “KOA,” “deep learning,” “X-ray,” “Kellgren–Lawrence,” “classification,” and “diagnosis”). Boolean operators were employed to refine the search queries, and no filters were applied regarding time or language. The full search strategy is shown in Table S1. In addition, reference lists and forward citations were screened to identify further relevant studies. Reference management software (EndNote v21.5) was used to organize citations and remove duplicates.
Inclusion and exclusion criteria
The inclusion criteria were based on the PIRTOS framework. Participants (P) were patients with KOA assessed using the K–L grading system. The index test (I) was a DL model applied to X-ray images for KOA grading. The reference standard (R) was the K–L grading system. The target condition (T) was the presence of a specific K–L grade, with other grades serving as comparators. Outcomes (O) included diagnostic metrics such as sensitivity and precision, which had to be explicitly reported or derivable from the included studies. The setting (S) included retrospective or prospective studies based on public databases or hospital datasets, with a prespecified minimum overall sample size of 10 patients or knees. This broad threshold was used to avoid excluding exploratory DL studies and studies with small grade-specific subgroups after stratification by K–L grade. However, this threshold did not influence the final evidence base, because none of the included studies had an overall sample size below 100 patients or knees; the smallest included study comprised 728 knees.
The exclusion criteria were as follows: (i) studies published in languages other than English; (ii) duplicate records; (iii) non-original formats (e.g. case reports, abstracts, letters, commentaries, reviews, or meta-analyses); (iv) studies lacking complete or accessible diagnostic data; and (v) studies focused solely on image segmentation without diagnostic classification.
Retrieval of relevant articles
Two reviewers independently screened titles and abstracts, assessed full-text eligibility, extracted data, and evaluated methodological quality. Disagreements were resolved through discussion and consensus, with consultation of a third reviewer when necessary. Formal inter-rater agreement statistics, such as Cohen's kappa, were not calculated. This systematic and rigorous selection process ensured the methodological integrity of the included studies and established a dependable foundation for the subsequent meta-analysis.
Quality assessment
We assessed the methodological quality of the included studies using the revised QUADAS-2 tool. Because the index tests were DL-based prediction or classification models rather than conventional diagnostic tests, selected PROBAST-informed considerations were used to supplement the assessment of AI-specific methodological issues within the Index Test and Flow and Timing/analysis-related domains. Specifically, these considerations guided judgments regarding model development, validation strategy, potential overfitting, handling of predictors or input features, data splitting, and completeness of performance reporting. These PROBAST-informed considerations were used to adapt the quality assessment to the AI prediction-model context and did not replace the core QUADAS-2 domains. The adapted framework retained the key QUADAS-2 domains of Patient Selection, Index Test, Reference Standard, and Flow and Timing, while incorporating AI-specific considerations relevant to model validation and reporting. Two reviewers independently assessed study quality using this adapted QUADAS-2 framework supplemented by selected PROBAST-informed considerations. Disagreements were resolved through discussion and consensus, with consultation of a third reviewer when necessary. Formal inter-rater reliability statistics, such as Cohen's kappa, were not calculated.
Data extraction
Two reviewers independently extracted data from all included studies, capturing diagnostic performance data and key methodological features: publication year, country, modeling strategy, data splitting method, dataset used, DL algorithm type, and the use of transfer learning, multi-task learning, joint training, and external validation. Data splitting strategies were categorized as hold-out validation or k-fold cross-validation. K-fold cross-validation may improve the stability of internal validation by repeatedly partitioning the dataset into training and validation subsets. Modeling strategies were classified as single-stage (integrated end-to-end models) or two-stage (sequential subtask models for detection and classification). Transfer learning was defined as the use of pretrained models, such as ResNet-101 and VGG16, for fine-tuning; multi-task learning was defined as the joint training of related tasks through shared feature representations; joint training referred to the simultaneous optimization of multiple subnetworks or learning objectives within an end-to-end framework; and external validation was defined as the use of independent datasets to assess generalizability. Missing or unclear information was verified from published materials and recorded as NR when unavailable.
Outcome measures
We systematically evaluated the diagnostic performance of DL models for grading KOA based on X-ray images, with a focus on sensitivity and precision. To reduce statistical dependency and potential patient overlap among multiple models from the same study, one representative DL model was selected from each study. When multiple models were reported, we prioritized the model with the highest reported AUC; if AUC was unavailable, the model with the highest sensitivity was selected. This approach was adopted because multiple algorithms within the same study were often trained and tested on the same or highly overlapping datasets. However, we acknowledge that this selection strategy may introduce optimistic bias by favoring better-performing models. We extracted true positives (TP), false positives (FP), and false negatives (FN): TP represents correctly identified positive cases, FP denotes negative cases incorrectly classified as positive, and FN refers to positive cases that the model failed to detect. Sensitivity (TP / (TP + FN)) quantifies the model's ability to accurately identify actual positive cases, reflecting its effectiveness in minimizing missed diagnoses. Precision (TP / (TP + FP)) indicates the proportion of TPs among all positive predictions, highlighting the model's capacity to reduce FPs. Specificity, negative predictive value, positive likelihood ratio, and negative likelihood ratio were additionally derived using a grade-specific one-vs-rest framework. For each study and each K–L grade, FN was calculated as TP + FN−TP, FP as TP + FP−TP, and TN as the total number of test cases within the corresponding study across all K–L grades minus TP, FN, and FP. These metrics were interpreted as reconstructed grade-specific estimates because K–L grading is a multi-class classification task.
Statistical analysis
In this study, we applied the Freeman-Tukey double arcsine transformation to proportion data to improve distributional symmetry and mitigate bias associated with small sample sizes. Given our expectation that differences in specific model algorithms may lead to variations in true effects, we employed a random-effects model for pooling estimates of sensitivity and precision in the radiographic diagnosis of KOA. We estimated model parameters using the restricted maximum likelihood (REML) method and calculated confidence intervals using the Jackson approach. We assessed between-study statistical heterogeneity using the I2 statistic, with values of 25%, 50%, and 75% representing low, moderate, and high heterogeneity, respectively. For outcomes exhibiting substantial heterogeneity (I2 >50%), we performed subgroup analyses and meta-regression to explore potential sources of heterogeneity. When heterogeneity was extremely high, particularly when I2 exceeded 95%, pooled estimates were interpreted cautiously as broad summary trends rather than precise estimates of diagnostic accuracy. In such cases, emphasis was placed on the direction and consistency of grade-specific performance patterns, subgroup findings, and potential sources of heterogeneity rather than on the pooled values alone. To evaluate the presence of publication bias, we utilized funnel plot symmetry and Egger's regression test. We performed all statistical analyses using STATA version 18, and a P value <0.05 was considered statistically significant.
Results
Search strategy and study selection
The literature search identified 1004 records. After the removal of 499 duplicates, 505 unique articles remained for screening. Title and abstract review led to the exclusion of 457 studies due to irrelevance or non-compliance with the inclusion criteria. An additional 16 studies were excluded for lacking complete diagnostic data (i.e. TP, FP, FN). Ultimately, 32 studies (7,15–45) met the predefined inclusion criteria and were included in the final analysis to evaluate the diagnostic performance of DL models in the X-ray–based grading of KOA. Fig. 1 shows the detailed study selection process.

PRISMA flow diagram of the study selection process.
Study description and quality assessment
Fig. 2 shows that most studies had low risk of bias and low applicability concerns across domains. However, several studies presented a high risk of bias or inadequate reporting in the Patient Selection domain (e.g. references 21,23,24,29,32,33,37,41,43), primarily due to inappropriate exclusion of patients and a lack of clarity regarding data sources. In the Reference Standard domain (e.g. references 18,20,24,29,33,37,40,41,43), issues arose from using predictive factors to determine outcome status. In addition, certain studies (e.g. 16,21) deviated from predefined definitions of DL models in the Index Test and Applicability domains. Despite some studies being classified as high risk, the overall assessment revealed a predominance of low-risk studies. Overall, most studies showed acceptable methodological quality, although several domains raised concerns regarding patient selection and reference standards. Because high-risk judgments were observed across multiple adapted QUADAS-2 domains and grade-specific data were limited, excluding all studies with any high-risk judgment would have substantially reduced the available evidence base and could have produced unstable pooled estimates. Therefore, a formal sensitivity analysis excluding all high-risk studies was not performed. Instead, study-quality concerns were incorporated into the interpretation of the findings as a descriptive robustness assessment. The main grade-specific interpretation remained unchanged, showing relatively better performance for K–L grade 0 and moderate-to-severe KOA but poorer performance for early-stage disease, particularly K–L grade 1.

Risk of bias and applicability concerns of the included studies assessed using an adapted QUADAS-2 framework supplemented by selected PROBAST-informed considerations for ai-specific issues in the Index test and flow/analysis-related domains. QUADAS-2, Quality Assessment of Diagnostic Accuracy Studies 2; PROBAST, Prediction model Risk of Bias Assessment Tool.
Table 1 provides an overview of the 32 studies included in this review. Most were published within the past 7 years. All studies employed retrospective designs and focused on the radiographic grading of KOA using X-ray imaging. In total, 14 studies adopted a single-stage modeling strategy (18,19,22,23,27,30,32,33,37–39,41,44,45), one utilized a multi-stage approach (43), while the remaining studies implemented a two-stage framework. cross-validation methods, such as k-fold cross-validation, were applied in three studies (18,24,29), and one study combined the hold-out method with cross-validation (43); all other studies relied solely on the hold-out approach. Regarding data sources, several studies used hospital-based datasets (7,18,20,21,23,28,30,42,44), whereas others employed publicly available datasets, including the Osteoarthritis Initiative (OAI), the Multicenter Osteoarthritis Study (MOST), and the Mendeley Dataset IV. Table 1 summarizes key methodological features across the included studies, with particular emphasis on the use of transfer learning, multi-task learning, joint training, and external validation. Five studies did not utilize transfer learning (15,20,23,24,39), while four incorporated multi-task learning frameworks (15,18,23,30). Joint training strategies were applied in eight studies (7,15,16,18,23,30,33,42) and external validation using independent datasets was performed in four studies (7,15,38,44).
Overview of study design, reported X-ray views, imaging protocol reporting, and DL methods in the included studies.
NR indicates that the X-ray view, imaging acquisition protocol, DL architecture, dataset composition, or preprocessing procedure was not reported in the original study.
AP, anteroposterior; FF, fixed flexion; FLSE, full-limb, standing, knee-extended positions; K–L, Kellgren–Lawrence grade; KX-E, knee X-ray–extended; LAT, lateral; MOST, Multicenter Osteoarthritis Study; NR, not reported; OAI, Osteoarthritis Initiative; PA, posteroanterior; PA-FF, posteroanterior with fixed flexion; WB, weightbearing.
Sensitivity and precision of DL in K–L classification of KOA
As illustrated in Table 2 and Figures S1–S10, the pooled sensitivity estimates for each grade were as follows: K–L grade 0 = 0.90 (95% confidence interval [CI] = 0.85–0.94), grade 1 = 0.66 (95% CI = 0.54–0.78), grade 2 = 0.80 (95% CI = 0.73–0.86), grade 3 = 0.87 (95% CI = 0.81–0.92), and grade 4 = 0.88 (95% CI = 0.82–0.92). The pooled precision estimates for each grade were as follows: K–L grade 0 = 0.87 (95% CI = 0.82–0.92), grade 1 = 0.71 (95% CI = 0.61–0.81), grade 2 = 0.81 (95% CI = 0.74–0.87), grade 3 = 0.86 (95% CI = 0.81–0.90), and grade 4 = 0.91 (95% CI = 0.86–0.95). Among all grades, pooled sensitivity showed a U-shaped pattern, with higher sensitivity for K–L grade 0 and advanced grades 3–4 but lower sensitivity for grades 1–2. K–L grade 1 showed the lowest pooled sensitivity and precision. This pattern suggests that DL models may more reliably distinguish normal knees and advanced structural disease than subtle early or mild radiographic KOA.
Pooled sensitivity and precision of DL models for radiographic KOA grading by K–L classification.
CI, confidence interval; DL, deep learning; K–L, Kellgren–Lawrence grade; KOA, knee osteoarthritis.
Additional grade-specific diagnostic metrics were reconstructed using a one-vs-rest framework based on the extracted TP, TP + FN, and TP + FP values. As shown in Table S3 (see supplementary material), specificity was high across all K–L grades, ranging from 0.877 for K–L grade 0 to 0.993 for K–L grade 4. NPV ranged from 0.909 for K–L grade 1 to 0.992 for K–L grade 4. Positive likelihood ratios were higher for advanced disease, particularly K–L grades 3 and 4, whereas K–L grade 1 showed a relatively higher LR−, consistent with its lower sensitivity and limited reliability for ruling out early-stage KOA. However, LR + estimates for advanced grades should be interpreted cautiously because very high reconstructed specificity and sparse FP counts may inflate likelihood ratio estimates in a one-vs-rest framework. These reconstructed metrics should therefore be interpreted as supplementary estimates rather than definitive binary diagnostic accuracy measures.
Heterogeneity in sensitivity and precision of DL for K–L classification
As shown in Table 2 and Table S4–S8 in the supplementary material, substantial heterogeneity in diagnostic performance was observed across different K–L grades. For sensitivity, I2 values were in the range of 93.38%–99.25% (all P <0.001), with the greatest heterogeneity noted in K–L grades 1 and 2 (I2 = 99.25%). Similarly, substantial heterogeneity was identified for precision, with I2 values in the range of 93.09%–98.91% (all P <0.001), and particularly pronounced in K–L grades 0 and 1, both nearing the upper bound (both I2 >98%).
To investigate potential sources of heterogeneity, subgroup analyses and meta-regression analysis were conducted. In the meta-regression analyses, subgroup-specific pooled sensitivity and precision estimates with 95% CIs were reported, and absolute between-subgroup differences (Δ) were also calculated to clarify the magnitude and direction of covariate effects across K–L levels. For K–L 0 classification, studies without transfer learning showed lower sensitivity and precision than those using transfer learning, with Δ values of −0.18 and −0.11, respectively. Transfer learning was also associated with higher precision for K–L 1 and K–L 2 classifications. Specifically, studies without transfer learning showed lower precision than those using transfer learning for K–L 1 (Δ = −0.24; P = 0.04) and K–L 2 (Δ = −0.29; P = 0.01). Most other covariates showed small or statistically non-significant absolute differences, suggesting that these subgroup findings should be interpreted cautiously. The results indicated that transfer learning significantly contributed to heterogeneity in K–L grade 0 sensitivity (P = 0.05), as well as in the precision of K–L grades 0, 1, and 2 (all P <0.05). In addition, the data splitting strategy emerged as a significant source of heterogeneity for K–L grade 1 precision (P <0.05), and potentially for K–L grade 4 precision (P = 0.05). Cross-validation yielded higher precision than hold-out validation for K–L grade 1 (0.83 [95% CI = 0.81–0.86] vs. 0.71 [95% CI = 0.59–0.81]; P = 0.02) and K–L grade 4 (0.97 [95% CI = 0.91–1.00] vs. 0.90 [95% CI = 0.85–0.95]; P = 0.05). However, these results should be interpreted cautiously because only three studies used cross-validation. Subgroup and meta-regression analyses stratified by external validation status showed no statistically significant differences in sensitivity or precision across K–L grades. However, only four studies included external validation; therefore, these findings should be interpreted cautiously because of limited statistical power. Other methodological variables—including modeling strategy (e.g. single-stage vs. two-stage), external validation, multi-task learning, and joint training—were not significantly associated with heterogeneity in either sensitivity or precision (all P >0.05). Additional covariates, including dataset size, single-center versus multicenter design, OAI versus hospital-based data sources, geographic region, and year of publication, were considered. However, these variables were inconsistently reported, unevenly distributed across subgroups, or had sparse category counts after stratification by K–L grade, limiting their suitability for reliable additional meta-regression analyses.
Assessment of publication bias in DL models for K–L classification
As shown in Figures S11–S20 in the supplementary material, the funnel plots revealed a generally symmetrical distribution of effect sizes, with most studies falling within the pseudo 95% CIs. This visual symmetry suggested no apparent evidence of funnel plot asymmetry. In alignment with this observation, Egger's regression tests returned non-significant results for all comparisons (P values in the range of 0.23–0.97), indicating no statistically significant small-study effects. However, these analyses cannot exclude selective outcome reporting within individual studies.
Discussion
DL algorithms showed grade-dependent diagnostic performance rather than uniformly high performance across all K–L grades. Specifically, pooled sensitivity showed a U-shaped pattern, with higher sensitivity for K–L grade 0 and advanced grades 3–4, but lower sensitivity for K–L grades 1–2, particularly K–L grade 1. Clinically, this suggests that current DL models may be better at distinguishing normal knees and moderate-to-severe structural KOA, whereas their ability to detect subtle early-stage radiographic changes remains limited. This pattern is consistent with earlier research indicating that more apparent degenerative changes, such as marked joint space narrowing, osteophyte formation, and subchondral bone sclerosis, provide more reliable imaging cues for model discrimination (46–48). By contrast, early-stage KOA, especially K–L grade 1, is characterized by doubtful or subtle radiographic findings, including mild osteophyte formation and minimal joint space narrowing, which may partly explain the reduced sensitivity observed at this grade (49–51).
Radiographic K–L grading is known to be affected by inter-observer variability, particularly for early or borderline grades. In this context, the relatively low sensitivity observed for K–L grade 1 suggests that DL models may encounter similar challenges to human readers when identifying subtle early radiographic changes. However, most included studies did not perform head-to-head comparisons between DL models and radiologists or report formal reader-agreement metrics, such as Cohen's kappa. Therefore, a direct quantitative comparison between pooled DL performance and published radiologist inter-observer variability was not feasible. Future studies should incorporate reader-performance designs, ideally multi-reader multi-case studies, to directly compare DL systems with radiologists under clinically relevant conditions and clarify whether these models outperform, match, or complement human experts in clinical practice.
The poor performance for K–L grade 1 has important clinical implications. Because K–L grade 1 represents doubtful or very early radiographic OA, low sensitivity at this stage may lead to FN classification, delayed monitoring, and missed opportunities for early lifestyle modification, symptom management, or preventive intervention. Conversely, FP or overestimated grading may result in unnecessary follow-up examinations, increased patient anxiety, and potential overtreatment. Therefore, the clinical value of DL-assisted K–L grading should be judged not only by overall accuracy but also by clinically meaningful error patterns, particularly around the distinction between normal knees and early KOA. To address this limitation, future studies should improve early-stage KOA detection through standardized radiographic acquisition, harmonized K–L labeling, integration of complementary clinical or imaging information, and clinically oriented validation.
Existing reviews, such as those by Akila et al. (52) and Tayyaba Tariq et al. (53) have summarized DL methods for automated K–L grading using X-rays. These studies covered model architecture, datasets, and classification strategies. However, they provided mainly qualitative descriptions. Key diagnostic metrics, such as sensitivity and precision, were not quantitatively synthesized, and potential sources of performance variation were not systematically explored. Earlier meta-analyses also had several methodological limitations. Mohammadi et al. (54) focused on binary classification and reported strong external performance (94% sensitivity, 91% specificity), but did not address K–L grades. Zhao et al. (55) assessed sensitivity across K–L 0–4 and reported higher diagnostic performance for K–L grade 4 (90.32%), but lower results for K–L 1–2. They also used sensitivity alone as the evaluation metric. In contrast, our study provided a more comprehensive grade-specific assessment by synthesizing both sensitivity and precision. Importantly, by synthesizing both sensitivity and precision, this meta-analysis provides a more clinically relevant assessment of DL model performance. We also used subgroup and meta-regression analyses to identify how factors like transfer learning and data splitting influence performance. These analyses improve the methodological depth, transparency, and clinical relevance of our findings.
The I2 statistics indicated pronounced heterogeneity in both sensitivity and precision across K–L grades, with I2 values in the range of 93.09%–99.25%. Although subgroup and meta-regression analyses identified transfer learning and data splitting strategies as potential contributors to heterogeneity, substantial residual heterogeneity remained. This residual heterogeneity may reflect inconsistently reported factors, including dataset source and size, acquisition protocols, labeling standards, validation design, and patient characteristics. Dataset source may be particularly important, as the included studies used diverse data sources, including public datasets such as the Osteoarthritis Initiative and the Multicenter Osteoarthritis Study, hospital-based datasets, and mixed datasets combining public and local clinical data. These datasets may differ in patient characteristics, image acquisition protocols, image quality, labeling procedures, disease severity distribution, and validation design. In addition, potential overlap of OAI-derived images or patients across studies may have violated the independence assumption and artificially narrowed confidence intervals. Because several studies used overlapping or mixed datasets and did not consistently report dataset-specific grade-level diagnostic outcomes, formal dataset-stratified pooled analysis was not feasible. Therefore, the pooled estimates should be interpreted as broad performance trends rather than precise measures of clinical diagnostic accuracy. Nevertheless, quantitative synthesis remains informative because it summarizes the overall direction and grade-dependent pattern of DL model performance, rather than providing a single definitive estimate of diagnostic accuracy.
Transfer learning is widely used in medical image analysis. In the included studies, several DL models initialized network weights using architectures pretrained on large natural image datasets, such as ImageNet. This approach enhances training efficiency and improves model performance, particularly in medical imaging scenarios with limited sample sizes (56–60). However, transfer learning from natural image datasets may be affected by domain shift. ImageNet contains natural RGB images, whereas knee radiographs are grayscale medical images with disease-relevant information represented by subtle anatomical structures, bone texture, osteophytes, and joint space changes. Therefore, features learned from natural images may not fully capture the radiographic patterns required for early KOA grading. Medical image pretraining, particularly using large-scale radiographic datasets, musculoskeletal images, or self-supervised learning on unlabeled X-ray images, may provide more domain-relevant feature representations and improve model generalizability. Such strategies may be especially useful for detecting subtle early-stage changes in K–L grades 1 and 2, where current models showed reduced sensitivity. Furthermore, the findings underscore the critical influence of data splitting strategies on diagnostic consistency. Cross-validation techniques, notably k-fold cross-validation, may provide more robust internal performance estimates than single hold-out methods by reducing performance fluctuations caused by random data splits. By repeatedly evaluating models across different training and testing partitions, cross-validation provides a more stable estimate of model performance and generalizability (61,62). In comparison, hold-out strategies are more susceptible to biases related to sample representativeness, and inadequate separation between training and testing sets may further lead to either overestimation or underestimation of the model's true diagnostic performance (62). Thus, repeated or stratified k-fold cross-validation may provide a more stable internal estimate of model performance, but it should not be regarded as a substitute for independent external validation.
In recent years, multi-task learning, joint training, and multimodal diagnostic frameworks have attracted increasing attention in medical image analysis (63–65). Multi-task learning may improve grading performance by jointly learning related tasks, such as classification, localization, or segmentation, through shared feature representations (63,64), whereas joint training may enhance feature representation and model generalizability through coordinated optimization of multiple models or learning objectives (65). Although few included studies explicitly adopted or reported multi-task or joint training strategies, and these factors were not identified as statistically significant sources of heterogeneity in the present analysis (15,18,23,30), this finding should be interpreted cautiously because the small number of relevant studies may have limited the statistical power to detect their effects. These strategies may still have potential value for improving model robustness and detecting subtle early-stage KOA. In addition, integrated diagnostic frameworks incorporating complementary information, such as MRI findings, biochemical indicators, and clinical manifestations, may further support early-stage KOA assessment, particularly given current concerns regarding dataset diversity, overfitting, interpretability, and validation rigor (20,29,35,39,42,45). However, these approaches require further evaluation in studies with clearer reporting of model design, training procedures, validation methods, and grade-specific outcomes.
Several limitations should be acknowledged. First, selecting one representative model from each study, particularly when prioritizing the model with the highest reported AUC or sensitivity, may have introduced optimistic bias and overestimated the real-world diagnostic performance of DL models. Although grade-specific diagnostic data were available for the selected representative models, most studies did not provide complete grade-specific diagnostic data for all reported algorithms. Therefore, formal sensitivity analyses based on all algorithms or study-level average performance could not be conducted without introducing additional assumptions or within-study dependency. In addition, although funnel plots and Egger's tests did not suggest statistically significant small-study effects, these methods cannot exclude selective outcome reporting within individual studies, such as preferential reporting of better-performing algorithms, selected K–L grades, or incomplete diagnostic metrics. Such selective reporting may have introduced reporting bias, limited the completeness of grade-specific evidence, and affected the interpretation of pooled diagnostic performance. Moreover, we did not perform a formal sensitivity analysis excluding all studies with high-risk judgments because high-risk assessments were distributed across multiple methodological domains and excluding these studies would have substantially reduced the already limited grade-specific evidence base. Instead, study-quality concerns were incorporated into the interpretation of the findings as a descriptive robustness assessment, and the pooled estimates were interpreted cautiously in light of potential study-quality bias. Second, substantial methodological heterogeneity was present across the included studies. Although random-effects models, subgroup analyses, and meta-regression analyses were used, the extremely high I2 values indicate considerable between-study variability. Several potential sources of heterogeneity, including dataset source and size, single-center versus multicenter design, patient characteristics, labeling standards, validation design, and reference standard variability, were incompletely or inconsistently reported and therefore could not be fully explored. Although no included study had an overall sample size below 100 patients or knees, some grade-specific testing subsets were relatively small after stratification by K–L grade, which may have led to unstable diagnostic estimates and increased uncertainty in grade-specific performance. Potential overlap of OAI-derived data across studies may also have affected the independence of pooled estimates. Third, reporting of imaging protocols and key methodological details was incomplete and heterogeneous. Although available X-ray view information was summarized, detailed acquisition parameters, including patient positioning, weightbearing status, knee flexion angle, beam angle, acquisition distance, device settings, and image preprocessing procedures, were often unavailable. In addition, several studies did not fully report CNN architecture, training dataset composition, or preprocessing procedures. Such incomplete reporting may limit reproducibility, contribute to between-study heterogeneity, increase uncertainty when interpreting pooled estimates, and restrict further subgroup or meta-regression analyses. Fourth, although specificity, negative predictive value, positive likelihood ratio, and negative likelihood ratio were reconstructed using a grade-specific one-vs-rest framework, these supplementary estimates should be interpreted cautiously. Because K–L grading is a multi-class classification task and some studies reported incomplete or slightly inconsistent grade-specific diagnostic counts, these estimates may be affected by reporting inconsistency, sparse FP counts, and reconstruction-related uncertainty. Fifth, the evidence base for clinical translation remains limited. Only four of the 32 included studies performed external validation, whereas most relied on internal validation strategies, including hold-out validation or cross-validation. This limits conclusions regarding model robustness and real-world generalizability across populations, imaging protocols, clinical settings, and healthcare environments. In addition, few studies performed head-to-head comparisons between DL models and radiologists or reported formal inter-observer agreement metrics, limiting our ability to determine whether these models outperform, match, or complement radiologists in clinical practice. Finally, formal inter-rater reliability statistics, such as Cohen's kappa, were not calculated for study selection, data extraction, or methodological quality assessment, although disagreements were resolved through discussion and consensus, with consultation of a third reviewer when necessary. In addition, this review was limited to English-language publications, which may have led to missed relevant studies.
Beyond diagnostic performance, several practical requirements must be addressed before DL-assisted KOA grading can be implemented clinically. These include regulatory approval, integration into routine radiology workflows, clinician training, continuous model monitoring, and periodic updating or recalibration to ensure safety, reliability, and sustained real-world performance. In addition, cost-effectiveness should be evaluated by weighing potential benefits, such as reduced radiologists’ workload, improved reporting consistency, earlier diagnosis, more appropriate monitoring, and reduced unnecessary imaging or referrals, against the costs of software deployment, workflow integration, maintenance, regulatory compliance, training, and model updating. Future studies should also assess real-world clinical impact, health-economic value, temporal validity, and equity using diverse multicenter cohorts, temporally independent datasets, periodic recalibration, and subgroup performance reporting.
In conclusion, DL models show promising performance for radiographic grading of moderate-to-severe knee osteoarthritis. However, their limited sensitivity for early-stage KOA, particularly K–L grade 1, indicates that current models remain insufficient for reliable early detection. Given the substantial heterogeneity and methodological variability across studies, these findings should be interpreted as broad performance trends rather than definitive estimates of real-world clinical diagnostic accuracy. Future studies should prioritize standardized reporting, complete diagnostic metric reporting, robust external validation, radiologist benchmark comparisons, and clinically oriented evaluation before routine implementation.
Supplemental Material
sj-docx-1-acr-10.1177_02841851261462501 - Supplemental material for Diagnostic performance of deep-learning algorithms in radiographic grading of knee osteoarthritis: a systematic review and meta-analysis
Supplemental material, sj-docx-1-acr-10.1177_02841851261462501 for Diagnostic performance of deep-learning algorithms in radiographic grading of knee osteoarthritis: a systematic review and meta-analysis by Xiaolu Ren, Xinyu Jin, Zhi Jun Wang and Ting Li in Acta Radiologica
Supplemental Material
sj-docx-2-acr-10.1177_02841851261462501 - Supplemental material for Diagnostic performance of deep-learning algorithms in radiographic grading of knee osteoarthritis: a systematic review and meta-analysis
Supplemental material, sj-docx-2-acr-10.1177_02841851261462501 for Diagnostic performance of deep-learning algorithms in radiographic grading of knee osteoarthritis: a systematic review and meta-analysis by Xiaolu Ren, Xinyu Jin, Zhi Jun Wang and Ting Li in Acta Radiologica
Supplemental Material
sj-docx-3-acr-10.1177_02841851261462501 - Supplemental material for Diagnostic performance of deep-learning algorithms in radiographic grading of knee osteoarthritis: a systematic review and meta-analysis
Supplemental material, sj-docx-3-acr-10.1177_02841851261462501 for Diagnostic performance of deep-learning algorithms in radiographic grading of knee osteoarthritis: a systematic review and meta-analysis by Xiaolu Ren, Xinyu Jin, Zhi Jun Wang and Ting Li in Acta Radiologica
Footnotes
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the University-level Key Project of Ningxia Medical University (grant no. XZ2024034).
Supplementary material
Supplementary material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
