Abstract
Background
High breast density is a strong risk factor for breast cancer. As such, high consistency and accuracy in breast density assessment is necessary.
Purpose
To validate our proposed deep learning (DL) model and explore its impact on radiologists on density assessments.
Material and Methods
A total of 3732 mammographic cases were collected as a validated set: 1686 cases before the implementation of the DL model and 2046 cases after the DL model. Five radiologists were divided into two groups (junior and senior groups) to assess all mammograms using either two- or four-category evaluation. Linear-weighted kappa (K) and intraclass correlation coefficient (ICC) statistics were used to analyze the consistency between radiologists before and after implementation of the DL model.
Results
The accuracy and clinical acceptance of the DL model for the junior group were 96.3% and 96.8% for two-category evaluation, and 85.6% and 89.6% for four-category evaluation, respectively. For the senior group, the accuracy and clinical acceptance were 95.5% and 98.0% for two-category evaluation, and 84.3% and 95.3% for four-category evaluation, respectively. The consistency within the junior group, the senior group, and among all radiologists improved with the help of the DL model. For two-category, their K and ICC values improved to 0.81, 0.81, and 0.80 from 0.73, 0.75, and 0.76. And for four-category, their K and ICC values improved to 0.81, 0.82, and 0.82 from 0.73, 0.79, and 0.78, respectively.
Conclusion
The DL model showed high accuracy and clinical acceptance in breast density categories. It is helpful to improve radiologists’ consistency.
Introduction
Excessive breast density has been associated with the increased risk of breast cancer and lowering the sensitivity of the mammogram (1–3). It is well-known that Chinese women in particular have denser breast tissue and tend to develop breast cancer earlier than women of other races (4,5). In order to emphasize the importance of the description of breast density categories, the 5th edition of the BI-RADS lexicon revised the categories of breast density and eliminates the quartile ranges of percentage tissue density (6). This more subjective assessment will therefore increase the variation in mammographic breast density assessments among radiologists (7,8).
Deep learning (DL) techniques based on convolutional neural networks are an emerging “computer vision” artificial intelligence technology, which has been used in diagnostic analyses for medical imaging (9). There have been several studies on the automatic classification of breast density based on DL, but only a few followed the 5th edition of BI-RADS standards (10,11). Furthermore, to the best of our knowledge, there is no similar study focusing specifically on a cohort of Chinese women (12–14). Therefore, we collected datasets that utilize the 5th edition of the BI-RADS classification standard, and also developed a DL model for mammographic density classification in a cohort of Chinese women. Due to the variance in assessment between the radiologists, the accuracy and the clinical acceptance were designed and used as evaluation metrics. The accuracy was defined as the proportion of cases that the DL model's assessment is consistent with the radiologist's assessment. Clinical acceptance was defined as the proportion of cases that the DL model's assessment is inconsistent with the radiologist's assessment, but radiologists think the DL model's diagnosis can be correct as well. We further evaluate it into the acceptance where the proportion of cases that radiologists accepted the evaluation results of the DL model.
The aims of the present study were to measure the accuracy and clinical acceptance of the DL model, and to explore the impact of the DL model on the density assessments made by clinical radiologists.
Material and Methods
The institutional review board of our university approved this retrospective study, and the requirement for informed consent was waived.
Development of the deep learning model
For the development of the DL model, 42,152 images describing breast density according to the ACR 5th edition of BI-RADS were collected from the Shenzhen People’s Hospital between 2014 and 2020, and the age range of the basic population was 32–80 years. We used a fivefold train and validation set during the training progress, and the independent testing set included 3732 cases for verifying the performance of the proposed method. The proposed DL model was implemented using a deep pyramidal residual network (PyramidResNet-101) with PyTorch version 1.1.0 (pytorch.org) and was trained using a Siamese network which shared trainable parameters of feature engineering. The mediolateral oblique views (L-MLO and R-MLO) account for both pectoral muscle and glandular backlog, which can interfere with the analysis. Therefore, a deep aggregation network for bilateral contextual information was used to integrate the extracting features from only two craniocaudal views (L-CC and R-CC). Instead of the traditional density classification methods using a single input image to a single output density value for each patient, a novel density classifier including two branches was proposed. As shown in Fig. 1, a feature extraction network followed by an attention module was designed to promote feature distinction. A spatial pyramid pooling layer was also utilized to consider both global and local features.

The framework of the proposed breast density classification model.
Patients for clinical validation
Data from consecutive screening and diagnosis mammograms were collected before and after the implementation of the DL model, which included 1738 patients from November 2018 to December 2018 (pre-DL period) and 2138 patients from January 2019 to March 2019 (post-DL period). A total of 144 cases (52 and 92 cases for the pre-period and post-period, respectively) were excluded due to prosthetic breast augmentation. Therefore, the final study cohort included 1686 pre-DL patients (mean age = 46.07 years; age range = 22–87 years) and 2046 post-DL patients (mean age = 45.54 years; age range = 18–79 years).
Mammographic evaluation of breast density
Five breast imaging radiologists were divided into two groups: the junior group (work experience <10 years [2, 3, and 5 years]), and the senior group (work experience >10 years [11 and 20 years]). All pre-DL cases (1686 cases from November 2018 to December 2018) were evaluated by these two groups of radiologists and the DL model. Each radiologist independently evaluated the mammogram images according to the 5th edition BI-RADS criteria and defined breast tissue density according to both the four-category (category a = almost entirely fatty, category b = scattered areas of fibroglandular tissue, category c = heterogeneously dense, and category d = extremely dense) and two-category (dense = including categories a and b, and non-dense = including categories c and d) assessments. All radiologists were blind to the original report as well as the outcome of the DL model. The two groups of radiologists reviewed the inconsistent cases with the DL evaluation, and the number of instances that the assessment of the DL model was accepted or not accepted in consensus was recorded. The post-DL cases (2046 cases from January 2019 to March 2019) were first tested by the DL model and the results were used as a reference to assist in the evaluation of breast tissue density by the two groups of radiologists. The evaluation of the results of each radiologist under the assistance of the model was also recorded.
Statistical analysis
Percentages were used to record the accuracy and clinical acceptance of the DL evaluation results. Statistical analyses were performed using SPSS version 22.0 (IBM Corp., Armonk, NY, USA). The Pearson's chi-square test was used to compare the results of the two groups of radiologists and the DL model evaluation, as well as the two groups of radiologists across the pre-DL and post-DL periods of assessment. A P value <0.05 was considered statistically significant. Linear-weighted kappa (K) and intraclass correlation coefficient (ICC) statistics were applied to analyze and compare the consistency between the two groups of radiologists and the model implementations made by all radiologists pre-DL and post-DL.
Results
Accuracy and clinical acceptance of the DL model
The 1686 screening mammograms in the pre-DL period were evaluated by two groups of radiologists as well as the DL model for categorization using the four- and two-category systems. The junior and senior groups of radiologists assessed 95.1% and 94.3% of cases as dense, and 4.8% and 5.7% of cases as non-dense, respectively, in the two-category assessment. The DL model assessed a similar proportion of cases as dense (95.0%) and non-dense (5.0%; P = 0.500) (Table 1).
Mammographic breast density assessment distribution before implementation of the DL model.
Values are given as n (%).
*Pearson's chi-square test.
DL, deep learning.
The DL model exhibited high accuracy. Compared with the evaluations of the junior group, the accuracy of the DL model for the four- and two-category assessments was 85.6% and 96.3%, respectively. Compared to the senior group, the accuracy of the DL model for the four- and two-category assessments was 84.3% and 95.5%, respectively.
In addition, both the junior and senior groups of radiologists had high clinical acceptance of the DL model density assessment. For the four- and two-category systems, the junior group expressed 89.6% and 96.8% acceptance, respectively, while the senior group expressed 95.3% and 98.0% acceptance, respectively (Table 2).
The accuracy and the clinical acceptance of the deep learning model in breast density assessment.
Values are given as n (%).
When the radiologists did not accept the heterogeneously dense results predicted by the DL model, the junior group evaluated 81.7% (134/164) of them as extremely dense, and the senior group evaluated 74.9% (131/175) of them as extremely dense. In comparison, only 18.3% (30/164) and 25.1% (44/175) of the cases were evaluated as scattered areas of fibroglandular density. When the radiologists did not accept the result of scattered areas of fibroglandular tissue predicted by the DL model, both groups evaluated 83.8% (31/37) of the cases as heterogeneously dense. On the contrary, only 16.2% (6/37) of the cases were evaluated as almost entirely fatty (Fig. 2).

Comparison of the junior and senior radiologist group assessments with the DL model for the (a, c) two- and (b, d) four-category mammographic breast density classification systems. Corresponding examples are provided of mammograms with concordant and discordant assessments by the radiologists and DL model. DL, deep learning.
Consistency between the radiologists’ pre- and post-dl model implementation
Breast density classifications evaluated by the two radiologist groups at the pre- and post-DL periods were compared. For the junior group, the proportion of heterogeneously dense assessments increased from 80.2% to 84.7% in the four-category assessments (P < 0.001), while the proportion of extremely dense assessments decreased from 14.9% to 11.3% (P < 0.01). For the senior group, the proportion of assessments using either four categories or two categories showed no significant differences (all P > 0.05) (Table 3).
Mammographic breast density assessment before and after implementing the DL model.
Values are given as n (%).
*Pearson's chi-square test.
DL, deep learning.
The comparison between these five radiologists’ assessment classifications of mammographic density pre- and post-DL implementation is shown in Fig. 3. The consistency within the junior group, the senior group, and among all radiologists improved after DL implementation, their K and ICC values improved to 0.81, 0.81, and 0.80 from 0.73, 0.75, and 0.76. And for four categories, their K and ICC values improved to 0.81, 0.82, and 0.82 from 0.73, 0.79, and 0.78, respectively (Table 4).

Comparison of the frequency of density cases classified by five radiologists before and after implementation of the DL model. The years of experience of radiologists 1, 2, 3, 4, and 5 are sorted from least to greatest. (a) Assessment using four categories before implementation of the DL model. (b) Assessment using four categories after implementation of the DL model. (c) Assessment using two categories before implementation of the DL model. (d) Assessment using two categories after implementation of the DL model. DL, deep learning.
Consistency between radiologists before and after implementation of the DL model.
Values in parentheses are 95% confidence intervals.
*Intraclass correlation coefficient value.
K value.
DL, deep learning.
Discussion
High mammographic density is a strong risk factor for breast cancer (1,2). As such, high consistency and accuracy in breast density assessment is necessary, especially in Asian women who have denser breast tissue compared to other races (4,5). Therefore, we constructed a dataset based on the 5th edition of the BI-RADS lexicon and trained a DL model to predict breast density. In this study, our DL model had high classification accuracy and clinical acceptance in breast density assessment categories. This model helped to improve consistency in the subjective visual assessment of breast tissue density by radiologists, which will promote risk judgment of breast cancer and provide better notification to patients in breast cancer screening.
Although there have been other studies on automatic classification of breast tissue density using DL, most of them are based on mammography images of European and American women (12–14). As we all know, the breasts of Asian women are generally denser than both of these populations. In the 1686 mammographic density classification evaluations included in this study, the two groups of radiologists with different levels of work experience assessed 95.1% and 94.3% of the cases as dense, and our DL model also assessed 95.0% as dense. Dontchos et al. reported in an external validation study of their DL model that radiologists assessed 39.3% and 42.8% of mammography images from two groups of patients as dense, while the DL model assessed only 34.9% and 34.1% as dense (14). In contrast, our DL model assessed the proportion of dense breasts more similar to that of the five participating radiologists. One reason for this may be that we proposed a novel algorithm of mammographic breast density classification using DL and integrated bilateral context information to decrease the effect of lesions in the image and improve performance. Another reason might be since the dataset used for model training consisted entirely of Chinese women with a large percentage of high-density classifications. In other words, the dense breast tissue in the dataset accounted for a more substantial proportion of the data, so the DL model was able to perform a better classification of denser breasts (15).
Clinical validation of the DL model is necessary to verify the performance of the algorithm in a real clinical environment (16). Before clinical implementation of the DL model, we conducted a study to simulate the form of clinical practice for verification of the DL model. Our DL model showed high accuracy in breast density classification using the two- and four-category assessments, reaching 96.3% and 85.6% in the junior group, respectively, and 95.5% and 84.3% in the senior group, respectively. Traditional density classification methods usually use a single image as the input for one patient. However, the large masses and abnormal regions in mammograms can disrupt proper density classification (12). As such, we proposed a novel density classification method including two branches: a single input branch for unique patients with a single breast, and a dual input branch in order to aggregate bilateral context information to reduce the perturbation of lesions, and to improve the classification performance of the model. The results of the present study suggest that the DL algorithm we designed properly classifies breast tissue density with high accuracy. Furthermore, two groups of radiologists in the present study expressed high acceptance of the DL model for the assessment of breast density, with an approval of 96.3% and 85.6% (junior group) and 95.5% and 84.3% (senior group) in the two- and four-category assessments, respectively. Overall, these results confirmed that the proposed DL model can support more widespread clinical use. It is important to note that the discordant cases evaluated by radiologists and the DL model were primarily in scattered areas of fibroglandular densities and heterogeneously dense classifications. Radiologists tend to define density as more dense assessment categories, which is consistent with previous findings (14). This indicates that the implementation of the DL model may potentially reduce the number of patients requiring supplementary screening tests and high-risk clinical evaluations, and help alleviate the psychological pressure and medical costs of female patients.
Improving the consistency of radiologists’ assessments of breast density classifications has been an essential topic in recent years (17). With the revision of the 5th edition of BI-RADS, there is more subjectivity in the qualitative assessment of breast tissue density classification by radiologists (18). Sprague et al. reported that there was considerable variation in qualitative BI-RADS density assessments across 83 radiologists, with a range of 6%–85% of mammograms being assessed as heterogeneously or extremely dense (8). The published research of DL in mammographic breast density evaluations mostly focuses on algorithm design and diagnostic performance (12–14,18). To the best of our knowledge, this is the first study to explore the integrated impact of a DL model on the consistency of radiologists’ assessments of breast tissue density classification. One similar study only measured the consistency between the evaluations of a DL model's evaluations and the radiologists’ evaluations but did not measure the changes in consistency between radiologists (13). As work experience and the number of mammograms interpreted could affect the consistency of the radiologists’ assessments (7), we divided five radiologists into either a junior or senior group based on their experience. Our findings indicate that the consistency across all radiologists improved after DL model implementation. The increase in consistency between radiologists when presenting the DL results would probably be related to a large acceptance of the AI results. Determining breast tissue category does not only impact the additional screening that follows but could also affect the possibility of false-negative mammography and breast cancer risk assessment. In addition, previous studies have reported considerable inconsistencies among radiologists, with K statistics in the range of 0.4–0.7 (7,19,20). Our findings show that the consistency among the radiologists in this work was satisfactory, with K and ICC statistics >0.7. This may be explained by the specialized training that our radiologists had all undergone in mammographic breast density classification, which was based on the 5th edition of BI-RADS.
There has been preliminary work with DL methods to assess breast density; however, none of these techniques have been implemented in clinical practice, raising questions about clinical acceptance by practicing radiologists and the effect on patient care. Therefore, it is not enough to use the DL system. Radiologists’ assessment of breast density mainly adopts subjective qualitative evaluation, while DL can automatically classify breast density with better generalization ability after being trained with a large amount of data. When radiologists cannot distinguish the two most variably assigned BI-RADS categories, i.e. “scattered density” and “heterogeneously dense,” the DL system can assist radiologists in assigning a BI-RADS category in the current clinical workflow, which will help improve the consistency of radiologists and better support consistent density notification to patients in breast cancer screening.
The present study has some limitations. First, this is a single-center study, and all the mammograms in this study were obtained from a single hospital, which is insufficient for the complete external validation of the DL model. Second, this was a retrospective study to simulate clinical practice, which cannot reflect actual clinical implementation and may lead to exaggerated estimates of the results. Third, this study only discusses the impact of the DL model on the consistency between radiologists, and the impact on the consistency within a radiologist still needs further research and analysis.
In conclusion, our DL model was developed for the classification of density in mammography, analyzing the breasts of Chinese women. We found that this DL model can accurately classify breast tissue density as it has demonstrated both high accuracy and clinical acceptance. Furthermore, the DL model can reduce the experience requirements related to subjectivity in clinical work and improve the consistency of radiologists in their breast density assessments.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship and/or publication of this article: This research was supported by the Shenzhen Science and Technology Research Fund (No. GJHZ20210705142208024).
