Abstract
Background
Volumetric breast density analysis is useful for quantitative mammographic assessment. However, there are few studies about clinical–radiologic factors contributing to discrepancies in the visual assessment by radiologists.
Purpose
To compare automated volumetric breast density measurement with BI-RADS breast density category by radiologists' visual assessments and to evaluate the clinical–radiologic factors affecting disagreement between two estimations.
Material and Methods
From February 2011 to September 2012, 860 patients (mean age, 54.7 ± 10.2 years) who had undergone digital mammography including fully automated volumetric breast density analysis, were enrolled. The agreement in breast density assessments between two radiologists, and between an experienced radiologist and the automated software were evaluated using a weighted kappa (k) value. Clinical–radiologic factors contributing to disagreement between the results obtained by a radiologist and the automated software were evaluated using univariate and multivariate analysis.
Results
Breast density assessments obtained by two different radiologists were in good agreement (weighted k statistics 0.835%; 95% confidence interval [CI], 0.8098–0.8608); breast density assessments obtained by an experienced radiologist versus automated software were in moderate agreement (weighted k statistics 0.799%; 95% CI, 0.7708–0.8263). Univariate analysis identified a difference in bilateral breast density and patient age as two factors that significantly contributed to disagreement between the two approaches (P = 0.0002, P = 0.019). Multivariate analysis only identified a difference in bilateral breast density as a contributing factor.
Conclusion
The automated volumetric breast density measurement showed good agreement with radiologists' assessment. The difference in bilateral breast density affected the disagreement between results from visual assessment and automated software.
Keywords
Introduction
Breast density in mammography depends on the relative amounts of fibroglandular and fatty tissues. Breast density is defined as the percentage of dense fibroglandular tissue in the entire breast area. Several factors can affect the composition of breast tissue (1). For example, the change and variability in breast composition depends on hormonal fluctuations including menarche, pregnancy, breastfeeding, or menopause in addition to genetic predisposition (2).
Mammographic density is important. First, dense breast fibroglandular tissue may obscure masses or calcification, thus lowering the sensitivity of mammography for detecting breast carcinomas (3,4). Second, mammographic parenchymal patterns have been shown to be associated with an increased risk of breast cancer; previous studies have also demonstrated a direct association between increased mammographic density and an increased risk of developing breast cancer (3,5–7).
Previous breast density assessments performed using the Breast Imaging Reporting and Data System (BI-RADS) (8) were based on subjective description, which has suboptimal reproducibility. Thus, several methods have been proposed for measuring mammographic density in a quantitative manner. The first such method, proposed by Wolfe et al. in 1987 (9) involved manual tracing of the dense white areas on mammograms, but was extremely time-consuming. More recently, computerized programs have been developed with the same goal, including fully automated programs as well as semi-automated programs (3,4).
However, few reports (4,10,11) to date have investigated whether breast densities obtained by radiologists correlate with those obtained by fully automated volumetric measurements. Therefore, the purpose of this study was to compare automated volumetric breast density measurement with BI-RADS breast density category as visually assessed by radiologists, and to analyze the clinical–radiologic factors contributing to discrepancies in the results obtained by these two approaches.
Material and Methods
Study population
This retrospective study was approved by the Institutional Review Board of our institution, and written informed consent was waived. Between February 2011 and September 2012, 877 women underwent screening or diagnostic mammograms, including fully automated volumetric breast density analysis. Seventeen patients with breast cancer and a history of breast-conserving surgery or mastectomy were excluded. Finally, 860 women (mean age, 54.7 ± 10.2 years; age range, 26–89 years) were included in this study. Out of all the patients, 760 had undergone mammography for screening purposes, whereas the remaining 100 patients had undergone mammography for diagnostic purposes. Routine craniocaudal (CC) and mediolateral oblique (MLO) views of films were obtained for each breast using dedicated full-field digital mammography (Senographe DS; GE Healthcare, Milwaukee, WI, US).
Mammographic evaluation by radiologists
One board-certified radiologist (SYM) with 6 years of experience in breast imaging and a third-grade radiology resident (LHN) independently reviewed mammographic densities using the BI-RADS breast density category 4 (8): breast tissue is less than 25% glandular (category 1), approximately 25–50% glandular (category 2), approximately 50–75% glandular (category 3), or more than 75% glandular (category 4). Additionally, a breast-dedicated radiologist (SYM) evaluated BI-RADS final assessment, based on the relevant mammographic findings. All radiologists were blinded to previous results and to patient information.
We divided the mammographic Volpara density grade (VDG) breast densities, obtained by Volpara, into two groups, fatty and dense to simplify the breast density category, similar to previous other studies (12,13). These groups were used to evaluate the associations with the mammographic BI-RADS final assessments. The fatty group included VDG categories 1 and 2, whereas the dense group included VDG categories 3 and 4.
After evaluation of the mammographic images of all patients, we reclassified agreement and disagreement groups according to agreement between BI-RADS density categories assessed by an experienced radiologist and VDG. We also evaluated the clinical–radiologic factors that could have affected any differences between the agreement and disagreement groups. The clinical–radiologic factors examined included patient age, indication for mammography (diagnostic or screening, according to the reason for performing the mammography), mammographic BI-RADS final assessment, and a difference in bilateral breast density as assessed by VDG categorization. We defined a difference in bilateral breast density as a situation in which the right and left breast exhibited not less than one grade difference in VDG category. We also defined the mammographic BI-RADS final assessment categories 0, 3, 4, 5, and 6 as positive findings and categories 1 and 2 as negative findings.
Mammographic evaluation by fully automated volumetric breast density measurements
For each mammographic image, we assessed volumetric breast density with commercially available software (Volpara®, version 1.5.1, Matakina Technology, Wellington, New Zealand). Volpara is a fully automated, volumetric breast density assessment software that provides volumetric breast density (VBD) and VDG. Volpara software can work out the mapping from each brightness in the image to a thickness of fibroglandular tissue and thickness of fat that must have been present between the pixel and X-ray source. Summing those across the image calculates the total volume of fibroglandular tissue, and dividing that by the volume of the breast (found from the compressed breast thickness and projected area) represents the volumetric breast density. Breast density results are provided per breast, obtained by averaging the craniocaudal and mediolateral oblique values. VDG categories were assigned automatically according to the relevant VBD values. A VBD value of 0–4.7% corresponds to VDG 1, 4.8–7.9% to VDG 2, 8.0–15.0% to VDG 3, and more than 15.1% to VDG 4 (14).
Statistical analysis
Inter-observer agreement between breast densities obtained by two different radiologists, and agreement between the BI-RADS density categories obtained by an experienced radiologist and VDG were assessed using weighted kappa (k) statistics with linear weighting. The strengths of agreement were expressed in k values: a value of 0.20 or less indicated poor; 0.21–0.40, fair; 0.41–0.60, moderate; 0.61–0.80, good; and 0.81–1.00, very good agreement (15). The association between automated volumetric breast densities, which were reclassified into fatty or dense, and mammographic BI-RADS final assessments, was analyzed with Fisher's exact test.
The clinical–radiologic factors that can affect differences between agreement and disagreement groups were evaluated by univariate and multivariate analysis. For univariate analysis, we performed t-tests for continuous variable such as age, and chi-square or Fisher's exact tests for non-continuous variables, that included indication for mammography (diagnostic or screenings), mammographic BI-RADS final assessment and a difference in bilateral breast density as assessed by VDG. We also performed multivariate analysis to estimate the odds ratios and 95% confidence intervals (CI). We performed multivariate analysis by including all variables that used in univariate analyses. The automated volumetric breast density measurement (VDG) was regarded as the reference standard. Spearman's correlation coefficient (ñ) was used to evaluate the correlation between the BI-RADS density category and the volumetric breast density, as assessed by the fully automated software. SAS (version 9.2, SAS Institute Inc., Cary, NC, USA) and Microsoft Excel (Redmond, WA, USA), were used to perform all statistical analyses. P values less than 0.05 were considered significant.
Results
Breast density measurements: radiologists versus automated volumetric measurements
The frequency of breast density categories, as assessed by two radiologists and automated volumetric breast density measurements, are listed in Table 1. Category 3 was the most frequently observed breast density according to both radiologists and Volpara (Table 2). Inter-observer agreement for breast density evaluation by BI-RADS density category was good (weighted k value = 0.835; 95% CI, 0.8098–0.8608), whereas the agreement between the BI-RADS density category according to a radiologist specialized in breast imaging (SYM) and the Volpara density grade was moderate (weighted k value = 0.799; 95% CI, 0.7708–0.8263). A significant positive correlation was observed between the BI-RADS density category, as assessed by an experienced radiologist, and the volumetric breast density as assessed by fully automated software (ñ = 0.8566, P < 0.0001) (Fig. 1).
Box and whisker plot shows a significant association between BI-RADS density categories, assigned by an experienced radiologist and volumetric breast density (VBD) assessed by a fully automated volumetric method using Volpara software (Spearman's ñ = 0.8566, P < 0.0001). Box indicates interquartile ranges (IQR). Thick line = median, O = outliers, * = extreme outliers. Frequency of breast density categories as assessed by radiologist 1, radiologist 2, and Volpara. Numbers in parentheses are percentages. Frequency of VDGs and BI-RADS density categories assessed by an experienced radiologist.
Volpara breast density that was recategorized as fatty or dense group showed significant association with the mammographic BI-RADS final assessment (P < 0.0001).
Univariate analysis
In total, there was agreement between the breast density category as determined by an experienced radiologist and the Volpara software in 688 cases. The remaining 172 cases were categorized in the disagreement group (Fig. 2a and b). An example of the results from the Volpara breast density analysis is shown in Fig. 2c. Table 3 shows the results from univariate analysis regarding five clinical–radiologic variables and their contribution to the agreement and disagreement groups. Age and a difference in bilateral breast density were significantly different between the agreement and disagreement groups. Specifically, the patients in the agreement group were younger than those in the disagreement group (54.3 ± 9.9 years vs. 56.3 ± 11.2 years, P = 0.0191). Patients with equal bilateral breast densities also exhibited greater agreement than those with different bilateral breast densities (82.9% vs. 17.1%, P = 0.0002). On the other hand, neither mammography indication (diagnostic or screening) nor breast density as categorized by Volpara (dense or fatty) was significantly different between the agreement and disagreement groups (P = 0.7903 and P = 0.0990, respectively).
CC (a) and MLO (b) views of mammography that showed disagreement with BI-RADS breast density category. Automated breast density (c) assigned grade 3, but an experienced radiologist read the mammography as grade 4. Clinical–radiologic characteristics between agreement and disagreement groups. SD, standard deviation.
Multivariate analysis
Multivariate logistic regression analysis for characteristics of 860 women by agreement and disagreement groups.
CI, confidence interval.
Discussion
Breast density is well established as an independent risk factor for breast cancer with a four-fold higher risk of breast cancer in women with dense breasts compared to women with lower breast densities (16,17). Because visual assessment of breast density according to BI-RADS density category is a subjective method and thus difficult to reproduce, various quantitative assessments of breast density have been proposed (18–20). One of these quantitative methods for breast density assessment is an area-based, computerized approach that can be either semi-automated or fully automated (18,19). However, the software used in this approach has some limitations regarding susceptibility of change according to the processing condition. On the other hand, recent technique of fully automated volumetric assessment of breast density demonstrated absolute reproducibility (10).
In the present study, we intended to compare automated volumetric breast density measurement with BI-RADS breast density category as assessed by radiologists, and moderate agreement with positive correlation was observed between the two approaches. Several previous studies (4,10,11,20,21) also reported moderate or good agreement between them, which was similar to the results from our study. Therefore, we suggest that fully automated volumetric breast density assessment is a potential substitute for visual assessment by a radiologist.
This study also investigated which clinical–radiologic factors contributed to the differences observed between the BI-RADS density categories and VDG. The disagreement group showed a significantly higher proportion of difference in bilateral breast density compared with the agreement group. However, this density difference was only a discrepancy of one grade. We assumed that bilaterally different breast density could be a hindrance that could prevent exact visual assessment of breast density. Whereas, the Volpara software may provide exact final VDG by averaging each breast density. Therefore, the use of automated software may help provide higher reproducibility, especially, for patients with different bilateral breast densities. Furthermore, a previous study by Zheng (22) demonstrated that bilateral mammographic density asymmetry is an important risk factor in predicting the likelihood of individual women developing high-risk breast abnormalities or cancer. Volpara software can provide bilateral breast densities separately, which may solve this problem.
Among the clinical–radiologic factors considered to potentially influence disagreement between the two approaches, we found there was significantly less possibility of disagreement, when visual assessment by a radiologist was category 3 and Volpara density assessed by Volpara was VDG 4. In a sense, our results are relevant to a previous study by Seo et al. (4) which revealed that the average breast density of the agreement group was significantly higher than that of the disagreement group. The authors explained that it may be difficult to evaluate scattered small amounts of tissue volume by visual assessment.
In the present study, age was not a significant independent factor by multivariate analysis, even though it did significantly contribute to differences between the agreement and disagreement groups in univariate analysis. In previous studies, a significant negative correlation between age and breast density has been well documented, consistent with the known transition of dense fibroglandular tissue to fatty tissue during aging (23). Therefore, age and agreement in breast density might have no relationship according to multivariate analysis. Seo et al. (4) also reported there was no difference between the agreement and disagreement groups with regard to age.
The present study provides an assessment of the associations between breast density of VDG and mammographic BI-RADS final assessment. To simplify density categorization, we recategorized breast density into fatty and dense groups according to VDG, and a significant association was observed with P value less than 0.0001. These results were comparable to those of many previous studies that have demonstrated significant positive associations between mammographic breast density and breast cancer risk (24).
In our study, inter-observer variability between two radiologists in the assessment of breast density showed good agreement. Many previous studies have also shown substantial agreement for intra-observer or inter-observer variability by visual assessments (4,25,26). However, a recent study indicated that fully automated volumetric breast density measure can provide higher reproducibility, particularly with serial mammograms, than visual assessment (11). Therefore, computerized volumetric breast density measurement may be the preferred approach for both serial and non-serial assessment of breast density imaging.
This study had several limitations. First, there was a higher-proportion of patients in the screening group compared with the diagnostic group. Thus, the indication for mammography had the much higher proportion in the screening group. Second, we only compared the BI-RADS density categories from an experienced radiologist to that of Volpara. However inter-observer agreement for breast density evaluation by two readers was good. Finally, many factors that could also affect breast density including race, body mass index (BMI), nulliparity, and menopausal stage were not considered. For example, previous studies have found that Asian women tend to have denser breast tissue than Western women (27). Indeed, the Asian women in our study tended to show denser breast tissue than Western women in the study by Ciatto et al. (28). Therefore, further study with various clinical factors should be performed. However, our study has the largest study population of any study of volumetric breast composition measurement analysis, performed thus far (4,10,23,26,29).
In conclusion, automated volumetric breast density measurement shows good agreement with radiologists' visual assessment of BI-RADS density category as a quantitative method. However, a difference in bilateral breast density can affect the disagreement between assessment by radiologists and by automated software.
Footnotes
Conflict of interest
None declared.
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
