Abstract
Background
Only few studies have assessed variability in the results obtained by the readers with different experience levels in comparison with automated volumetric breast density measurements.
Purpose
To examine the variations in breast density assessment according to BI-RADS categories among readers with different experience levels and to compare it with the results of automated quantitative measurements.
Material and Methods
Density assignment was done for 1000 screening mammograms by six readers with three different experience levels (breast-imaging experts, general radiologists, and students). Agreement level between the results obtained by the readers and the Volpara automated volumetric breast density measurements was assessed. The agreement analysis using two categories—non-dense and dense breast tissue—was also performed.
Results
Intra-reader agreement for experts, general radiologists, and students were almost perfect or substantial (k = 0.74–0.95). The agreement between visual assessments of the breast-imaging experts and volumetric assessments by Volpara was substantial (k = 0.77). The agreement was moderate between the experts and general radiologists (k = 0.67) and slight between the students and Volpara (k = 0.01). The agreement for the two category groups (nondense and dense) was almost perfect between the experts and Volpara (k = 0.83). The agreement was substantial between the experts and general radiologists (k = 0.78).
Conclusion
We observed similar high agreement levels between visual assessments of breast density performed by radiologists and the volumetric assessments. However, agreement levels were substantially lower for the untrained readers.
Introduction
The analysis of mammographic breast density has attracted considerable attention because it supplies the referring physicians with information regarding the mammographic sensitivity and the relative risk of developing breast cancer (1). Breast density depends on the ratio of fibroglandular tissue and fat. According to the BI-RADS lexicon, heterogeneously dense and extremely dense breasts are considered as dense, whereas breasts with scattered areas of fibroglandular density and almost entirely fatty breasts are considered non-dense (2). The risk of breast cancer for women with dense breasts is four to six times higher than that for women with non-dense tissue; the relative risk is greater than traditionally considered risk factors such as nulliparity or early menarche (3). Because of the importance of breast density in screening mammography, approximately half of U.S. states have implemented breast density reporting laws.
Ideally, inter-reader variability rates in these invaluable assessments should be low. However, in clinical practice, mammography is interpreted by radiologists with widely varying experience levels. Visual assessments of BI-RADS breast density category are subjective and agreement level between readers varies (4). With a wide spread of digital mammography, automated quantitative measurement tools, such as Volpara (Volpara Solutions, Wellington, New Zealand), for breast density assessment have been developed and are used in the routine clinical practice. According to a recent study, the mammographic density measured employing the full-field digital mammography is significantly associated with breast cancer risk when assessed by radiologists using the BI-RADS classification and employing fully automated computer-assisted methods (5).
To our knowledge, only few studies have assessed variability in the results obtained by the readers with different experience levels in comparison with automated volumetric breast density measurements (6,7). The purpose of this study was to examine the variation in breast density assessments using BI-RADS categories, performed by readers with different experience levels. We compared the results with those of the automated quantitative measurements.
Material and Methods
Our institutional review board approved this retrospective study and the requirement for informed consent was waived. A set of 1000 screening mammograms, obtained in our institute from 1000 consecutive healthy participants between January 2016 and June 2016, was used in the study. All the participants were women; the participants with a history of breast surgery, augmentation, or foreign body injections were excluded. Routine craniocaudal and mediolateral oblique views were obtained for each breast using dedicated full-field digital mammography (Senographe DS; GE Healthcare, Milwaukee, WI, USA).
Visual assessment
Mammographic density assignment according to BI-RADS categories was performed by six readers. They had three different experience levels; two were breast-imaging experts with more than five years of experience in reading mammograms, two were general radiologists with fewer years of experience in reading mammograms, and two were medical students without clinical experience in breast imaging. Two medical students were trained to read total of 80 mammogram set comprised of 20 mammograms per each Volpara density categories. All readers were blinded to the results of automated volumetric breast density measurements and read the mammograms independently. Each radiologist read the mammograms twice (phase 1 and phase 2) with an interval of two months, blinded to the previous results. The following BI-RADS categories for breast density were used for the interpretations: category a, tissue almost entirely fatty; category b, scattered areas of fibroglandular density; category c, heterogeneously dense (which might obscure small masses); and category d, extremely dense (which might lower the sensitivity of mammography) (8). In addition, we used two groups of breast density: non-dense and dense. BI-RADS assessment categories a and b were considered non-dense, and category c and d were considered dense.
Automated volumetric breast density measurement
For fully automated volumetric analysis, Volpara software was used. The output from each program consisted of the volume of fibroglandular tissue (cm3), breast volume (cm3), and the volumetric breast density (%) for each image. The Volpara density grade (VDG) was extracted automatically. The grades were obtained according to the percentage volumetric breast density: VDG 1 = < 4.5%; VDG 2 = 4.5–7.5%; VDG 3 = 7.5–15.5; and VDG 4 = ≥15.5% (9). BI-RADS assessment categories a, b, c, and d were regarded to correlate with VDG 1, 2, 3, and 4, respectively.
Data and statistical analysis
Intra-reader agreement on BI-RADS density category was analyzed using the readings made by reader pairs in phase 1 and 2. The agreement between visually assessed BI-RADS categories and automated volumetric breast density measurements was assessed using the weighted kappa statistics. For inter-reader analysis, the reader with better intra-reader agreement was chosen from each group. The intra-reader and inter-reader analyses were also performed for the two broader categories: non-dense and dense. The kappa values were interpreted as suggested by Landis and Koch (10): a kappa value of ≤ 0.20 indicates slight agreement; 0.21–0.40, fair agreement; 0.41–0.60, moderate agreement; 0.61–0.80, substantial agreement; and 0.81–1.00, almost perfect agreement. Statistical comparisons were performed using the independent or paired t test or the Pearson test for continuous variables and the chi-square or Fisher’s exact test for categorical variables. Statistical analysis was performed using SPSS statistical analysis software (PASW Statistics, version 21.0.0; SPSS Inc., Chicago, IL, USA), and P < 0.05 was considered indicative of a statistically significant difference.
Results
Baseline characteristics
Results of automated volumetric measurements for 1000 women.
Version 1.5.12 (Mātakina Technology, Wellington, New Zealand).
SD, standard deviation.
Agreement analyses: BI-RADS category
Intra-reader agreement in assessments of breast density.
Data are weighted kappa values (95% confidence intervals).
Inter-reader agreement between visual and volumetric assessments of breast density.
Data are weighted kappa value (95% confidence intervals).
Agreement analyses: non-dense versus dense
Intra-reader agreements on the non-dense and dense group classification among the breast-imaging experts, general radiologists, and students were almost perfect or substantial (k = 0.76–0.95) (Table 2). The agreement between visual and volumetric assessments of breast density for these two groups is shown in Table 3. The agreement between visual assessments of the breast-imaging expert and volumetric assessments by Volpara was almost perfect (k = 0.83). The agreement was substantial between visual assessments of general radiologist and volumetric assessment by Volpara (k = 0.73), and between visual assessment of the expert and general radiologist (k = 0.78). The agreement between visual assessments of the students and volumetric assessments by Volpara was slight (k = 0.01).
Comparison of breast density categories
Comparison of breast density categories assigned by the readers and volumetric measurement method.
Except for P values, the data represent the number (%) of examinations.
Readers assessed breast density according to the BI-RADS density categories, whereas the Volpara software (version 1.5.12, Mātakina Technology, Wellington, New Zealand) assigns density grades as Volpara density grade.

Representative case of dense breast. Screening mammogram including both craniocaoudal (a) and mediolated oblique view (b) is shown which was categorized as BI-RADS category c by experts and general radiologists. According to Volpara software (c), density grade obtained was Volpara density grade 3 (9.9%) which correlates with BI-RADS category c.
Discussion
Our study supports the common assumption that the agreement between visual assessments performed by readers and obtained by volumetric measurements of breast density differs depending on the reader experience levels. The agreement improves with increasing reader experience.
The data from the Breast Cancer Surveillance Consortium, collected in USA, show that on 934,098 negative screening mammograms (during 1994–2008) the distribution of tissue densities was 9.0%, 44.1%, 38.3%, and 8.6% for fatty breasts, breasts with scattered areas of fibroglandular tissue, heterogeneously dense breasts, and extremely dense breasts, respectively (11). In our study, the distribution of tissue density was different; the proportion of dense breast tissue was notably higher, comprising 68.6% of the participants by volumetric measurements. The population in our study was only comprised of Asian and it reported that the mammographic breast density is substantially higher in Asian women than in African American and white women (12–14).
The consistency of the breast density classification has implications in clinical practice, including the development of stratification models for breast cancer risk and decision-making for asymptomatic women (15,16). Some studies have reported inter-reader agreement for BI-RADS classification but showed rather variable results (9,17–19). Furthermore, according to a systematic review of current evidence on the reproducibility of BI-RADS breast density, density may be re-categorized on serial screening mammograms (20). In our study, there was substantial or almost perfect agreement between visual and volumetric assessments of breast density (using the BI-RADS density categories or two category groups, non-dense and dense) for breast-imaging experts and general radiologists, with a slightly better agreement for the experts. There was also a moderate agreement between the experts and general radiologists. We found that the general radiologists categorized more mammograms as dense breast tissue in the two-group classification and more as category d than category c in the BI-RADS classification. We assumed that their lower confidence level might have affected the result.
There was a slight agreement between visual assessments by untrained students and volumetric measurements by Volpara (both for the BI-RADS density categories and for the two category groups). Even though their intra-reader agreement was substantial or almost perfect, the scores for accurate density assignment of the untrained students were far below expectation. Previous reports have shown that untrained radiologists significantly overestimate the percentage density in comparison with trained individuals; the accuracy of mammographic density assignments improves after training (6,7). These data, as well as our study results, support the recommendation of the U.S. Mammography Quality Standards Act, which requires that readers should interpret an average of 480 mammograms per year. The high interpretive volumes should increase experience levels and improve accuracy (21).
There are several limitations to our study. First, all the mammographic examinations were performed in a single mammographic unit, with only one specific kind of automated quantitative measurement to be used for comparisons. However, employing a unified equipment and software might have increased the data reliability. Second, the number of the readers was small and they were all trained at the same institution. However, we tried to assess the differences between the readers with different experience levels, which would reflect the situation often found in clinical practice. Finally, the automated volumetric measurement was used as a reference standard. The revised fifth edition of BI-RADS no longer indicates the ranges of the percentage of dense tissue and emphasizes the changes in mammography sensitivity. There is no other standard reference for mammographic density assignment in clinical practice.
In conclusion, we found a similar, high level of agreement between visual assessments by radiologists and volumetric assessments of breast density by Volpara. For untrained readers, agreement level was substantially lower.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
