Abstract
Disaggregated statistics help improve the description of the society. However, survey estimates are subject to larger uncertainty at finer levels than at higher levels, and often not even available at fine levels. The Behavioral Risk Factor Surveillance System (BRFSS) is considered the nation’s premier system collecting health data from individuals in the US using telephone surveys. Among the BRFSS official statistics, state-level estimates are available for two related health prevalence quantities: the prevalence of having a personal doctor and the prevalence of having health insurance coverage. No county-level BRFSS estimates are released for these quantities. In addition, county-level estimates for the prevalence of having health insurance coverage are also available from the US. Small Area Health Insurance Estimates (SAHIE) program. This article addresses the disaggregation of the state-level prevalence of having a personal doctor to the county level, by using the state-level relationship between the two BRFSS prevalence variables and the county-level bridge between the BRFSS and the SAHIE prevalence of having health insurance coverage. Using 2018 public-use data, county-level model estimates are produced for both prevalence variables and on both BRFSS and SAHIE scales, improving the usability of the BRFSS public-use data.
Keywords
Introduction
Disaggregated statistics help improve the description of the society and inform funds allocation and policy decision making. Small area estimation (SAE) is a field that has grown rapidly in the recent decades to offer solutions to challenging estimation problems at fine levels of aggregation where survey data alone is not sufficient to produce reliable official statistics (Rao and Molina[11]). In a typical model-based SAE approach, sparse survey data from one source are combined with auxiliary data from multiple sources at the aggregation level of interest (Fay and Herriot[9]) or lower (Battese et al.[1]). The availability and the quality of these data components dictates the quality and usability of the resulting SAE data. The sparser the survey data are, the more one needs to rely on the model and the auxiliary data.
Multivariate SAE models address estimation for two or more related quantities of interest and help improve the precision of the model-based estimates, under joint modeling versus separate modeling. In the absence of reliable auxiliary data, the multivariate SAE models borrow strength from the relationship between the variables modeled jointly and across the small areas. These models often start with the survey data collected on these quantities of interest from one source at the aggregation level of interest or lower (Erciulescu and Opsomer[6]).
When multiple data sources include information on the same quantity, a bridging model can be developed to describe the relationship between the two sources. Such models have been recently considered for the purpose of maintaining trend and constructing of estimates at fine levels of aggregation (Erciulescu et al.[8]), for increasing the number of domains with available estimates (Erciulescu et al.[7]), and for describing the link between official statistics and constructing estimates at fine levels of aggregation (Erciulescu[4]). All these models were specified as multilevel models and Bayes inference was adopted. A review of statistical methods on measuring discontinuities in time series obtained with repeated sample surveys is available in van den Brakel et al.[13] Data blending methods that assume one data source to be less biased than the rest have also been studied by Raghunathan et al.[10] and Lohr and Brick (2012)[14] for domain estimation purposes, and by Balgobin et al.[2] for domain forecasting purposes.
In this article, we address estimation of one quantity of interest when only survey estimates at aggregated levels are available for it. A multivariate model is assumed at a high level of aggregation for the quantity of interest and one other related quantity, both collected from the same source. The other related quantity is also collected from a second source and made available at the fine level of aggregation of interest. A bridge model is assumed between the two data sources, and together with the multivariate model at the high level of aggregation, used to disaggregate the quantity of interest. None of the two data sources is assumed to be a gold standard, rather the various sources of error are accounted for by the model.
Erciulescu et al.[7] also assumed a multivariate model for two related variables collected from the same source and a bridge model between two data sources containing information on one of the variables, but they worked with data available at the same level of aggregation across sources. The bridging studies in Erciulescu et al.[8] and Erciulescu[4] addressed estimation for one quantity of interest only. So, this article builds on these previous studies by addressing estimation for another quantity of interest at the finest level of aggregation available from the two sources. The proposed model is also specified as a multilevel model and Bayes inference is adopted. A summary of all the scenarios mentioned thus far is available in Figure 1.
Literature Scenarios and Contribution of this Article. Bolded Information Represents New Data Produced under a Scenario. The Prediction Set is the Set of Domains of Interest.
Literature Scenarios and Contribution of this Article. Bolded Information Represents New Data Produced under a Scenario. The Prediction Set is the Set of Domains of Interest.
Publicly available health data are used to fit the proposed model. The two data sources comprise the United States (US) Behavioral Risk Factor Surveillance System (BRFSS) and the US Small Area Health Insurance Estimates (SAHIE) program. In the US, the Behavioral Risk Factor Surveillance System (BRFSS) is considered the nation’s premier system collecting health data from individuals in the US using telephone surveys. The SAHIE program is used to analyze health insurance coverage by geography and population groups across the nation. It is one of the first US program created to develop model-based official statistics.
Among the BRFSS official statistics, state-level estimates are available for two related BRFSS health prevalence quantities of interest for this article: the prevalence of having a personal doctor and the prevalence of having health insurance coverage. No county-level BRFSS estimates are released for these quantities because their quality does not pass publication standards. In addition, county-level estimates for the prevalence of having health insurance coverage are available from SAHIE, along with associated population totals allowing for aggregation of the county-level prevalence estimates to higher levels of aggregation such as state, census division, and nation. So, the quantity of interest available from both sources is the percent uninsured, and the related quantity of interest available from BRFSS only is the percent without personal doctor. This article addresses the disaggregation of the state-level prevalence of having a personal doctor to the county level, by using the state-level relationship between the two BRFSS prevalence variables and the county-level bridge between the BRFSS and the SAHIE prevalence of having health insurance coverage. County-level model estimates are produced for both prevalence variables and on both BRFSS and SAHIE scales, improving the usability of the BRFSS public-use data.
The rest of the article is organized as follows. In Section 2, we provide the background and motivation for this work, including a description of the available data and the initial estimation process, as well as selected results motivating the assumptions made in the disaggregation model presented in Section 3. Selected results are included in Section 4 and a discussion in given in Section 5.
In this article, we work with the 2018 BRFSS public-use data (Center for Disease Control and Prevention[3]) and with the 2018 SAHIE public-use data (United States Census Bureau[12]). The prediction set is defined as the set of all the counties in the 50 US states and the District of Columbia, except Kalawoa County in Hawaii, that is, 3,141 counties. The Kalawoa County is excluded from the SAHIE public-use data and noted as not applicable in the SAHIE data documentation, hence we exclude it from the prediction set, too. The target population of interest is defined as the adult individuals with age between 18 and 64. Following Erciulescu,[4] the public-use data are restricted to this target population of interest and to a common definition for the prevalence of having health insurance coverage, and the BRFSS survey weights are calibrated to the SAHIE population totals.
Initial Estimation
Let
Let
As shown in Erciulescu,[4] the BRFSS point estimates of the percent uninsured, and their associated uncertainty are larger than the corresponding SAHIE estimates for all the census divisions and for most of the states. In addition, a strong linear relationship is observed between the BRFSS and SAHIE point estimates of the percent uninsured at both the state and the census division levels. These observations motivated the linear bridge assumed in the models in Erciulescu,[4] and the same linear bridge is assumed here. Finally, a strong linear relationship is observed between the BRFSS point estimates of the percent uninsured and the percent without personal doctor. This observation motivates the joint modeling of the two BRFSS percentages. The linear relationships described for the state-level quantities are illustrated in Figure 2.
Relationships Observed Between the State-level Initial Estimates.
Relationships Observed Between the State-level Initial Estimates.
A disaggregation model is specified as a multilevel model with two levels for the initial estimates available from the two sources, two bridging levels to link the two sources and the two quantities of interest, and a smoothing level to improve the precision of the final model estimates. The smoothing level includes two nested latent effects corresponding to counties and states. According to Erciulescu,[4] using one latent effect would result in greater uncertainty in the final model estimates and using three nested latent effects does not have any noticeable benefits.
The first two levels of the model comprise a bivariate level for the BRFSS initial estimates and a univariate level for the SAHIE initial estimates, specified as follows:
Then, a latent linear bridge level is assumed between the percent uninsured measured on the BRFSS and SAHIE scales,
and another latent linear bridge level is assumed between the percent uninsured and the percent without personal doctor measured on the BRFSS scale,
Note that there are no initial county-level estimates for the percent uninsured and the percent without personal doctor measured on the BRFSS scale, that is, the components of
The final level of the disaggregation model is the smoothing level with two nested latent effects corresponding to county and state,
To use Bayes inference, we adopt independent weakly informative priors for the model parameters. The normal distribution with mean zero and variance 10,000 is assumed for the bridge levels regression coefficients and the smoothing level mean
Then, the models are fit in R JAGS using Markov chain Monte Carlo (MCMC), with three chains consisting of 20,000 draws with the first 5,000 as burn-in and with the remaining 15,000 thinned to every 10th draw to reduce dependence. The joint posterior distribution of all unknown model parameters is approximated by the resulting resulting
Model-based county-level predictions are constructed using the posterior samples for
where
Results
The state-level and census division-level estimates are illustrated in Figures 3 and 4, for the percent of uninsured and the percent without personal doctor, respectively. As the legend reflects, the different plotting symbols correspond to the different type of estimates: initial estimates or model-based predictions on both BRFSS and SAHIE scales. The initial estimates in Figure 4 are only on the BRFSS scale because the percent without personal doctor was not estimated in the SAHIE program. On the corresponding scales, the model predictions represent smoothed versions of the initial estimates. The magnitude of the difference between the initial estimate and the model-based prediction depends on the estimated variances of the initial estimates and on the discrepancy between the point estimates for the percent uninsured from the two sources, with these factors taken into account for the given geography (county, state, census division) as well as across geography. As expected, the model predictions on the BRFSS scale are closer to the BRFSS initial estimates and higher than the SAHIE initial estimates when the latter are available, and the model predictions on the SAHIE scale are closer to the SAHIE initial estimates for the percent uninsured. The model predictions are less variable than the initial estimates when compared for the same quantity and on the same scale. Specifically, the ratio between the model prediction variance and the initial estimate variance ranges from 1.5 percent to 25 percent for the state-level percent uninsured on the BRFSS scale, from 1.6 percent to 1.8 percent for the state-level percent without personal doctor on the BRFSS scale, and from 5.8 percent to 1 percent for the county-level percent uninsured on the SAHIE scale.
Census Division-level and State-level Estimates for the Percent Uninsured. Top: Point Estimates and Uncertainty Intervals for the Two Data Sources and the Two Scales Versus Census Division. Bottom: Point Estimates and Uncertainty Intervals for the Two Data Sources and the Two Scales Versus State.
Census Division-level and State-level Estimates for the Percent Uninsured. Top: Point Estimates and Uncertainty Intervals for the Two Data Sources and the Two Scales Versus Census Division. Bottom: Point Estimates and Uncertainty Intervals for the Two Data Sources and the Two Scales Versus State.
Census Division-level and State-level Estimates for the Percent Without Personal doctor. Top: Point Estimates and Uncertainty Intervals for the Two Data Sources and the Two Scales Versus Census Division. Bottom: Point Estimates and Uncertainty Intervals for the Two Data Sources and the Two Scales Versus State.
As a validation measure for the county-level model-based predictions of percent without a personal doctor, we compare their map against the map provided in Figure 2 in Erciulescu et al.[5] The latter estimates were produced using proprietary record-level BRFSS data and auxiliary data from various sources, using a small area estimation model. Despite the possible differences in the scales, note the similarities in the overall pattern across the US observed in the two maps in Figure 5.
County-level Estimates for the Percent Without Personal Doctor. Top Map is Produced using the Disaggregation Model. The Source of the Bottom Map is Erciulescu et al.[5] The Scales of the Maps may not Coincide, but the Ranges of the Scales are Similar, as Illustrated in the Legends.
Building on the bridging work of Erciulescu,[4] we extended the statistical data integration model to allow for joint estimation between two related quantities. For this, a disaggregation model is developed to link initial estimates from BRFSS and SAHIE on percent uninsured and initial BRFSS estimates for percent uninsured and percent without personal doctor, and construct model predictions for percent uninsured and percent without personal doctor at the county level. State-level BRFSS estimates and county-level SAHIE estimates served as the initial estimates and the bridge levels of the disaggregation model were specified as latent at the county level. The final model predictions have smaller uncertainty than the corresponding initial estimates, when compared for one quantity of interest and on one scale, as a result of including a smoothing level in the disaggregation model. The model can be improved by incorporating auxiliary data that better describe the quantities of interest at the county level. Similarly to the geographic disaggregation addressed in this manuscript, the proposed methodology could be applied to address data disaggregation by population groups defined by age, race, ethnicity, gender, disability, income, or other important demographic variables that help better describe society.
Footnotes
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author received no financial support for the research, authorship and/or publication of this article.
