Abstract

There is a growing demand to extract information from multiple data sources to address emerging challenges in survey and data science. Often, there is a need for estimates at disaggregated levels beyond what traditional survey estimation techniques can support. This is exactly the realm of Small Area Estimation (SAE). The international conference SAE 2022: Small Area Estimation, Surveys and Data Science was held at the University of Maryland College Park campus on 23–27 May 2022. This conference served as a bridge among statisticians, survey methodologists, engineers, mathematicians, computer scientists, and others interested in combining information from multiple data sources to develop reliable inference at granular levels. In addition to traditional topics in SAE the conference covered a few emerging topics in surveys and official statistics (e.g., nonprobability sampling, probabilistic record linkage, data fusion, etc.).
To elucidate the current state of research involving emerging topics in small area estimation, surveys, and data science, we discussed the possibility of publishing a special issue highlighting the topics covered in the SAE 2022 international conference. We thank the editors of the Calcutta Statistical Association Bulletin (CSAB) for agreeing to devote two special issues of CSAB to the conference, and for inviting us to be their editors. We agreed that anyone, including the participants of SAE 2022, could submit papers for possible publication in the special issues, and that all papers would go through a thorough review process.
Of the multiple papers that were submitted by our deadline of 15 January 2023, we finally accepted fourteen papers, after a rigorous refereeing and revision process. These accepted papers will be published here and in the first issue of 2024. The two special issues should serve as excellent references, and we hope they will inspire future research in these exciting and challenging topics. We now briefly describe the papers published in this issue.
Out of the seven papers published in this issue, the papers by (a) Erciulescu, (b) Bandyopadhyay and Jiang, (c) Tang and Ghosh, (d) Pratesi, Marchetti, Giusti, and Salvati, and (e) Sverchkov and Pfeffermann advance research on SAE in several different directions. The papers by (f) Das, Salvati and Chambers and (g) Hirose and Mano address estimation problems in finite population sampling. The models and/or tools used in the latter two papers can also be related to the SAE literature. The seven papers published in this special issue collectively advance knowledge in surveys and data science.
Researchers have been developing SAE methodology using models applied on aggregates (e.g., estimates) or on observations at the ultimate unit (e.g., person) level. Papers (a)–(d) use models on estimates while paper (e) uses models on observational micro units. Inclusion of spatial correlations is explored in papers (c) and (d). Paper (a) explores a small area model whose aim is to disaggregate estimates obtained for larger population domains. The proposal is driven by a specific important application. Paper (b) advances the Observed Best Prediction (OBP) methodology by incorporating a benchmarking criterion that ensures model predictions will yield the direct estimators at higher levels of aggregation. Paper (c), unlike other papers in this issue, advances Bayesian estimation for spatial aggregates by using a Global-Local prior that accounts for wide variation in the random effects. Paper (d) covers models with spatially structured random effects, models based on nonparametric spatial splines, and models with spatially varying regression coefficients. Using Italian data, the authors compare efficiency of various predictors in two scenarios that differ in the predictive power of the auxiliary information. Paper (e) addresses a few outstanding issues pertaining to unit level models, including informative sampling, not missing at random nonresponse mechanisms and selection of the response model.
Using a nested error regression model, paper (f) develops predictors of a finite population distribution function. The authors suggest that their theoretical framework could potentially be applied to solve small area estimation problems. Finally, like paper (f), paper (g) uses a superpopulation model to derive asymptotic uniformly minimum variance unbiased estimators of disclosure risks. Their method develops an adjusted maximum likelihood method, used earlier in SAE that guarantees estimates in the interior of the parameter space.
We would like to thank the authors for submitting their papers to this special issue. Thanks are also due to the anonymous referees who offered many constructive suggestions to improve the quality of the original submissions. We would like to thank Professors Manisha Pal and Tathagata Bandyopadhyay, Editors-in-Chief, for encouraging us to take the lead on this project. We appreciate all the help we received from Professor Gauranga Chattopadhyay, Coordinating Editor of CSA Bulletin, and the SAGE editorial staff, especially Neha Bahuguna and Himani Raghav. Without their enormous help, we would not have these high-quality special issues.
