Abstract
Secondary data analysis can benefit researchers of advanced academics by providing large sample sizes and a variety of data on multiple topics. However, using secondary data comes with unique challenges. This article will outline how gifted education researchers can find, access, and use secondary data. Data are available on children from birth to adulthood and are typically accessed through the Inter-university Consortium for Political and Social Research (ICPSR) or the National Center for Education Statistics (NCES). The majority of data sources have public-use files available, but some sensitive data may require special permissions. This article includes examples of advanced academic research that used popular databases along with software options for utilizing available data. We conclude with considerations researchers should take into account when considering using secondary data analysis, such as computer memory and technical skills.
In the past few decades, advanced academic research has typically included descriptive or correlation analyses rather than longitudinal or causal types of analyses (Dai, Swanson, & Cheng, 2011; Jolly & Kettler, 2008; Warne, Lazo, Ramos, & Ritter, 2012). Although some areas of gifted education have significant depth (e.g., acceleration, prescriptive curricular models), there are many areas of research that need attention such as social and emotional development of students and addressing the racial disparities in gifted education (Plucker & Callahan, 2014). One way to increase empirical research on these topics is through secondary data analysis. Secondary data analysis occurs when a researcher uses data gathered by someone else for another purpose to address the researcher’s own specific questions (Greenhoot & Dowsett, 2012). Analysis can be conducted to investigate new questions, extend previous research, or to replicate earlier findings from related research. Given the expanding use of secondary data analysis in gifted education, this article will address how secondary data can benefit advanced academic researchers and detail how to find, access, and use these datasets.
Benefits of Using Secondary Data Analysis
Secondary data, collected through governmental or other organizations, can provide many benefits to the field of gifted education. These data include information from hundreds or thousands of people on an assortment of variables that facilitate interdisciplinary research (Mueller & Hart, 2010). Arguably, the collaborative sharing of data and analysis is essential to advancing a body of knowledge or a field (Greenhoot & Dowsett, 2012). Publicly accessible data also facilitate researcher transparency, thorough documentation, clear justification of analytic processes, and verification of results prior to publication, which are good practices that uphold the scientific process (Donnellan, Trzesniewski, & Lucas, 2011). Finally, the open-access secondary data may accelerate the pace of research dissemination and advance the field in a more timely manner (Johnston, 2014).
The use of existing secondary data has multiple advantages for gifted education researchers. First, researchers do not have to gather data but can benefit from the financial resources and time invested by the primary researchers in collecting the data. Although all researchers benefit from this, secondary data may be an especially great resource for graduate students, early-career researchers, or underfunded researchers (Hakim, 1982; Johnston, 2014). Second, secondary data from national databases are generally of higher quality compared with what could be collected by individual researchers and are typically accompanied with technical manuals that provide complete documentation on the assessment design, data collection, and quality control (Donnellan et al., 2011). When done on a large scale, designers of these studies employ rigorous methods, such as complex sampling techniques to examine trends using nationally representative samples (Lohman & Marron, 2008; Mueller & Hart, 2010). As such, the findings have greater external validity and are more likely to be generalizable (Donnellan et al., 2011; Greenhoot & Dowsett, 2012). Third, large-scale public datasets may also provide access to specific subpopulations, including some subpopulations that are purposefully oversampled (Greenhoot & Dowsett, 2012). Fourth, longitudinal data collected in one study or from aggregating multiple studies for trend analysis allow researchers to examine changes over time. Finally, large-scale datasets are one way that researchers can utilize statistical modeling to address complex research questions (Greenhoot & Dowsett, 2012; Plucker & Callahan, 2014) and move beyond the descriptive or correlational research that has characterized the majority of research in gifted education (Dai et al., 2011).
Thus, the use of secondary data analysis provides researchers with tools to advance psychological science in ways that would otherwise not be possible (Donnellan et al., 2011). Given these advantages of using secondary data, researchers should consider exploring the vast possibilities to find, access, and use secondary datasets.
Finding Secondary Data
There are two popular sources for finding data for advanced academic research. The first is the Inter-university Consortium for Political and Social Research (ICPSR; 2019) website: https://www.icpsr.umich.edu/icpsrweb/. Based at the University of Michigan, the ICPSR archives over a quarter of a million data sources and provides topic-specific collections relevant to advanced academic researchers. On the website, members can easily search for data, specific variables, or relevant publications from the search bar at the top of the page. To maximize search capabilities, one should use Boolean operators for a more refined search. For example, a simple search for advanced returns over 12,000 results for variables but a search for “advanced placement” returns 195 results, making the search more manageable. Each result describes the research question, details the number of respondents for each answer selection, and links to the data. Boolean operators can help researchers not only find variables that are relevant to their topic but also help them find the most appropriate variable. For example, researchers can choose the data that ask students about whether or not they are in advanced placement classes (yes or no) or choose a detailed question about the number of advanced placement classes they have taken. By clicking on the “Search Series” tab using the same “advanced placement” search, researchers can find data sources that incorporate questions about advanced placement which can be helpful for finding alternative data when one dataset does not have other crucial variables for analysis.
Another resource for finding advanced academic data is the National Center for Education Statistics (NCES; n.d.) website: https://nces.ed.gov/. This website provides federal research pertaining to American education and links to any relevant international datasets. As all of the data pertain to education, researchers can select surveys based on the age of participants (i.e., early childhood, elementary, postsecondary, adult). If researchers simply want to search for data on a topic, there is an ED Data Inventory (https://datainventory.ed.gov/) website that is similar to searching on ICPSR. Unlike the ICPSR website, the NCES website provides a variety of online analytical tools to begin working with the data, which will be detailed later in this article.
Accessing Secondary Data
In the past, data were delivered via CD files; now almost all researchers can access the secondary data as well as a wealth of helpful associated resources via the aforementioned websites. While each dataset has its own materials, most include the raw data, a codebook, and a technical manual. A codebook details the variable names, variable descriptions, and response options to use when interpreting the raw data files. Some datasets have online codebooks that one can examine before downloading anything (e.g., https://nces.ed.gov/onlinecodebook). Technical manuals are essential for understanding how the instruments were developed, how the sample was collected, and how issues with data collection were resolved. Some datasets also have accompanying variable lists, publications, and training tools. Publications and products can be helpful in determining what others have done with the data and what gaps remain. However, not all publications have been posted on the websites; therefore, additional searches should be conducted to find related research utilizing the same data. Researchers should explore one or more of the data documents before digging into the data as the raw data are a large download that can be complicated to use without context.
There are two important types of labels for secondary data: public-use files (PUF) and restricted-use files (RUF). A PUF label indicates that anyone can download the data and begin to analyze it. Data with a PUF label are less detailed than RUF and may hide variables that could be used to identify schools or individuals. Restricted or hidden variables may include information regarding special education flags, income, or location, for example. If researchers need this detailed information, they can apply for a restricted-use data license. For more information on applying for a restricted-use license with NCES data, see https://nces.ed.gov/pubsearch/licenses.asp. To obtain an NCES license, researchers must be affiliated with an organization, such as a university, and detail how they will protect the data. Protecting the data typically requires a security officer, a locked room with a password-protected standalone computer, and an individual with specific rankings to serve as the project officer. Once the license is approved, only seven individuals may access the data. Although RUF files can provide researchers with more detailed data, the researchers may need to submit an institutional review board (IRB) proposal to their institution before conducting the project. Because PUF files make it practically impossible to identify the participants, projects do not generally require a full IRB review. When using RUF data, researchers much also get approval from NCES before sending work out for publication to ensure that the researchers adhered to the confidentiality procedures.
In addition, some surveys have multiple waves of data collection. ELS, for example, had a 2002 wave but also collected data throughout 2016. Many times these files can be merged on the same variables to examine trends or researchers can opt to use the most recent version of the data as it may include different variables. Not all waves will have the same variables, so it is best to understand which files a research team will need and have the ability to access before committing to using a particular dataset.
Using Secondary Data
For researchers conducting secondary data analysis in the field of gifted education, one of the primary decisions is how to operationally define giftedness within the constraints of the data collected. Decisions on data usage from well-known assessments should prioritize conceptual understandings of giftedness. No datasets, to our knowledge, have the ideal parent, student, and teacher reports in conjunction with multiple types of assessments supporting an advanced ability or gifted label. Instead, researchers often target a variable representing some traditional giftedness measure (e.g., identified as gifted by school, top 10% on an achievement or ability test, etc.). Multiple variables may also be used to enhance validity. In other words, there must be one or more variables in the secondary data set that can represent the construct or the researcher’s operational definition of giftedness. As a result, the term gifted could mean something different in each study.
Other practical considerations, such as sample size and ceiling effects, may also impact the operational definition of giftedness for a study using secondary data. As an illustration, some studies may need to widen the cutoff criteria, say to the top quartile, to have an adequate sample size for analyses. When examining exceptionally talented individuals, researchers must understand that standardized assessments, such as the SAT or ACT, have ceiling effects, making it difficult to differentiate between advanced and highly gifted populations (Park, Lubinski, & Benbow, 2007).
In addition to defining advanced ability or giftedness, researchers must identify research questions that can be addressed using the variables contained in the secondary data set. Each data set has a finite set of variables and the manner of collection may or may not support using the item as an independent or dependent variable, which limits the types of relationships that can be examined (Mueller, 2009). If a data set does not have adequate items to answer the questions of interest, the researchers must look for another dataset to meet their needs.
As with any research, authors and consumers of research should be aware that operational definitions, the selection of variables used to measure any construct, and the data analysis directly influence the conclusions and the generalizability of a study’s findings. For example, findings based on an operational definition of giftedness, such as the top 10% achievement on a version of the Peabody Picture Vocabulary test (Barber & Mueller, 2011), are unlikely to generalize to gifted students identified based on products or performances, creativity assessments, and parent or teacher recommendations. Furthermore, if the variables selected do not include an indicator that a teacher was formally certified in gifted education, it would be inappropriate to make inferences about the gifted education teacher preparation programs. Implications for policy would also differ depending on the population utilized. Researchers should clearly identify these limitations in each study.
Table 1 provides a short description of 14 popular datasets. A publication related to advanced academics or gifted education using secondary data is given as an example for each dataset. In addition, the variables used to operationally define advanced or gifted are reported. The table shows the variety of the variables used, from student to school level, in relation to advanced academic research. Depending on the type of analysis and data source, researchers can either use online programs to answer their research questions or download the data and analyze the data with their preferred software. For example, the NCES DataLab allows users to search for relevant data and, depending on the data, use QuickStats, PowerStats, or TrendStats to examine the desired variables. QuickStats enables users to compute simple descriptive tables, PowerStats allows users to compute more inferential statistics such as linear regression, and TrendStats uses the same variables from multiple datasets to show change over time. Similarly, the International Data Explorer (https://nces.ed.gov/surveys/international/ide/) allows users to compute descriptive statistics and run simple regressions for the larger international studies. However, these programs may limit users’ analysis options, and therefore, users may want to download the data for use on their own computers.
Examples of Advanced Academic or Gifted Education Research Using Secondary Data Analysis.
Note. More studies, including dissertations, are available through the NCES Bibliography Search Tool (https://nces.ed.gov/bibliography/). NCES = National Center for Education Statistics.
Both the ICPSR and NCES websites provide data files that can be downloaded in a variety of formats depending on the user’s needs. Popular software for using these analyses include R, SAS, SPSS, and Stata because the data websites often provide plug-ins to help with complicated aspects of the data (i.e., sampling weights). For example, the International Association for the Evaluation of Educational Achievement (IEA) developed a downloadable SPSS application called the International Database Analyzer (IDB) that includes the data and tools necessary to work with international large-scale assessments such as the Program for International Student Assessment (PISA) or the Trends in International Mathematics and Science Study (TIMSS). The tools are free but do require registration. For users interested in using complex data but do not have access to expensive software, we recommend the EdSurvey package in R. This package assists with issues such as plausible values with multiple national and international datasets (https://nces.ed.gov/nationsreportcard/researchcenter/software.aspx). Researchers can also download the text files of most data to upload in their software of choice.
Certain study creators also provide trainings for interested researchers at little or no cost. IEA and NCES also have long-distance training available on topics related to using a specific data source, utilizing the online programs or navigating difficulties with issues such as a complex sample design. These can be found at https://www.iea.nl/training#IDB_Analyzer_Video_Tutorials and https://nces.ed.gov/training/datauser/. Some datasets have training or grants available to encourage the use of the data and conferences to present findings from the databases; researchers should check the websites regularly and sign up for email updates if interested.
Considerations for Using Secondary Data
Although large datasets are enticing to researchers, they have technical and methodological considerations that do not arise with a typical project. For example, many of the large datasets included in this article have thousands of variables and cases, which may create additional issues with data storage and manipulation. Some statistical programs hold the entire dataset in Random Access Memory (RAM); large datasets can fill RAM quickly and slow a computer’s processing speed. Other programs designed to handle the size of the datasets may be limited in the range of analyses possible; the IDE and IDB provide point and click access to data and descriptive statistics, but they are not currently capable of performing advanced statistical analyses. Reducing the size of the dataset by deleting unnecessary variables may help especially if the researcher is interested in converting the dataset to a different type of file for use in a specific statistics program.
The last consideration, and possible limitation, for large datasets is the technical skills needed for analyzing this type of data. A working knowledge of statistics and access to helpful personnel resources for conducting quantitative research is needed. For example, large surveys, like Early Childhood Longitudinal Study (ECLS) K-11, use complex sampling frames and weighting—this allows for generalizability to the represented population. Sampling units represent the individuals who will be asked to participate in the survey; these units are placed within a sampling frame, or context, that represents all possible sampling units. Sample weights are used to ensure the frequency of specific groups from the sample also reflects that group’s frequency within the larger population—creating a representative sample. Because of this complexity, researchers need to apply sampling weights to most large-scale datasets. Analyzing the data without applying these weights will result in incorrect and often misleading estimates.
In addition to sample weights, researchers may need to apply replicate weights to correct the standard errors on estimates. Replicate weights are typically included in large-scale datasets that used complex sampling frames. To apply replicate weights correctly, researchers are urged to read the technical manual or user’s manual for their dataset. Replicate weights function like multiple imputation—researchers run the analysis for a given number of times and use the provided replicate weights to adjust the standard errors. Without replicate weights, results from these datasets may suffer from increased Type 1 error leading to false positives on significance.
If the dataset includes cognitive measures, plausible values may be used to provide population-level estimates of performance. Like sample and replicate weights, plausible values are used to help researchers generalize a limited set of information to a larger set of information. In this case, we are generalizing from a small number of cognitive items completed by a single individual to a range of possible scores that individual could have received if he or she completed the entire cognitive assessment. Sets of plausible values, or possible scores, are generated to account for the measurement error that results from generalizing performance on a few items to performance on all the items. Using plausible values incorrectly can reduce the variance in scores, which will directly impact the ability to see differences in performance between groups. Researchers are urged to learn more about plausible values if the data they selected use them.
Conclusion
Research in gifted education is moving toward more advanced methods, making use of secondary data critical for increasing the rigor of the field (Plucker & Callahan, 2014). Researchers can utilize secondary data to help answer new questions related to advanced academics or help replicate the findings of the past (Plucker & Callahan, 2014). Popular sources for finding data include ICPSR and NCES, two websites with varying search and analysis capabilities. Some data are available to the public, but researchers may also opt to apply for restricted-use files for more specific analyses. By using these datasets, researchers have access to quality data and resources, but researchers should maintain caution when using complex data without the proper tools or training.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
