Abstract
Data science is an emerging field that provides new analytical methods. It incorporates novel data sources (eg, internet data) and methods (eg, machine learning) that offer valuable and timely insights into public health issues, including injury and violence prevention. The objective of this research was to describe ethical considerations for public health data scientists conducting injury and violence prevention–related data science projects to prevent unintended ethical, legal, and social consequences, such as loss of privacy or loss of public trust. We first reviewed foundational bioethics and public health ethics literature to identify key ethical concepts relevant to public health data science. After identifying these ethics concepts, we held a series of discussions to organize them under broad ethical domains. Within each domain, we examined relevant ethics concepts from our review of the primary literature. Lastly, we developed questions for each ethical domain to facilitate the early conceptualization stage of the ethical analysis of injury and violence prevention projects. We identified 4 ethical domains: privacy, responsible stewardship, justice as fairness, and inclusivity and engagement. We determined that each domain carries equal weight, with no consideration bearing more importance than the others. Examples of ethical considerations are clearly identifying project goals, determining whether people included in projects are at risk of reidentification through external sources or linkages, and evaluating and minimizing the potential for bias in data sources used. As data science methodologies are incorporated into public health research to work toward reducing the effect of injury and violence on individuals, families, and communities in the United States, we recommend that relevant ethical issues be identified, considered, and addressed.
Data science is rapidly transforming public health practice. 1 Unlike traditional analytical methods, such as biostatistics and epidemiology, which focus on structured data, data science uses a confluence of methods from computer science, statistics, and other disciplines to analyze large or novel structured or unstructured data sources.1-4 Through applications such as machine learning and natural language processing, data science holds promise for generating timely, accurate, detailed, and specific analyses that can provide new insights or a deep understanding of public health trends. 1
Data science holds great promise for injury and violence prevention. Injury- and violence-related mortality rates have been rising in the United States, and unintentional injuries, suicides, and homicides are in the top 10 leading causes of death for people aged 1 to 44 years.5-7 Insights from data science methods can allow public health agencies and injury and violence organizations to better focus resources and prevention efforts and mitigate these trends.1,8 For example, machine-learning modeling can provide estimates of weekly suicide fatalities in the United States, and such estimates can then be used to distribute public health resources toward prevention efforts. 9
Injury and violence encompass a range of issues (eg, opioid overdose, motor vehicle accidents, intimate partner violence, sexual violence) that affect all people, regardless of age, sex, race, and socioeconomic status. 10 Notably, injury and violence data often involve potentially sensitive, intimate, or identifiable information, which, if handled or used incorrectly, can lead to unintended harms for members of the public and public health practitioners. These harms can include loss of privacy, loss of employment, discrimination or social stigma, and loss of institutional trust.9,11 A cautious and thoughtful approach is needed when applying emerging data science methods to minimize unintended harms and promote public beneficence. 12
Insights from the fields of bioethics and public health ethics, when applied to public health data science projects in injury and violence prevention, may minimize such risks of inadvertent harm. In this article, we distill key concepts from foundational bioethics and public health ethics literature into 4 ethical domains relevant to the practice of data science in injury and violence prevention. For each domain, we also share a list of ethics questions to consider when planning or conducting injury- and violence-related projects. In doing so, we aim to provide the data science workforce with a framework for ethical enquiry into the design, implementation, and evaluation of injury and violence prevention projects to minimize harms and promote public health and safety. 13
Methods
We first reviewed foundational bioethics and public health ethics literature to identify key ethical concepts relevant to public health data science. Bioethics sources included the Belmont Report, 14 Beauchamp and Childress’s biomedical ethics framework, 15 and reports issued by the Presidential Commission for the Study of Biomedical Ethics (2010-2016).12,16-20 Public health ethics sources included the American Public Health Association’s (APHA’s) Public Health Code of Ethics. 13
Our review of bioethics sources revealed key ethics concepts: autonomy (respect for people), beneficence, nonmaleficence, justice and fairness, privacy, responsible stewardship, intellectual freedom and responsibility, and democratic deliberation.12,14-20 We identified additional concepts through review of public health ethics sources: professionalism and trust, health and safety, health justice and equity, interdependence and solidarity, human rights and civil liberties, and inclusivity and engagement. 13
After identifying these ethics concepts, a core group of authors (N.I., E.B., P.C., L.O., J.B., L.N., R.L.) held a series of discussions to organize these concepts under broad ethical domains especially pertinent to the practice of public health data science in injury and violence prevention. Within each domain, we then examined relevant ethics concepts from our review of the primary literature (eg, autonomy, transparency). Lastly, core authors developed questions for each ethical domain to facilitate the ethical analysis of injury and violence prevention projects.
Discussion
We identified 4 ethical domains: privacy, responsible stewardship, justice as fairness, and inclusivity and engagement. Each domain carries equal weight, with no consideration bearing more importance than the others.
Privacy
Privacy is a general concept that “includes confidentiality, secrecy, anonymity, data protection, data security, fair information practices, decisional autonomy, and freedom from unwanted intrusion.” 16 Its foundational role in ethical biomedical practice was instantiated through the Health Insurance Portability and Accountability Act of 1996, 21 whose privacy rule established boundaries for disclosing protected health information and using data such as names, small geographic subdivisions, and elements of dates.22,23
Data used to confirm a person’s identity (personally identifiable information) can include direct or sensitive identifiers (eg, name, medical records, mailing address) and publicly accessible indirect or nonsensitive identifiers that can be used to reconstruct identity, especially when combined with other data. For example, linking health data to publicly available data, such as data involving a tax assessor, which contains personally identifiable information, could have social consequences for the person whose data are linked. Data privacy and informed consent are 2 concepts related to privacy that are particularly relevant to injury and violence projects.
Data privacy
Revealing a person’s private information when working with sensitive injury and violence topics could have social consequences and lead to diminished trust between data stewards and the people from whom the data are collected. Current best practices include identifying and removing personally identifiable information, routinely ensuring that the security standards of hardware and software are up-to-date, and restricting user access to ensure privacy. Privacy should be maintained through the entire pipeline of data use, from collection through analysis and dissemination. Data scientists should be aware that the risk of reidentification is not inconsequential when analyzing or reporting a person’s direct or indirect personal information, especially if using combinations of multiple publicly available data sources.
Considerations for data scientists:
• Will the project involve personal, sensitive, or intimate identifiable information? 16 How can the data be anonymized?
• Are people at risk of reidentification through external sources or linkages if publicly available data do not meet confidentiality requirements? How can we minimize this risk? How will we inform people of this risk? The risk of reidentification evolves with new tools and new connections between data sources and, hence, requires ongoing assessment and evaluation.24,25
• What data-sharing and data-use agreements need to be put in place to ensure that private data are protected? Can we verify that entities or groups with whom we share data are able to respect data privacy?
• What processes, protocols, and systems are in place to ensure the confidentiality of personal data (eg, data encryption, data access restrictions, processes to securely dispose of data)?
Informed consent
Informed consent is a fundamental concept in clinical and research ethics and reflects the primacy of a person’s right to autonomy in choosing to participate in research. It involves 3 elements: voluntariness, information disclosure, and decision-making capacity. 16 Informed consent for data use has notable implications in data science, which have been complicated by rapid technological developments such as the increased capacity for widespread aggregation of publicly available internet-based data (eg, social media data, forum data) and the rise in interoperability of different data systems. 26
Considerations for data scientists:
• How can informed consent be meaningfully obtained from publicly available data sources before data are collected?
• If the project uses publicly available information, surveillance data, or data collected or aggregated by private entities, how were data gathered and for what uses?
Responsible Stewardship
Responsible stewardship is a shared duty to represent communities or to stand in where they might not have an opportunity to represent themselves; it embraces the core APHA public health ethics values of professionalism and trust.12,13 To be responsible data stewards, data scientists should conduct data security evaluations and data safety assessments and serve as a fiduciary throughout the lifetime of each data science project.5,27 These roles are particularly important in areas of emerging methods, such as generative artificial intelligence or complex data science models, where existing ethical frameworks may not keep pace with new analytic techniques. 12 Responsible stewardship fosters transparency, integrity, trustworthiness, and teamwork among all parties involved in a project, including members of the public.13,28-30 Responsible stewards ensure privacy and weigh harms and benefits before and during the sharing of any data products. Responsible stewards need to understand how the data are being collected and used in their data science models and critically examine analysis results and any information that is being shared. Areas of responsible stewardship particularly relevant to injury and violence projects are transparency and data security.
Transparency
Transparency is an ethical concept that, in data science, refers to the practice of making information and data readily accessible, understandable, and open to scrutiny by relevant partners. It reassures all parties that analysts are acting responsibly and ethically and that findings are reproducible. Data scientists need to consider how best to ensure transparency and reproducibility. Notably, while transparency should be maximized, it should not infringe on individual- or population-level privacy needs. Adherence to privacy standards and guidelines during and beyond a project’s life cycle should be continuously monitored to prevent and promptly respond to privacy breaches. Additionally, the appropriateness of fulfilling individual-level data requests should be evaluated in accordance with data-use agreements, policy, regulations, and privacy considerations. Transparent and standardized processes can facilitate trustworthiness and open communication between data scientists and the public.
Considerations for data scientists include:
• How can a project’s algorithms, code, documents, processes, results, and data be audited and shared with other data scientists or the public?
• How can the accuracy and validity of models implemented on small scales be ensured, as in studies performed with limited segments of the population or specific demographic groups or communities? How can accuracy and validity be preserved if data science models are scaled up or applied to other populations?
• If the models produce incorrect, inaccurate, or invalid outputs, how can errors be identified and corrected in a timely way?
• How and when should data scientists not involved in the project be engaged to independently verify the analysis and results?
• Is the data science model opaque (ie, a black box model)? If so, are any alternate models available that may facilitate greater transparency? If not, what efforts can be taken to explain the methods and results in an accessible way to other researchers and the public?
Data security
Data security is paramount for responsible stewardship. Data stewards can advance security in the field of data science by protecting data, metadata, relevant computer systems, and the infrastructure hosting or housing data from harm, theft, and unauthorized entry. Computing and technology systems are growing and advancing exponentially; therefore, maintaining the security of public health data and information systems depends on keeping abreast of these advances. Data-use and data-sharing agreements are vital prior to doing any analysis and must ensure the safety and security of data and their output.
Considerations for data scientists include:
• Is the system protected from outside intruders and hackers? How can the system’s security be strengthened and augmented?
• Can the system be turned off when it is showing an error?
• Can this technology be attacked or abused? 31 If so, what are the remedies?
• Is there a plan to protect, secure, and store user data during the project and after its completion?
Justice as Fairness
The APHA Public Health Code of Ethics understands health justice not only in terms of the equitable distribution of scarce resources but also as the “remediation of structural and institutional forms of domination that arise from inequalities.” 13 This latter concern is especially pertinent to problems of justice and fairness in developing and deploying algorithmic models. Although a familiar characterization of justice is treating equals as equals, scholarship on the presence of bias in the development of these models has demonstrated not only how they can treat equals unequally but also how they can impose disproportionate benefits or burdens (eg, loss of employment) on distinct populations who have experienced a history of socioeconomic and health inequalities.32,33 For example, models can be tainted by biased datasets that reflect centuries of racial and ethnic prejudice. For racial and ethnic minority populations, this results in poor access to health care, inaccuracies or omissions in medical records, and poor measurement calibration for some groups—such as inferior detection or identification of dermatologic lesions among dark-skinned people or entry of derogatory descriptors (eg, “noncompliant,” “challenging”) into patients’ electronic health records. This also results in the social justice dilemma of moral trade-offs presented by a model, such as its predictive accuracy or, alternatively, its error rates favoring one affected population over another.33-41 Data quality assessments on important variables such as race, ethnicity, age, sex, disability, or geography can be vital for ensuring project fairness, and they must be sustained through the life of a model. 42 While not all differences or disparities between populations are attributable to bias, models should be assessed to determine whether predictions may be affected by these factors. Many suggestions have been offered for mitigating bias or unfairness in models, but the future will likely witness a robustly democratic process of ethical evaluation bearing on “justice as fairness” by all community members, especially patients or others in the public who are affected by them. 43
Considerations for data scientists include:
• What are the known and potential biases in the data (eg, historical measurement, misrepresentation, aggregation, proxy)? Did the project team explore any potential bias in the data with community members?
• Could the data be used or interpreted to produce biased results (eg, ageism, sexism, racism)?
• Were any potential factors missed that could affect the results, especially over time?
• Does diversity exist in the backgrounds, experiences, beliefs, and perspectives of the project team and other engaged partners? Was this diversity reflected in the design of the project?
• Have the training data been tested to ensure that they are fair and representative? 31
• Could the model or data drift be considered to remain fair over time? 31
• How are the benefits and burdens of this project distributed? Are certain populations unfairly benefited or disadvantaged?
• Will communities be harmed by this project in the short term? In the long term?
• Is communication between community members and the project team transparent in ways that mitigate any new risk of biases or error and ensure scientific integrity?
Inclusivity and Engagement
Inclusivity and engagement (encompassing ethical concepts such as democratic deliberation; transparency; and justice, fairness, and impartiality) encourage data scientists and representatives of affected groups to collaboratively participate in dialogue, listen to and meaningfully consider opposing perspectives, and negotiate the appropriate boundaries of data science projects.12,13,18,44 Community partners should be included in the decision-making process when creating, understanding, and deciding necessary ethical boundaries for data science projects.13,18 A collaborative community involvement process should be incorporated in data science projects from collection to analysis so that disseminated results can foster understanding, validation, and trust in public health.13,45 Without such intentional collaboration, public health practitioners might overrely on data science while excluding community partners; therefore, engaging with the community throughout the process is necessary. 46 As an example of early engagement, injury-related data science projects might involve asking people who have experienced an injury (eg, adverse childhood experiences, drug overdose, suicide attempts or suicidal thoughts, community violence) about relevant factors that could be considered prior to embarking on a related data science project.
Considerations for data scientists:
• How was the project’s purpose determined? Who was involved in determining the project’s purpose?
• Are the project’s purpose and scope clearly defined?
• Have affected communities been engaged to hear their perspectives? How can community engagement be measured or assessed? What is the best way to identify people from these communities who can inform the project?
• Do the communities understand the risks and benefits of the project? Could communities be harmed in the short or long term because of this project? What are the methods of reducing or eliminating the risk of harm in the short or long term of the project?
Public Health Implications
Data science is evolving quickly and is increasingly applied in public health projects on injury and violence. Although data science methods hold tremendous promise, incorporating ethical considerations in injury and violence projects can help to ensure that these methods are applied and interpreted responsibly. When applied early in the process, ethical considerations can provide a critical examination of assumptions and a framework for practitioners to strengthen data and analytical processes. Data scientists working in the field of injury research can benefit from explicitly ensuring that ongoing ethical discussions and considerations occur before, throughout, and beyond the life cycle of research projects. Simultaneously, advancing ethical considerations will not only complement and enhance scientific training but may lead to better public health data scientists. 46
Footnotes
Acknowledgements
The authors thank Amy Wolkin, DrPH (Centers for Disease Control and Prevention [CDC]); Ira Bedzow, PhD, MA (Emory University); Mary Leinhos, PhD, MS (CDC); Chad Heilig, PhD (CDC); Tom Savel, MD (CDC); and Walter Wietzke, PhD (University of Wisconsin–River Falls) for contributing feedback to this project. The authors Isaac Asimov and Ursula K. Le Guin provided inspiration for this project.
The conclusions, findings, and opinions expressed by the authors contributing to this article do not necessarily reflect the official position of the US Department of Health and Human Services, the US Public Health Service, CDC, or Emory University.
Institutional review board approval for this project was not required by the Centers for Disease Control and Prevention because it considered the project to not have involved human subjects research.
Dedication
We dedicate this article to the memory of our esteemed friend, mentor, colleague, and coauthor Leonard Ortmann, PhD, who died May 4, 2023.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
