Abstract
Background:
There has been growing interest in the application of artificial intelligence (AI) in thyroid cancer care, given its potential to enhance diagnostic accuracy, predict patient outcomes, and personalize treatment plans. However, bias introduced during the development of AI algorithms used for thyroid cancer care poses a significant challenge, as biased datasets can lead to disparities in diagnosis and treatment recommendations, particularly in underrepresented populations. This systematic review evaluates the current landscape of AI models for thyroid cancer, focusing on demographic representation and potential biases.
Methods:
This systematic review was registered on PROSPERO (ID: CRD42024519238) and conducted in accordance with the Cochrane handbook and reported in accordance with Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. A literature search was performed on EMBASE, PubMed, and Google Scholar up to January 2024. Studies were included if they involved AI models for thyroid cancer management and provided demographic details. Data extraction and risk-of-bias assessments were conducted by two independent reviewers.
Results:
A total of 197 studies were included in the review, with the majority focusing on diagnosis (n = 133) and prediction/prognosis (n = 47). Most studies predominantly involved participants from China (n = 124) and the United States (n = 26), with more female participants (n = 12,410) than males (n = 4222). Ethnicity data from 197 studies (248,896 participants) revealed a significant underrepresentation of East Asians (14.6%) compared with their global thyroid cancer prevalence (18.7%), while White (26.8%) and Black participants (26.8%) were overrepresented relative to their global prevalence (20.7% and 11.3%, respectively). Socioeconomic factors, marital status, and race/ethnicity were less frequently considered in the models.
Conclusion:
The findings highlight significant gaps in the diversity and representativeness of data used in thyroid cancer AI models. Current models align with epidemiological trends but lack comprehensive demographic inclusion. As such, more representative AI models are required that account for all aspects of a patient’s demographics and sociocultural background. Future research should focus on developing and validating more equitable AI models to improve thyroid cancer care across diverse populations.
Introduction
Thyroid cancer is the most common endocrine malignancy worldwide, 1 with significant variation in incidence across different ethnic groups and regions. 2 As the medical community increasingly explores artificial intelligence (AI) models to enhance cancer diagnosis and treatment, it is crucial to address the issue of bias in AI models. Bias can arise from underrepresentation of certain populations in training datasets, leading to disparities in accuracy and effectiveness of AI-driven tools for minority groups. 3 In the context of thyroid cancer, where ethnic and regional factors play a role in disease prevalence and outcomes, it is essential to ensure that AI models are designed to mitigate bias and promote equitable health care across diverse populations.
AI has rapidly emerged as a transformative tool in health care, with applications ranging from medical imaging analysis to personalized treatment recommendations. 4 In thyroid cancer, AI models are being developed to assist clinicians in early detection and diagnosis based on ultrasonographic appearance of thyroid nodules,3–5 risk stratification, 6 management, 7 and streamline care pathways. 8 However, the success of these models is highly dependent on the quality and diversity of the data they are trained on. Bias can infiltrate AI systems when the training datasets are skewed toward certain populations, typically those with more health care access, such as Caucasian or high-income groups, resulting in reduced performance for underrepresented groups. For thyroid cancer, this bias may lead to less accurate diagnoses or suboptimal treatment recommendations for minority populations, further widening the gap in health care equity.
Bias in AI has been extensively studied for certain conditions, such as melanoma, which demonstrated that AI-based diagnostic tools often perform as well as or better than dermatologists but are limited by a lack of generalizability across diverse skin types and populations.7–9 In fact, a meta-analysis of AI algorithms created for melanoma diagnoses showed that AI tools achieve high accuracy in detection with sensitivities exceeding 80%. However, nominal refinements in training photos are required for further performance boosts across varied demographics. 10 Nevertheless, unlike dermatological conditions, where differences in skin types are more visually apparent, demographic differences in thyroid cancer—such as ethnicity, and genetic variants— can influence multiple stages of care, from diagnosis to prognosis.11,12 For instance, studies have shown that racial and ethnic minority groups often face disparities in access to timely diagnostic evaluation, are less likely to receive guideline-concordant treatment, and may experience poorer oncologic outcomes due to systemic barriers in care delivery.13,14 These demographic factors influence disease characteristics and treatment responses in thyroid cancer and should be reflected in AI models training and data sets.15,16 For example, Asian populations have a higher incidence of thyroid cancer, and genetic variants like the BRAFV600E allele are more prevalent in this ethnic group, which can increase tumor aggressiveness. While a model that accurately predicts genetic mutations like BRAFV600E may function well independent of racial mix, demographic representation still matters for validating the model across populations.17,18 Without subgroup-level analysis, it’s unclear whether performance is consistent across all groups, especially when other variables such as image quality, comorbidities, or access to care may vary and affect downstream outcomes. Thus, representation enables auditing for hidden disparities and ensures broader applicability. As such, to ensure accurate predictions for diagnosis, treatment, and prognosis, AI models must be trained and validated on data sets that reflect the epidemiology and significant influencing factors of thyroid cancer. Therefore, using diverse data sets is essential to improve model generalizability and optimize accuracy.
Failure to include diverse ethnic groups in AI model development can lead to biases that disproportionately affect underrepresented populations, reducing the accuracy and generalizability of these models. 11 In thyroid cancer, where incidence, genetic predisposition, and treatment responses vary significantly across ethnic groups, a lack of inclusivity can result in AI models that misclassify malignancies or provide suboptimal treatment recommendations. This systematic review examines how demographic factors, such as ethnicity, sex, and socioeconomic status, are represented in AI models for thyroid cancer screening, diagnosis, treatment, and prognosis. We also propose a framework for AI and machine learning (ML) experts to enhance model design by incorporating diverse ethnic data, ensuring more equitable and clinically relevant AI applications in thyroid cancer care.
Methods
This systematic review was registered on PROSPERO (ID: CRD42024519238). Results were reported in accordance with Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines and conducted in accordance with the Cochrane handbook.
Literature search
A literature search was conducted on EMBASE, PubMed, and Google Scholar from inception to January 06, 2024. The search strategy used a combination of MeSH terms and keywords related to thyroid cancer, bias, equity, diversity, inclusion, ML, and AI (Supplementary Appendix A1 and A2). The search strategy was designed and validated by a trained research librarian (R.S.) to identify relevant studies on AI models for thyroid cancer screening, diagnosis, treatment, and prognosis.
The inclusion criteria for screening consisted of publications that included information or directly compared with patient demographics when creating the AI/ML model and studies using ML/AI/predictive models for various aspects of thyroid cancer management. Studies using traditional rule-based natural language processing methods without any AI, ML, or predictive modeling components were excluded, as our focus was on models incorporating data-driven learning or predictive capabilities. Non-English publications, conference abstracts, and studies with AI models not involving human subjects or clinical data were omitted.
Screening
Selected papers underwent title and abstract screening, followed by full-text screening. Each article underwent screening based on the specified inclusion and exclusion criteria by at least two independent reviewers (S.G.B. or S.G.S. or M.M. or D.M. or C.H.). Discrepancies were resolved with discussion till consensus was reached or with consultation of a third-party expert (R.R. or E.G.). To assess inter-rater reliability during the abstract screening phase, Cohen’s kappa statistic was calculated based on the number of agreements and disagreements between the reviewers.
Data extraction
Data extraction was conducted using a standardized online form, distributed to six co-authors (R.R., S.G.B., S.G.S., M.M., D.M., and G.A.-O.) who participated in the data extraction process. This standardized form was designed to facilitate the collection of essential information from each study. The form included questions about key study details, methodological specifics, and details regarding the training and validation sets of the models mentioned in the studies, as well as the clinical application of AI models. For this review, clinical applications of AI models were categorized based on the stage of thyroid care targeted, including screening (e.g., early detection of nodules), diagnosis (e.g., malignancy prediction, cytology classification), treatment planning (e.g., surgical decision support), and prognosis (e.g., recurrence prediction and survival estimation). Each article underwent data extraction by two independent student reviewers.
Risk of bias
In this study, an adapted version of the Risk of Bias in Non-randomized Studies of Interventions tool was used to assess for bias contributing to methodological limitations, as based on PRISMA guidelines and the Cochrane handbook criteria. Each study’s risk of bias was evaluated by two independent reviewers, with discrepancies resolved through discussion or consultation with an expert reviewer. Based on the risk-of-bias assessment, studies were categorized as having low, moderate, or high risk of bias. High and moderate risk-of-bias studies were included in the analysis to encapsulate the extent of bias in study training and validation datasets.
Statistical analysis
Analyses were conducted using Python version 3.10 and R version 4.1.2. Descriptive statistics were used to summarize the demographic characteristics and study features, including counts and percentages for categorical variables, and means with standard deviations for continuous variables.
The distribution of ethnicities in AI model datasets was compared with both the U.S. incidence of thyroid cancer and the global incidence of thyroid cancer regardless of study origin, as AI models often included diverse international populations that were not limited to the region in which the study was conducted. Specifically, U.S. incidence data were derived from the Surveillance, Epidemiology, and End Results Program 18 databases. 19 This was done to assess the representativeness of training and validation sets used in AI models in reference to the epidemiology of disease on a regional and international scale.
Global prevalence estimates were calculated using thyroid cancer incidence data from Magreni et al. 12 (Table 1). For each ethnic group, male and female incidence rates were summed to determine the total incidence per 100,000 individuals (e.g., White: 6.3 + 20.4 = 26.7 per 100,000). These values were then aggregated across all groups to compute the total global incidence per 100,000, and the proportion for each ethnicity was determined by dividing the group-specific incidence by the overall total. The Magreni et al. study did not provide separate reporting for South Asian, Middle Eastern, or Pacific Islander populations. 12 Additionally, an “other” ethnicity category was included in the original dataset, accounting for 20.5% of cases, but was excluded from the comparative analysis as it could not be mapped to a specific ethnic group defined in this study. This exclusion was accounted for in the figure description to clarify that when included, the total would sum to 100%. Studies were then grouped broadly for comparison into “AI training and validation models” and “Clinical Applications.”
Strategies to Mitigate Bias in AI Applications to Thyroid Cancer
AI, artificial intelligence.
Results
A total of 1681 publications were identified, 21 duplicates were removed, and 1660 publications underwent title and abstract screening by two independent reviewers. Following title and abstract screening, a total of 863 papers underwent full-text screening by two independent reviewers, after which 199 articles progressed to the data extraction stage (Fig. 1). The Cohen’s kappa was 0.84, indicating strong agreement between reviewers. After adjustment for risk-of-bias, ultimately 197 studies were included in the analysis of this review (Supplementary Appendix A3).

PRISMA flowchart of systematic review methodology. PRISMA, Preferred Reporting Items for Systematic Reviews and Meta-Analyses.
Of the 197 papers included in this review (Supplementary Appendix A4), the overall risk-of-bias was considered low for 60.1% (n = 120) and moderate for 39.9% (n = 76). There were no studies where the risk-of-bias was deemed to be unclear (Supplementary Appendix A4).
From the included investigations, the study designs included: 163 retrospective, 22 prospective, 5 quasi-experimental, 4 cross-sectional, 3 observational, and 2 case–control studies. The studies employed various AI models. Most used deep learning approaches (n = 98), including convolutional neural networks, artificial neural networks, and deep neural networks. Others used traditional ML models (n = 54), such as decision trees (n = 12) and logistic regression (n = 17). A subset (n = 10) did not specify the model type.
The applications of these AI models were categorized into six primary areas: diagnosis (n = 133), screening (n = 2), risk stratification (n = 2), treatment (n = 3), prediction/prognosis (n = 47), and other uses (n = 10). The training and validation dataset sizes for the different uses of AI models were analyzed (Supplementary Appendix A5). The “other” category included applications such as automatic segmentation of high-risk populations (n = 3), prediction of malignancy in nodules (n = 3), analysis of diagnostic accuracy (n = 3), and prediction of malignancy based on radiation exposure (n = 1).
The average age of participants ranged from 44.7 to 55.6 years. The majority of studies included in this review were in China (study n = 124, participant n = 420,193), followed by the United States (study n = 26, participant n = 415,874). Across all AI applications, more female participants were included compared to males. The greatest discrepancy was observed in treatment models (6570 females vs. 2742 males) and prediction/prognosis models (3066 females vs. 1069 males). The smallest discrepancy in sex distribution was seen in risk stratification models (196 females vs. 92 males), likely due to the fewer number of studies in this category.
Ethnicity
The composition of AI training datasets directly impacts model performance and generalizability. In thyroid cancer, where incidence varies across ethnic groups, imbalanced datasets can lead to biased predictions. To assess representation, we analyzed the ethnic distribution of participants in AI model training and validation datasets and compared these to real-world thyroid cancer prevalence. We then examined how these disparities translate to clinical applications, identifying gaps that may affect AI-driven decision-making in thyroid cancer care.
AI training and validation models
The distribution of described ethnicities can be seen in Figure 2. Scarce studies included Middle Eastern, Pacific Islander, and South Asian ethnicities in AI models. Additionally, several studies did not specify the ethnicity of participants, which was most notable in models for thyroid cancer diagnosis (8.3%) and prediction/prognosis (8.8%).

Ethnic distribution of studies involving participants in the United States relative to distribution of 2014 ethnicities in the American population and thyroid cancer prevalence in America. American ethnicity distribution and thyroid cancer prevalence rates from Weeks et al., 2018. 19
The AI training and validation models that included ethnicity as a variable were analyzed. For example, in models created for the diagnosis of thyroid cancers, 76.0% (n = 8 studies) of participants were East Asian, while the percentage of other ethnicities, such as Black was 5.1% (n = 4 studies) and White was 2.3% (n = 2 studies). Similarly, for screening and risk stratification, East Asian participants constituted 100% of the study population. The treatment category showed a more diverse distribution, with 44.4% East Asian, 33.3% Hispanic, 11.1% White participants, while 11.2% were unspecified. There was no representation of Black participants in the treatment category (Fig. 4).
Ethnic representation in thyroid cancer AI models demonstrates significant disparities when compared with both the overall U.S. population and thyroid cancer prevalence by ethnicity in America (Fig. 2). White populations are underrepresented in AI models (26.8% each) relative to their prevalence among the U.S. thyroid cancer cases (67.7%) and the U.S. general population (67.6%). A similar trend is seen for Hispanic individuals, who comprise of 15.0% of thyroid cancer cases in the United States but are notably underrepresented in AI model datasets (4.9%).
Disparities are also seen when comparing ethnic distribution to both global thyroid cancer prevalence and the overall global population distribution (Fig. 3). White and Black populations are overrepresented in AI models (26.8% each) relative to their global thyroid cancer prevalence (20.7% and 11.3%, respectively), with Black populations also exceeding their global population distribution (16.7%). In contrast, East Asian representation in AI models (14.6%) is lower than its thyroid cancer prevalence (18.7%) and its global demographic share (22.2%). Furthermore, Indigenous populations (7.3%) appear in AI models at rates lower than their global thyroid cancer prevalence (11.7%) but higher than their global demographic proportion (5.3%). South Asian, Middle Eastern, and Pacific Islander populations could not be assessed for thyroid cancer prevalence due to a lack of segmented data for global thyroid cancer prevalence. However, their representation in AI models (4.9%, 2.4%, and 12.2%, respectively) is notably lower than their overall global population distribution (25.4%, 5.6%, and 0.1%, respectively). Finally, Hispanic populations, comprising 8.3% of the global population and 16.9% of thyroid cancer cases, are notably underrepresented in AI models (4.9%). These discrepancies highlight potential biases in AI model training data and underscore the need for more diverse datasets to enhance global applicability and fairness in AI-driven thyroid cancer management.

Graph on distribution of participants’ ethnicity across different AI applications for thyroid cancer care relative to the global population and global thyroid cancer prevalence. Data on global ethnic distribution were obtained from the World Population Data Sheet,
20
and global incidence of thyroid cancer data were sourced from Magreni et al.
12
Segmental data on global thyroid cancer prevalence were not reported for South Asian, Middle Eastern, or Pacific Islander ethnicities. For panel

Distribution of participants’ ethnicity across different AI applications for thyroid cancer care:
Clinical applications of AI models
Ethnic representation in thyroid cancer AI models demonstrates notable disparities across various clinical applications (Fig. 3). Particularly, East Asians and Hispanics are disproportionately represented in screening and treatment studies, while Black, Indigenous, and other ethnic groups are underrepresented across all categories.
Factors to train models
The frequency of factors used in AI model training was analyzed (Fig. 5), categorized into patient, image, tumor, and sonographic characteristics. Among patient factors, age (n = 76) and sex (n = 68) were the most frequently considered, while race/ethnicity (n = 13), marital status (n = 7), and geographical location (n = 18) were rarely included. Image features such as aspect ratio (n = 50) and calcification (n = 29) were common, whereas radiomics (n = 9) and image composition (n = 17) were less frequently used. Tumor characteristics such as shape (n = 33) and size (n = 38) were the most considered, followed by TNM stage (n = 23) and thyrotropin level (n = 6). Models trained on ultrasound images prioritized modality-specific features such as ultrasound composition (n = 31) and echogenicity range (n = 40) over patient demographics.

Demographics of participants in thyroid cancer AI models including
Validation models largely mirrored training sets, with 128 studies reporting identical factor distributions. However, key clinical variables, including genetics, environmental exposure, comorbidities, and socioeconomic status, were infrequently incorporated. Genetics were mentioned in 15 studies, environmental exposure in 3, and comorbidities in 14. Dietary and substance use factors were included in only 2 studies each, while radiation exposure was considered in 9. Socioeconomic status, a critical determinant of thyroid cancer outcomes, was assessed in just 2 studies.
Discussion
This review of 197 studies revealed the growing application of AI in thyroid cancer management. The analysis demonstrated an emphasis on diagnosis (67.5% of studies) and prediction/prognosis (23.8%), with deep learning, neural network, and random forest emerging as predominant techniques. The majority of studies were retrospective, primarily conducted in China (n = 124) and the United States (n = 26). While AI shows promise across the thyroid cancer care continuum, there are significant disparities in dataset sizes, demographic representation, and the factors considered in algorithm development.
Interestingly, while White and Black individuals are the most represented ethnicities in thyroid cancer AI models overall (Fig. 2), East Asians appear to be the most represented group in specific AI applications such as risk stratification and treatment (Fig. 3). These representation disparities may stem from a combination of systemic and structural factors. First, data availability varies significantly by region, with high-volume AI research often concentrated in countries with robust digital infrastructure and centralized health records. Second, research funding priorities in high-income countries may drive AI model development based on locally available datasets, even when those datasets are not demographically representative. Third, stringent data privacy regulations, such as Health Insurance Portability and Accountability Act in the United States and General Data Protection Regulation in Europe, may limit international data sharing and collaboration, hindering the development of multicenter, diverse datasets. Finally, there is currently no standardized framework guiding researchers to consider demographic representativeness in relation to disease epidemiology during AI model design.
Moreover, other patient characteristics such as age (n = 76) and sex (n = 68) were frequently considered in thyroid cancer AI models. Thyroid cancer epidemiology varies by age, country, and sex, with a mean diagnosis age of 51 years and three times as many women affected as men. 21 Hormonal differences may influence tumor behavior in females, while younger patients often have distinct genetic drivers.10,15,16 Current models generally reflect these trends, with more data from younger individuals and females. To improve generalizability, future models should incorporate broader representation across countries and age groups or be tailored to specific populations. Additionally, concept drift, the change in statistical properties of target variables over time, must be considered in model development. 22 This can be addressed through periodic validation, monitoring performance metrics, or tools like the population stability index. When drift is detected, models may require retraining or recalibration to maintain clinical reliability.
A notable concern highlighted in many studies is the potential for selection bias in the radiological or ultrasound images used to train AI models. Often, only the largest or most suspicious lesions are fed into the model, which may skew the model’s learning process.23,24 Relying on surgical pathology as the “gold standard” also excludes less suspicious, likely benign lesions that were not resected, introducing a bias based on initial human judgment. This selective inclusion can limit the model’s ability to accurately identify and assess smaller or subtler lesions, reducing its diagnostic utility. Additionally, data imbalances within the training sets, with a higher proportion of certain lesion types or limited cases with specific conditions, can lead to biased models that overlook or misclassify less common but clinically important variations.25–27 Furthermore, this approach can contribute to overfitting, where the model overlearns specific features and mistakes irrelevant details for significant patterns.
There was limited inclusion of other established risk factors for thyroid cancer in the reviewed AI models, with only 15 studies incorporating genetic data, 3 addressing environmental exposures, and 2 considering socioeconomic status (SES). This represents a substantial gap, particularly given the well-documented influence of SES on thyroid cancer outcomes. For instance, individuals with lower income, limited education, or lack of insurance are more likely to present with advanced-stage disease and experience prolonged delays between diagnosis and treatment. 28 A multisite retrospective study by Siu et al. further demonstrated that SES is closely tied to both the stage at diagnosis and overall disease incidence, with both low and high SES associated with more advanced presentation and higher incidence rates, respectively. 29 Similarly, genetics and environmental exposures are critical in shaping thyroid cancer risk and progression, and their underrepresentation in AI algorithms may constrain the predictive power and clinical utility of these models.30–32
Studies have shown that AI models for conditions like melanoma, where skin lesions of Caucasian people are predominantly used, have been extensively scrutinized for bias, leading to improvements in their design and application.24,27 Similarly, AI algorithms for diabetic retinopathy have faced criticism for underrepresenting minority groups, affecting the accuracy of retinal disease detection.25,26 Cardiovascular disease models have been scrutinized due to sex bias, often resulting in less accurate predictions for women compared with men.33–35 While these examples highlight the importance of addressing bias in AI development, similar scrutiny is less common for thyroid cancer AI models. A lack of representation in AI models can lead to biases, making the models less accurate and generalizable to the broader population. Matching the input of AI models to the real-life prevalence and distribution of thyroid cancer ensures that the models are robust and effective, especially across different demographic groups. The outcome of well-trained and diverse models can not only improve patient care and outcomes but also contribute to improving clinical efficiency. Such diversity is crucial for the immediate needs of the current patient population, but also prepares the models to adapt to future demographic shifts, ensuring sustained relevance and accuracy in clinical applications.36–38
Reducing bias in AI algorithms for medical applications that lack equal representation of populations can be particularly challenging. Highlighting the importance of shared worldwide data is crucial to addressing this issue. By emphasizing the need for shared worldwide data and employing these strategies, it is possible to reduce bias in AI algorithms and create more equitable and accurate medical applications. This can be achieved through the incorporation of several strategies to mitigate bias under these circumstances (Table 1). While transfer learning and data augmentation are standard practices in deep learning workflows, particularly for convolutional neural networks (CNNs), current evidence suggests that these techniques alone do not eliminate demographic bias. They may help with model generalization, especially when original datasets are small, but do not fully correct for underrepresentation of specific populations unless those populations are meaningfully represented in the source or augmented data.39,40 Moreover, while we propose certain strategies, such as global data sharing, routine audits, and inclusive model training, we acknowledge that these approaches may require substantial technical, financial, and regulatory resources. This may limit their feasibility in lower-resource settings or institutions without infrastructure to support large-scale data governance and model maintenance.
Limitations and future directions
This systematic review has several limitations that may impact the validity of its findings. The literature search was limited to studies published in English, potentially excluding relevant studies in other languages (n = 12). Additionally, the database search and subsequent reference screening may have led to the omission of pertinent studies available in other databases.
Furthermore, while this review aimed to evaluate the representativeness and equity of AI models for thyroid cancer, we acknowledge several limitations in our perspective and approach. First, the majority of included studies originated from China, yet our author group does not include contributors from that region. This may limit the cultural and clinical contextualization of the findings and introduces the possibility of interpretive bias. Second, although we advocate for improved demographic representation in AI model development, we recognize that global generalizability may not always be a practical or necessary goal. In many health care settings, locally optimized models may be more appropriate, especially when tailored to the epidemiology and infrastructure of a specific region.
Based on the findings of this systematic review, further studies are needed to develop and validate AI models using more balanced and comprehensive datasets that include diverse demographic and clinical factors. Many of the data sources in this review were from single sites, which do not capture the full variety of patients within a broader geographical area. Future studies should aim to include prospective, multi-site data collection to ensure models are more representative and capable of addressing the diverse needs of thyroid cancer patients. In addition, future studies should report model performance stratified by demographic subgroups, including ethnicity, sex, and age. This would provide critical insights into whether existing disparities in representation also translate into disparities in model accuracy or clinical utility. Additionally, there is a need for longitudinal studies to assess the long-term performance and equity of these models in real-world settings. By addressing these gaps, future research can contribute to the development of more robust, equitable, and generalizable AI models for thyroid cancer care.
Conclusion
This systematic review provides a comprehensive analysis of the current state of AI models for thyroid cancer care, highlighting areas for improvement. The findings underscore the importance of ensuring training and validation data sets used for AI development are representative of the epidemiological distribution of participants for which the AI model is being created. Doing so will enhance the accuracy of the model for diagnosis and prediction as well as improve its generalizability. Moreover, while current models predominantly focus on age, sex, tumor, and ultrasound characteristics, there is a need to investigate the necessity for a broader range of factors such as socioeconomic status, race/ethnicity, and marital status. Furthermore, the review identifies the potential for selection bias and data imbalances in existing studies, emphasizing the necessity of balanced and comprehensive datasets. By addressing these limitations and incorporating more diverse and representative data, future AI models will hopefully be able to improve their diagnostic and prognostic utility across various clinical settings. Future studies should focus on creating inclusive datasets, considering concept drift, and ensuring the regular revisions of AI models on an as-needed basis.
Authors’ Contributions
Conceptualization: R.R., A.E., and J.M.S. Data curation: R.S. Formal analysis: R.R., E.G., and L.C. Methodology: R.R., E.G., S.G.B., S.G.S., M.M., D.M., C.H., G.A.-O., and R.S. Project administration: R.R. Supervision: L.C., E.J.P., N.E.W., J.D.W., A.E., and J.M.S. Writing—original draft: R.R., E.G., N.E.W., A.E., and J.M.S. Writing—review and editing: R.R., E.G., S.G.B., S.G.S., M.M., D.M., C.H., G.A.-O., L.C., E.J.P., N.E.W., J.D.W., A.E., and J.M.S.
Footnotes
Author Disclosure Statement
All authors have no competing interests or disclosures.
Funding Information
This study did not receive any funding.
Supplemental Material
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
