Abstract

KEY POINTS
In a large multicenter validation cohort (n = 4572), a deep learning system deployed as a secondary check (“goalkeeper”) for nodules already flagged for biopsy by radiologists successfully identified 86.8% of benign nodules, theoretically reducing the unnecessary fine-needle aspiration rate from 68.5% to 9.1%.
In a direct comparison (n = 260), the artificial intelligence (AI) system achieved a benign nodule identification rate of 81.4%, significantly outperforming both junior (40%) and senior (55%) radiologists, with an area under the curve of 0.88 versus 0.43 (P = 0.002) and 0.63 (P = 0.003), respectively.
The AI system misclassified 8.6% of malignant nodules as benign. Although the majority were low-risk papillary thyroid microcarcinomas, the presence of BRAF V600E mutations (74%) and extrathyroidal extension (20.3%) in these missed cases necessitates careful active surveillance protocols for AI-benign nodules rather than outright dismissal.
SUMMARY
Background
The global incidence of thyroid cancer has tripled over the past three decades, a phenomenon largely driven by the overdiagnosis of small, indolent papillary thyroid cancers. 1 Although high-resolution ultrasonography has increased detection sensitivity, the specificity in distinguishing benign from malignant nodules remains suboptimal. Current risk-stratification systems, such as the American College of Radiology Thyroid Imaging Reporting and Data System (ACR TI-RADS) and the American Thyroid Association (ATA) guidelines, rely on subjective interpretation of sonographic features. 2 This reliance results in significant interobserver variability and a high rate of diagnostic thyroidectomies and unnecessary fine-needle aspirations (FNAs). This diagnostic cascade imposes substantial economic burdens, estimated to reach $3.5 billion annually in the United States by 2030, and patient anxiety.3,4 Artificial intelligence (AI)—specifically, deep learning (DL)—offers a potential solution by standardizing feature extraction. 5 The study by Ni et al. evaluates an AI system not as a primary screening tool, but as a “goalkeeper” to reevaluate nodules initially deemed suspicious by radiologists, aiming to curb unnecessary interventions. 6
Methods
This retrospective multicenter study used the ITS100 (MedAI Technology) DL system to reevaluate thyroid nodule ultrasound images. The study analyzed a validation cohort (Dataset 1) of 4572 nodules from two institutions and a comparison cohort (Dataset 2) of 260 nodules from a third center. Inclusion was restricted to adults with nodules that had already undergone FNA based on radiologist suspicion (standard of care) and had definitive pathologic results (surgical histology or Bethesda II cytology with >12 months of stable follow-up). Nodules with indeterminate cytology lacking surgical confirmation or poor image quality were excluded. The primary outcomes were the theoretical reduction in unnecessary FNAs (specificity) and the rate of missed malignancies (false-negative rate). In the comparison cohort, AI performance was benchmarked against 21 radiologists (10 senior, 11 junior) in a blinded review.
Results
The primary validation cohort consisted of 4572 thyroid nodules from the same number of patients, with a mean age of 49.4 years and a female predominance of 75.3%. In terms of sonographic risk stratification, the vast majority of nodules were classified as low suspicion (76.7%) by the American ATA classification, while intermediate suspicion (16.0%) and high suspicion (7.3%) nodules constituted smaller proportions of the sample. Cytopathologic assessment revealed a bimodal distribution, with 3106 nodules (67.9%) categorized as benign (Bethesda II) and 971 (21.2%) as malignant (Bethesda VI); fewer cases fell into the indeterminate or suspicious categories (Bethesda III–V). Final diagnoses confirmed that 68.5% of the nodules were benign and 31.5% were malignant.
In this validation cohort, the AI system correctly identified 2719 (86.8%) of the benign nodules, thereby reducing the theoretical unnecessary FNA rate from 68.5% to 9.1% (panel A of Figure). The system achieved an area under the curve of 0.91, with a sensitivity of 91.4% and a specificity of 86.8%. Performance was notably stratified by risk level; the AI successfully avoided biopsy in 73.7% of ATA low suspicion nodules compared to only 26.9% of ATA high suspicion nodules. Regarding safety, the system misclassified 123 malignant nodules (8.6%) as benign. Subgroup analysis of these missed malignancies showed that 56.9% were papillary thyroid microcarcinomas, 74% harbored the BRAF V600E mutation, and 20.3% exhibited extrathyroidal extension. In the separate comparison cohort of 260 nodules, the AI significantly outperformed human readers, identifying 81.4% of benign nodules, compared to 40% for junior radiologists (P = 0.002) and 55% for senior radiologists (P = 0.003) (panel B of Figure).

The impact of artificial intelligence (AI) on thyroid nodule management.
Conclusions
The AI system demonstrated high efficacy as a clinical “goalkeeper,” significantly outperforming radiologists in specificity and potentially preventing nearly 90% of unnecessary biopsies in the study cohort. However, the misclassification of a subset of malignancies underscores that AI-benign designations should trigger active surveillance rather than clinical discharge, particularly in those with ATA intermediate- and high-risk nodules.
COMMENTARY
The integration of AI into thyroidology is transitioning from theoretical promise to clinical application. 5 As an endocrinologist focused on both AI technology and health care resource stewardship, I view the study by Ni et al. as a significant contribution to the literature on diagnostic de-escalation. 6 The “goalkeeper” paradigm, using AI as a secondary filter for nodules already flagged for biopsy, addresses one of the most pressing challenges in our field: the overdiagnosis of indolent disease and the resulting diagnostic cascade of unnecessary interventions.1,4
The study highlights a substantial disparity in specificity between the AI systems and human radiologists (81.4% vs. 40–55%). This gap quantifies the defensive medicine bias often observed in practice: clinicians, wary of missing a cancer, may lower their biopsy threshold, upgrading benign-appearing nodules to intermediate risk to ensure tissue sampling. 7 The AI, adhering strictly to learned probability densities without liability fears or cognitive fatigue, acts as a rationalizing force in this workflow.
However, moving beyond validation studies reveals a complex landscape of real-world implementation. A growing number of institutions are actively integrating AI assistive systems into their daily thyroid nodule interpretation and reporting workflows. As these sites begin generating dual scores (human TI-RADS and AI TI-RADS), we face an experience bifurcation. Novice clinicians risk automation bias, deferring to the algorithm without question, while experts may suffer from anchor bias, dismissing valid AI downgrades due to established practice patterns. 8 Furthermore, we must recognize that these are operator-dependent algorithms. An AI prediction is only as accurate as the image frame selected by the human operator; poor acquisition technique inevitably yields unreliable output, potentially generating false negatives. 9 Additionally, the current generation of AI, including the system used in this study, largely relies on unimodal imaging data. It analyzes pixels but lacks the clinical context, such as symptoms, family history, or patient age, that a human clinician synthesizes. 5 This limitation likely contributes to the system’s blind spots regarding biological aggressiveness.
Finally, the study’s 8.6% false-negative rate is a necessary caution. Therefore, identifying discordance should not be viewed as a failure, but as a specific safety signal. If a radiologist recommends a biopsy and the AI suggests benignity, the compromise is neither immediate discharge nor reflexive biopsy, but discordance-based active surveillance. This approach leverages the AI’s high specificity to curb immediate intervention, particularly for ATA low-risk nodules, while maintaining a safety net for the minority of “AI-missed” malignancies, which the study noted frequently harbored BRAF V600E mutations or extrathyroidal extension. The “goalkeeper” does not replace the clinician; it ensures that intervention is reserved for the most appropriate cases.
