Abstract
Background:
Artificial intelligence (AI) chatbots are increasingly being used by patients to obtain medical information. Comparison between platforms with specialty-specific physician assessment remains limited. This study compares the quality, factual accuracy, readability, and consistency of responses generated by four publicly available AI chatbots when answering patient-centered questions about thyroid radiofrequency ablation (RFA).
Methods:
We conducted a cross-sectional analysis of chatbot-generated responses using 20 standardized clinical questions about thyroid RFA. Responses from ChatGPT-4, Gemini, Copilot, and Perplexity were evaluated by six blinded physician reviewers experienced in thyroid RFA using 5-point Likert scales for global quality and factual accuracy. Higher Likert scale scores indicated better performance. Readability and response length were analyzed with established metrics. Statistical significance was defined as p < 0.05.
Results:
Gemini achieved the highest mean scores for global quality (4.08 ± 0.87) and accuracy (3.76 ± 1.05), with significantly better performance than ChatGPT and Copilot (p < 0.005). ChatGPT responses were significantly longer and more readable. Score variability across questions was lowest for Gemini. Copilot and Perplexity ranked lowest across most domains. Question-level analysis identified specific prompts that best discriminated between platforms.
Conclusions:
AI chatbot performance varied across platforms for thyroid RFA queries. Chatbots were generally reliable for straightforward factual information but were less dependable for judgment or context-dependent assessments. These AI tools should supplement, not replace, clinician-vetted patient education and institutional materials.
Keywords
Introduction
Obtaining medical information from a patient perspective has markedly changed in recent history due to advancements in technology. Patients no longer rely solely on brochures, clinician discussions, or traditional internet searches. Rather, with the rise of artificial intelligence (AI) and large language models (LLMs), patients are seeking answers to their medical questions using publicly available chatbots. As of 2025, nearly 60% of adults in the United States report usage of some form of technology supported by AI. 1 ChatGPT, Gemini, Copilot, and Perplexity exist as some of the most widely used AI chatbots. 2 All four of these chatbots have platforms that are free to the public, accessible through basic web browsers, and require no subscription. Built using LLMs, which are large datasets trained to mimic human language patterns, chatbots are enticing to patients as they offer both easy accessibility and real-time responses. Chatbots have the potential to offer language more readable and comprehendible than physician-derived responses or resources. Some studies have even reported higher levels of empathy from AI-generated responses than physician responses to prompts. 3 Additional work has demonstrated that AI can lower required reading levels without reducing informational depth. 4 Despite these reported benefits, there are concerns in health care about AI accuracy and reliability for patient education. 5 Research has identified fabricated content, termed “AI hallucination,” and blatant omission of important details pertinent to medical advice. These represent reasonable concerns for medical providers and patients. 6 As such, understanding how these models respond is an important matter of public health.
Prior literature has begun to explore application of LLMs to thyroid diseases. Campbell et al. evaluated ChatGPT responses to thyroid nodule questions and found that content was generally appropriate but often presented at reading levels above recommended patient-education targets. 7 Recently, Guo et al. compared ChatGPT-4 responses with those of junior and senior thyroid specialists for common thyroid-related questions and reported higher ratings for accuracy, comprehensiveness, compassion, and satisfaction for ChatGPT. This group noted that performance on complex reasoning remains uncertain. 8
The goal of this study was to evaluate the accuracy, quality, readability, and consistency of AI chatbot responses to standardized questions about thyroid radiofrequency ablation (RFA). Thyroid RFA is a minimally invasive outpatient procedure approved by the Food and Drug Administration in 2018 to treat primarily benign thyroid nodules. As an alternative to surgery, RFA effectively reduces nodule volume with the goal of relieving symptomatic compression or improving cosmesis. Major complications are rare in RFA, and local anesthesia is typically used over general anesthesia, resulting in a shorter recovery period than traditional thyroid surgery. Uniquely, public awareness of RFA is limited even with its favorable outcomes, low complication risk, and growing support across the country. Low baseline public knowledge of RFA presents a valuable area for analyzing AI responses to patient questions regarding this therapy.
As chatbots continue to provide information on procedures like thyroid RFA, and studies increasingly reveal inconsistency and unreliability in AI outputs,9–11 specialty-specific evaluation of AI material becomes more necessary. Using subspecialty thyroid surgeons who were experienced in thyroid RFA as blinded evaluators, we hypothesized that significant differences would emerge between platforms in accuracy, readability, and consistency when answering standardized patient questions about thyroid RFA.
Materials and Methods
This study evaluated the performance of four AI chatbots: ChatGPT-4 (OpenAI), Gemini (Google), Copilot (Microsoft), and Perplexity (Perplexity.ai). A total of 20 patient-centered questions were developed based on commonly asked clinical questions (Table 1), and each chatbot was queried with the same questions in a fixed, sequential order. Responses were anonymized and randomized for scoring by expert reviewers, which included six board-certified otolaryngologists who are high-volume thyroid surgeons and not only perform but also teach thyroid RFA at a national level. To our knowledge, this represents one of the largest panels of subspecialty physician reviewers used in a comparative AI chatbot evaluation to date. Each reviewer independently scored responses using structured metrics, with all reviewers blinded to chatbot identity.
Initial Prompt and the 20 Patient Questions Used for Chatbot Evaluation
The same initial prompt and set of 20 questions were submitted to each chatbot. Questions were designed to reflect common patient concerns regarding thyroid radiofrequency ablation.
Responses were reviewed by each rater independently, and two domains were assessed for each response: Global Quality Score (GQS) and factual accuracy. Both were scored using 5-point Likert scales, with higher scores corresponding to higher quality and correctness. All reviewers rated each response without coordination or consultation. This is consistent with prior evaluations of online and AI-generated health information.5,11 In addition, objective readability metrics were calculated for each AI response. These included Flesch Reading Ease Score and Flesch–Kincaid (FK) Grade Level, which estimate reading difficulty based on sentence and word structure and total word count.12,13 Standard formulas were applied uniformly across all chatbot outputs.
All statistical analyses were performed in R (v4.5.1), and figures were made using GraphPad Prism (v10.6.1). Descriptive statistics computed the mean and standard deviation (SD) of GQS and accuracy for each chatbot. Shapiro–Wilk tests were used to assess normality of the data. Given the ordinal and non-parametric nature of the data, Kruskal–Wallis tests were used to evaluate differences between chatbots. Significant results were followed by pairwise comparisons using Dunn’s test with Bonferroni correction. Pairwise comparisons of readability metrics were conducted using Tukey’s Honest Significant Difference (HSD) following one-way analysis of variance. Pairwise comparisons of word count were evaluated using Mann–Whitney U tests with Bonferroni correction. Pearson correlations were used to assess associations between Flesch Reading Ease and FK Grade Level for each chatbot. Effect sizes were calculated using epsilon squared (ε2) to quantify observed differences across chatbots. SDs were calculated across the 20 items for each chatbot to assess consistency. Identifying prompts producing wide variation in chatbot scores allowed us to analyze question-level discrimination. A mean rank analysis was conducted to compare chatbot performance ranks across metrics. Statistical significance was defined as p < 0.05. This study did not involve human subjects or protected health information and was exempt from Institutional Review Board review.
Results
Descriptive statistics
There was variability between platforms in both quality and accuracy. Mean GQS and accuracy scores differed across the four chatbot platforms. Gemini had the highest average scores for both quality using GQS (mean 4.08 ± 0.87) and accuracy (3.76 ± 1.05). Perplexity ranked second, followed by ChatGPT and Copilot. Copilot scored lowest across both metrics (Table 2). Kruskal–Wallis tests confirmed that chatbot responses differed significantly for both quality and accuracy (GQS: p < 0.001, accuracy: p < 0.001 [Table 3]). Dunn’s post hoc tests with Bonferroni correction identified specific pairwise differences. Gemini significantly outperformed ChatGPT (GQS: p = 0.0038, accuracy: p = 0.0032) and Copilot (GQS: p < 0.001, accuracy: p < 0.001). Perplexity also outperformed Copilot for GQS (p = 0.0115). Comparisons between ChatGPT and Perplexity did not reach significance for either score. Effect size analyses using epsilon squared (ε2) indicated medium effects by η2 conventions, showing that observed differences were meaningful rather than trivial. 14 GQS had ε2 = 0.078, and Accuracy had ε2 = 0.075 (Table 4).
Mean GQS and Accuracy for Each Chatbot
GQS is scored on a 5-point Likert scale (1 = very poor, 5 = excellent). Accuracy scored from 1 (factually incorrect) to 5 (completely correct).
GQS, Global Quality Scale; SD, standard deviation.
Pairwise Comparisons of Global Quality Score and Accuracy Between Chatbots Using Dunn’s Test with Bonferroni Correction
Global Kruskal–Wallis tests were significant for both GQS and Accuracy (p < 0.001). Reported p-values reflect Bonferroni-adjusted pairwise comparisons using Dunn’s test.
Dunn, Dunn’s multiple comparisons test; GQS, Global Quality Scale.
Kruskal–Wallis Effect Size (ε²) for Global Quality Score and Accuracy
Epsilon squared (ε²) values were calculated to estimate effect sizes from Kruskal–Wallis tests. Interpretations follow established thresholds: >0.14 = large, 0.06–0.14 = moderate, <0.06 = small. Both GQS and Accuracy demonstrated moderate effect sizes.
GQS, Global Quality Scale.
Score variability and question discrimination
Score stability was evaluated using the SD of scores across the 20 questions (Table 5). Gemini had the most consistent GQS ratings (SD 0.219). Followed by Copilot, Perplexity, and ChatGPT, Gemini had the lowest variability in accuracy (SD 0.380), and Copilot had the highest (0.417). Gemini’s scores were higher on average but also more consistent across different clinical scenarios. Questions 6, 13, and 20 showed the greatest variability in accuracy and addressed: (Q6) whether thyroid RFA requires anesthesia, (Q13) whether patients must take medications after the procedure, and (Q20) whether thyroid nodules disappear completely following RFA (Table 1). These three questions were the most effective at distinguishing differences in answer quality.
Standard Deviation of Accuracy and GQS Scores Across All Questions by Chatbot
Lower standard deviation reflects greater consistency in chatbot responses. Gemini had the most consistent GQS scores, while Copilot had the highest variability in Accuracy scores.
GQS, Global Quality Scale; SD, standard deviation.
Readability and length metrics
Objective readability metrics varied across platforms. Higher Flesch Reading Ease scores indicate simpler text, and scores differed significantly between chatbots (Fig. 1A). ChatGPT-4 yielded the most readable responses (mean score 45.5), while Gemini had the least readable output (28.2). Tukey’s HSD tests showed that ChatGPT-4 had significantly higher Flesch Reading Ease scores than Gemini (mean difference –17.30, p < 0.001) and Copilot (–12.84, p = 0.0005). No significant differences were found among the remaining comparisons.

Readability and length of chatbot responses. Each chatbot provided answers to 20 patient questions (n = 20 per chatbot per metric). Data are shown as Tukey box-and-whisker plots (center line, median; box, interquartile range [IQR]; whiskers, 1.5×IQR) with individual responses overlaid. Group differences were assessed with Kruskal–Wallis tests with Dunn’s post-hoc correction. (
FK Grade Level scores followed a similar trend, but higher FK Grade levels reflect increased reading difficulty (Fig. 1B). Gemini had the most complex output (mean FK 13.9), while ChatGPT-4 produced language that was easiest to read (mean FK 10.7). Across platforms, higher Reading Ease was strongly associated with lower FK Grade (r = –0.76 to –0.94, all p < 0.01).
Word counts differed significantly across platforms (Fig. 1C). ChatGPT-4 produced the longest responses (mean word count 365.7), and Copilot produced the shortest (124.0). Pairwise comparisons with Bonferroni correction showed that ChatGPT-4 and Gemini generated significantly longer responses than Copilot (p < 0.0001). Perplexity also produced significantly longer responses than Copilot (p < 0.001).
Consistency across questions
Figures 2 and 3 shows chatbot performance trends across all 20 patient questions. Gemini demonstrated the most consistent performance for both GQS and accuracy, with relatively narrow score bands and fewer outliers. ChatGPT and Copilot exhibited greater score variability across the question set, with individual outliers lowering their average rank. These fluctuations contributed to the lower consistency and overall rank for these models.

Global Quality Score and accuracy ratings for AI chatbot responses to patient-centered questions about thyroid radiofrequency ablation. (

Accuracy ratings (1–5) for the three questions with the largest between-chatbot variance. Panels show per-question accuracy for
Discussion
LLMs have recently developed AI chatbots that are often used by patients to access real-time information for answering medical questions. Answers are often delivered rapidly and in an easily digestible manner. While early application of AI to medicine was constrained by logical rules, LLMs helped fast-track the transition from template behavior to dialogue. They draw from broad data that include both scientific content and layperson content while operating through likelihood and probability. Models are extremely capable3,4 and have even been shown to pass the board exam for the American Board of Otolaryngology Head and Neck Surgery. 15 However, performance by AI models when responding to head and neck cancer queries varied considerably depending on the type of question and complexity of the topic.15,16
Published studies of medical chatbots often focus on a single platform or output. Early LLMs such as ClinicalGPT and ChatDoctor were trained on domain-specific data but lacked reviewer evaluation.17,18 Clinical testing and formal scoring, accounting for patient complexity, were much needed.19,20 We previously found that ChatGPT response quality varied significantly with prompt structure for thyroid nodules and obstructive sleep apnea questions.7,21 Given increasing reliance on AI-generated health care advice, we evaluated the accuracy, quality, readability, and consistency of chatbot responses to standardized thyroid RFA questions. Direct comparisons between platforms using identical clinical questions remain scarce, and direct comparisons between platforms using physician reviewers are even more uncommon. In studies that have done so, the topics span hypertension education, 5 medication safety and drug–drug interaction, 6 postoperative questions following functional rhinoplasty, 22 and patient-education content for aesthetic facial plastic surgery. 23 All of these studies report meaningful differences in readability and factual accuracy between models such as ChatGPT and Gemini.
In the current study, Gemini outperformed ChatGPT, Copilot, and Perplexity across most domains. ChatGPT demonstrated higher readability and accuracy scores, but it was not superior in response length or completeness.
Our study addresses a key gap in the literature, which compares how chatbots fare in head-to-head comparisons when graded by experts in a specialized field. Directly comparing multiple AI chatbot queries about thyroid RFA allowed our team to assess AI response and delivery of information. The incorporation of surgeons who are experts in RFA as blinded reviewers added a level of academic rigor not previously seen. We found that ChatGPT responses were longer and had higher readability scores than the other platforms, while Gemini responses tended to be more concise and were rated higher for factual accuracy. Copilot outputs were variable in quality and often brief, consistent with prior benchmarking studies.11,15,16,24 This suggests that certain chatbots may be appropriate to supplement clinician-provided education on thyroid RFA.
Furthermore, RFA technology, technique, and indications continue to evolve in moving-shot technique, hydrodissection strategies, device capabilities, and selection criteria. LLM and chatbot training systems are inherently bound by training cutoff dates that lag behind current practice. These lags are not necessarily a defect of AI but a consequence of rapid medical innovation outpacing model updates. Differences in training, retrieval, and alignment can also influence factual accuracy, stability, and response style in medical domains.17,19,24 Models incorporating frequent updating or retrieval generation may capture evolving procedural indications better compared with models optimized for conversation generating more language but less focused responses. The logic by which these platforms generate information (training data, retrieval policies, and reinforcement heuristics) is also proprietary and undisclosed. The “black-box” design can limit interpretability and external verification, and material could be modified without notice, curbing reproducibility and trust. RFA experts continue to remain the most reliable source for up-to-date guidance, while chatbots should supplement counseling for thyroid RFA.
Our data represent a pattern seen across medical literature on LLMs, as models performed well on basic factual recall but were not reliable when clinical judgment was required.25,26 Low-scoring categories included post-procedural care, complications, or decision-making trade-offs. Questions 6, 13, and 20 had the highest variability between chatbots. Fact-based or definitional items provided uniform scoring. Chatbots may address the straightforward facts of RFA (what it is, how it is performed), but they are less reliable for judgment or context-dependent assessments.
Gemini had the lowest SD, meaning the most consistent performance, in both quality and accuracy, while ChatGPT exhibited greater variability with a wider spread of scores. This finding is similar to other evaluations of head and neck content—where chatbots perform well on average but continue to generate responses with low factual content.11,16 Clinicians should be aware that reliability is not uniform, and caution should be used when relying on AI tools to support pre-procedural education. For instance, even the highest-performing platform (Gemini) did not reach perfect accuracy. No chatbot achieved “completely correct” (5/5) ratings across questions, and relative to clinician-derived current standards and expert judgment, chatbot outputs should be treated as supportive rather than definitive.
There are inherent limitations to the current study. First, the chatbot responses used were generated in June 2025. Performance may change over time as each model is updated. Second, the answers were produced using a single prompt structure with 20 patient-centered questions on thyroid RFA. This does not capture variation in performance that may occur with different prompt styles, question phrasing, or topic complexity. Although questions were designed for common patient inquiries, there are no RFA-specific health literacy instruments.
Future studies could evaluate whether chatbots tailored to specific specialties, such as endocrinology or otolaryngology, have improved performance. Our 20-item question set was intentionally designed around common patient queries about indications, technique, risks, and follow-up for thyroid RFA. This mainly tested factual recall and general counseling rather than complex decision-making from LLMs. Subsequent work could investigate scenarios that require multi-step clinical reasoning, comorbidity personalization, and integration of clinical or imaging data that build on thyroid-focused and oncologic LLM evaluations.7,8,24 While the current study focused on a relatively new and rapidly evolving technique in thyroid RFA, evaluations across other procedures would provide a broader picture of chatbot reliability. Reliability is potentially more volatile for RFA than for mature procedures, as RFA device capabilities, techniques, and indications are changing quickly. In the future, we anticipate specialty-specific benchmarks to help compare chatbots more fairly. Finally, combining AI models with trusted clinical sources, or purely academic LLMs, will improve accuracy and reliability without losing natural language.
A head-to-head comparison of AI chatbots with board-certified specialty-specific surgeon reviewers represents one of the more robust evaluations of AI performance to date. We found meaningful differences across chatbots: Gemini delivered the most accurate and consistent answers, while ChatGPT offered the best readability but more variability. For physicians and surgeons, especially those introducing newer treatments like thyroid RFA, AI tools should supplement—but not replace—trusted education materials.
Authors’ Contributions
J.B.: Methodology (equal), data curation (lead), formal analysis (lead), investigation (lead), visualization (lead), and writing—original draft (lead). L.E.E.: Conceptualization (equal), methodology (lead), data curation (equal), formal analysis (equal), investigation (equal), and writing—review and editing (equal). D.C.: Conceptualization (equal), methodology (equal), validation (lead), data curation (equal), and writing—review and editing (equal). V.K.D.: Resources (equal) and writing—review and editing (equal). D.G.: Resources (equal) and writing—review and editing (equal). J.N.: Resources (equal) and writing—review and editing (equal). L.A.O.: Resources (equal) and writing—review and editing (equal). M.R.: Resources (equal) and writing—review and editing (equal). J.O.R.: Resources (equal) and writing—review and editing (equal). C.F.S.: Resources (equal) and writing—review and editing (equal). E.E.C.: Conceptualization (equal), methodology (lead), supervision (lead), project administration (lead), and writing—review and editing (equal).
Footnotes
Author Disclosure Statement
Dr. Julia Noel is a consultant for Pulse Biosciences. No other conflicts to disclose.
Funding Information
No author has funding to declare.
Data Availability Statement
All data analyzed during this study are included in this article or available from the corresponding author on request.
