Abstract

To the Editor,
Stevenson et al. recently asked whether artificial intelligence can replace biochemists by comparing ChatGPT and Google Bard with practising biochemists in the interpretation of thyroid function test patterns. 1 Their finding that chatbots correctly interpreted only a minority of scenarios, and at times offered unsafe advice, is an important empirical warning as large language models (LLMs) increasingly intersect with routine laboratory workflows. 2
Yet the ‘replacement’ framing risks overshadowing a more critical challenge for laboratory medicine: building biochemist-led, evidence-based decision-support systems in which AI is supervised, context-aware and embedded within laboratory processes rather than deployed as decontextualised, free-text interfaces. Recent analyses emphasise that the most mature and safe AI applications in laboratory medicine are narrow, task-specific tools integrated into laboratory information systems – supporting quality control, autoverification, demand management and operational risk detection – rather than autonomous clinical decision-making.3–5 Viewed this way, the study by Stevenson et al. is best interpreted as a stress-test of inappropriate AI deployment rather than an indictment of AI’s potential value in clinical biochemistry. 1
Their work highlights several important gaps that merit attention from the journal. 1 First, the chatbots evaluated lacked access to structured laboratory data – including local reference intervals, historical deltas, pre-analytical flags or guideline-anchored interpretive frameworks – all of which can significantly influence model performance and safety.3–5 Second, safety was assessed in terms of point-in-time correctness, whereas current policy and regulatory frameworks emphasise longitudinal monitoring, transparency, explainability, and safeguards against inequitable performance across diverse populations. 6 Third, in many low- and middle-income countries, severe workforce shortages risk creating an environment where unsupervised chatbot use bypasses clinical biochemists entirely, potentially amplifying global inequities in test interpretation.2–5
Building on Stevenson et al., we outline three future directions for a solution-oriented research agenda. 1 (i) Inside-laboratory AI co-pilots: LLM-enabled tools can assist with flag interpretation, drafting context-sensitive comments, and enabling reflex or reflective testing – always under biochemist supervision.3–5 (ii) Industry 5.0-aligned human–AI collaboration: AI should offload repetitive analytical tasks, enabling biochemists to prioritise complex interpretive and consultative responsibilities.5,7 (iii) Global-scale evaluation and governance: multinational, prospective studies evaluating accuracy, workflow impact, patient comprehension and unintended consequences are essential for designing safe and equitable AI-supported pathways. 6
Rather than asking whether AI can replace biochemists, a more productive question – one well aligned with the mission of the Annals of Clinical Biochemistry – is: how can biochemists lead, design and govern AI systems that demonstrably improve patient care while preserving professional accountability? Stevenson et al. provide a valuable foundation; the next step is a global, evidence-driven agenda that frames biochemists not as subjects of automation but as architects of trustworthy laboratory AI.1–7
Footnotes
Acknowledgements
None.
Ethical approval
Not applicable. This manuscript is a
Guarantor
MVS (M. Vijaya Simha) is the guarantor for this article and takes full responsibility for the integrity and accuracy of the work, including the reference list.
Contributorship
MVS conceptualised the work, reviewed relevant literature, drafted the manuscript, performed all revisions, and approved the final submitted version. No other individuals contributed to authorship.
Use of Generative AI
A language model was used only for
