Abstract
Objectives
Pressured healthcare resources make risk stratification and patient prioritisation fundamental issues for the investigation of colorectal cancer (CRC) in symptomatic patients. The present study uses machine learning algorithms and decision strategies to improve the appropriate use of colonoscopy.
Design
All symptomatic patients in a single health board (2018–2021) proceeding to colonoscopy to investigate for CRC were included. Machine learning algorithms (NeuralNetwork, randomForest, Logistic regression, Naïve-Bayes and Adaboost) were used to risk-stratify patients for CRC using demographics, symptoms, quantitative faecal immunochemical test (qFIT) and haematological tests. Decision curve analyses were performed to determine the optimal decision strategies.
Results
3776 patients were included (median age, 65; M:F,0.9:1.0) and CRC was identified in 217 patients (5.7%). qFIT > 400 μg Hb/g was the most important variable (%IncMSE = 78.5). RandomForrest had the highest area under curve (0.91) and accuracy (0.80) for CRC. When utilising decision curve analysis (DCA), 30%, 46% and 54% of colonoscopies were saved at accepted CRC probabilities of 1%, 2% and 3%, respectively. RandomForrest modelling had superior net clinical benefit compared to default colonoscopy strategies.
Conclusions
MLA-derived decision strategies that account for patient and referrer risk preference reduce colonoscopy demand and carry net clinical benefit compared to default colonoscopy strategies.
Introduction
Establishing efficient and sensitive referral pathways for the investigation of colorectal cancer (CRC) is essential when colonoscopy resources are limited but depends upon the prioritisation of high-risk patients.1,2 Utilising decision curve analysis in accordance with machine learning algorithms (MLA) could address such a challenge by recognising the relative importance of clinical factors and by acknowledging patient risk tolerance, neither of which are addressed by current guidelines.3–7
The National Institute for Health and Clinical Excellence (NICE) suggests urgent referral for suspected cancer on the basis of high-risk symptoms, iron deficiency anaemia or elevated quantitative faecal immunochemical test (qFIT).8–11 Relying on such guidelines recognises individual variables that should prompt referral for colonoscopy, but does not acknowledge overall risk as estimated by MLA. Clinical variables (e.g. symptoms) are important but their significance in conjunction with other variables (e.g. qFIT) must be put in perspective.3,12–20
A patient's and referrer's tolerance towards risk should be considered and weighed up against the risk of CRC, yet both current guidelines and published risk models do not support this aspect of decision making. More comorbid patients who are opposed to undergoing procedures may have a higher threshold for investigation than a less comorbid patient. Decision curve analysis (DCA) is a method for evaluating the benefits of a diagnostic test across a range of patient's possible risk tolerance towards over-investigation and under-investigation. In the case of CRC, DCA could be used to identify the best colonoscopy strategy to ensure that colonoscopy is used appropriately, guided by risk preference.
The present study aims to use MLA techniques to predict CRC in symptomatic patients and to use DCA to improve the appropriate use of colonoscopy.
Methods
All symptomatic patients within a single health board, of population of approximately 400,000 people, who presented to primary health services with possible CRC between June 2018 and October 2021 and proceeded to colonoscopy were considered for inclusion. Patients who were investigated as part of the national bowel screen program were not considered.
In our health board all referrals by primary care with the suspicion of CRC had to undergo a qFIT prior to referral for colonoscopy. Patients were identified through a regional database of all qFIT results and followed up for colonoscopy, if applicable. Patient data was collected using data-linkage methodology and those who did not have a full set of data recorded (demographics, symptoms, a qFIT result, haematology tests and colonoscopy) during diagnostic work-up were excluded (n = 130, 3.3%).
Symptom data were retrieved from the electronic referral at the time of referral for colonoscopy and from the symptoms at the time of qFIT request. In the case of multiple symptoms, all were recorded. Patients’ haematology results were analysed and iron deficiency anaemia (IDA) was defined as a low Hb (<130 g/L for males; <120 g/L for females) and either a low ferritin (<30 mcg/L) or a high transferrin (>3 g/L). 16 A qFIT of >7 μg Hb/g was considered elevated.
Colonoscopy outcomes were recorded; as well as CRC, additional significant bowel disease (SBD) were recorded and included colonic polyps and inflammatory bowel disease. Upper gastrointestinal malignancy was also recorded for those who proceeded to have an upper GI endoscopy in addition to colonoscopy.
Statistical analysis
Characteristics of those with CRC were compared with those without CRC using Mann-Whitney U and Chi-squared tests.
MLA were used to generate models predicting CRC using pre-colonoscopy data (demographics, symptoms, qFIT result and haematology tests). A five-fold cross-validation process using ‘R Studio’ was performed using the following models: artificial neural network (ANN; package ‘neuralnet’), randomForrest (RF; package ‘randomForrest’), multivariate logistic regression (MLR; package ‘GLM’), Naïve-Bayes (NB; package ‘naivebayes’) and AdaBoost (ADA; package ‘Ada’). Each model aimed to predict the likelihood of CRC as a binary classifier (0/1). The number of ‘nodes per layer’ and the number of ‘hidden layers’ for ANN were each set as four guided by the number of variables. In RF, 501 decision trees were run and the number of variables sampled at each tree (mtry) was set to four. In ADA, the number of iterations was set to the default of 50 iterations and the ‘nu’ shrinkage parameter was set to the default of 1.
With the aim of predicting CRC, each model included the following variables: demographics [age (<40, 40–59, 60–79, ≥80), gender (M/F)], qFIT value (0, 7–19; 20–49; 50–99; 100–199; 200–299; 300–399; ≥400 μg Hb/g), IDA, symptoms (abdominal pain, altered bowel habit, rectal bleeding, weight loss, other). The cohort was split at random and models were derived using 60% of the cohort and then tested with the remaining 40%. Receiver operating characteristic (ROC) curves were generated for each model (packages “R0RC” and “pROC”) with CRC as the outcome measure and the area under curve (AUC) was reported. Confusion matrices were then generated using the package ‘caret’ for each model and the diagnostic accuracy measurements were reported (e.g. sensitivity, specificity, positive predictive value (PPV) and negative predictive values). The sum of the sensitivity and specificity for each model iteration were optimised and the final model accuracies were reported. The above methodology was repeated for each model using only the ‘qFIT’ variable to determine the relative predictive ability of qFIT in comparison to all variables.
With the aim of predicting SBD, the analyses were repeated and SBD was considered the outcome measure. SBD included CRC, colonic polyps and inflammatory bowel disease.
Relative importance of variables
The importance of each variable (e.g. qFIT > 400 μg Hb/g) for the prediction of CRC was calculated using the importance function in the ‘randomForest’ package and reported as the mean decrease accuracy (MDA) and mean decrease Gini (MDG) values. MDA reports to what degree the accuracy the model suffers from eliminating individual variables and MDG is a measure of the contribution from each variable to the homogeneity of the nodes in the model.
Decision curve analysis
DCA were generated for each MLA using the package ‘dcurves’ to help determine the optimal MLA and to determine the number of colonoscopies that could be saved compared to current practice. The DCA determines the clinical benefit of utilising the MLAs in comparison to default strategies such as ‘colonoscopy for all’ or ‘colonoscopy for none’. DCA demonstrates clinical value by reporting ‘net benefit’ of tests or models at certain ‘threshold probabilities’. The ‘net benefit’ is a weighted calculation of true and false positives that accounts for the risk of cancer and the risk of colonoscopy at a certain ‘threshold probability’. The ‘threshold probability’ is the accepted degree of risk estimated by the model that a patient or referrer would accept before agreeing to a colonoscopy. For example, a 1% risk threshold would assume a risk of 99 unnecessary colonoscopies as equivalent to the risk of missing one CRC. Depending on the accepted risk preference or ‘threshold probability’ a net benefit in true positives can be gained using models relative to the ‘colonoscopy for none’ approach. Alternatively, a percentage of colonoscopies can be saved relative to the ‘colonoscopy for all’ approach.
The optimal MLA via decision analysis was determined and patients were divided into thirds (high risk, intermediate risk and low risk for CRC) and different decision analysis strategies were compared based on risk profile.
All statistical tests were performed using R.Studio 2022.22.1 and p-value <0.05 was regarded as statistical significance. All data retrieval was carried out in accordance to local guidelines. No experimental protocols were used. As a local audit informed consent from the subjects and ethical approval were not required in accordance with local guidance.
Results
Background
In total, 3776 patients were included (median age, 65; M:F, 0.9:1.0) (Table 1) and CRC was identified in 217 patients (5.7%). Patients with CRC were older (73 vs. 65; p < 0.001), were more likely to be male (p < 0.001) and had a higher qFIT (p < 0.001). The CRC group had a lower Hb, ferritin, mean corpuscular volume, mean corpuscular haemoglobin, transferrin, transferrin saturation and higher platelets (p < 0.05) (Table 1). The diagnosis of alternative SBD following colonoscopy included non-cancerous colonic polyps in 736 patients (19.5%) and inflammatory bowel disease in 36 patients (1.0%). There were also 8 patients identified with upper gastrointestinal cancer (median age, 69; M:F, 1:1.7) identified on gastroscopy, 5 of which had a positive qFIT (range, 10–>400 μg Hb/g) and 3 of which had IDA.
Background data of primary cohort.
*Where the reason for qFIT was ‘anaemia’, the patient may not have had iron-deficiency anaemia by the accepted definition.
CRC: colorectal cancer; Hb: haemoglobin; MCH: mean corpuscular haemoglobin; MCV: mean corpuscular volume; TSAT: transferrin saturation; qFIT: quantitative faecal immunochemical test.
Of those with CRC, 93.1% (201/217) were qFIT positive compared with 55.6% (1979/3560) of those without CRC. The rate of CRC increased from 1% (15/1595) in the qFIT-negative group to 19.5% (136/697) in the qFIT ≥400 group. There was no significant difference in symptoms between the CRC and non-CRC groups (p > 0.05).
Prediction of colonoscopy outcome using MLA models
Using the five-fold cross-validation method (ANN, RF, MLR, NB and ADA), ROC curves were generated (Figure 1) for the prediction of CRC and the AUC were reported. Diagnostic accuracy measurements were reported (Table 2). The RF model had the highest AUC (0.91) and the highest accuracy (0.80) when aiming to maximise the sum of sensitivity and specificity.

Receiver operating characteristic curves for each model iteration using machine learning algorithms. AUC: randomForrest (RF), 0.91; neural network (ANN), 0.84; logistic regression (MLR) 0.85; Naïve-Bayes (NB), 0.75; AdaBoost (ADA), 0.75. ADA: AdaBoost; ANN: artificial neural network; AUC: area under curve; MLR: multivariate logistic regression; NB: Naïve-Bayes; RF: randomForrest.
Diagnostic accuracy for each model iteration to predict CRC using (A) all variables and (B) qFIT only.
ADA: AdaBoost; ANN: artificial neural network; CI: confidence interval; CRC: colorectal cancer; NPV: negative predictive value; PPV: positive predictive value; qFIT: quantitative faecal immunochemical test; MLR: multivariate logistic regression; NB: Naïve-Bayes; RF: randomForrest.
The analyses were repeated using only qFIT and for each MLA, the AUC of the all-variable models were superior to that of the qFIT-only models (RF, 0.91 vs. 0.81; ANN, 0.84 vs. 0.80; MLR, 0.85 vs. 0.80; NB, 0.75 vs. 0.72; ADA, 0.75 vs. 0.65, respectively). Similarly, for each model, the diagnostic accuracy measurements were superior using all variables compared to only qFIT (Table 2).
The five MLA models were used to predict SBD: ROC curves for each model were plotted and AUC for each model were determined: RF, AUC = 0.66; ANN, AUC = 0.67; MLR, AUC = 0.63; NB, AUC = 0.58; ADA, AUC = 0.58 (Figure 2). Diagnostic accuracies for each model were reported (Table 3).

Receiver operating characteristic curves for each model to predict SBD: RF, AUC = 0.66; ANN, AUC = 0.67; MLR, AUC = 0.63; NB, AUC = 0.58; ADA, AUC = 0.58. ADA: AdaBoost; ANN: artificial neural network; AUC: area under curve; MLR: multivariate logistic regression; NB: Naïve-Bayes; RF: randomForrest; SBD: significant bowel disease.
Diagnostic accuracy for each model to predict SBD using all variables.
ADA: AdaBoost; ANN: artificial neural network; CI: confidence interval; NPV: negative predictive value; PPV: positive predictive value; qFIT: quantitative faecal immunochemical test; MLR: multivariate logistic regression; NB: Naïve-Bayes; RF: randomForrest; SBD: significant bowel disease.
Variable importance
The importance function in the most accurate model (RF), was used to determine the most important variables for the prediction of CRC (Figure 3 and Table 4). qFIT over 400 μg Hb/g was the most important variable (%IncMSE = 78.5), followed by IDA (%IncMSE = 14.9), age > 80 (%IncMSE = 14.2) and age 60–79 (%IncMSE = 13.6). Of all symptoms, altered bowel habit was the most important variable. Variables which did not improve the predictive ability of the model included weight loss, qFIT 20–49 μg Hb/g, qFIT 50–99 μg Hb/g, qFIT 100–199 μg Hb/g and qFIT 300–399 μg Hb/g.

Importance of individual variables using ‘randomForest’ MLA reported as mean decrease accuracy (%IncMSE) and mean decrease Gini (IncNodePurity).
Importance of individual variables using ‘randomForest’ MLA reported as mean decrease accuracy and mean decrease Gini.
qFIT: quantitative faecal immunochemical test; MLR: multivariate logistic regression.
Decision-curve analysis
The net benefit of using each model is demonstrated across the threshold probabilities (Figure 4). RF was the superior model with greatest net benefit across the threshold probability range.

Decision curve of ANN, RF, MLR, NB, ADA; net benefit vs. threshold probability. ADA: AdaBoost; ANN: artificial neural network; MLR: multivariate logistic regression; NB: Naïve-Bayes; RF: randomForrest.
The cohort was split into thirds (‘high risk’, ‘intermediate risk’ and ‘low risk’) groups guided by the RF scores (PPV for CRC of 15.9%, 1.1% and 0.08%, respectively). Colonoscopy strategies (e.g. colonoscopy for high risk only; colonoscopy for intermediate and high risk; colonoscopy for RF score > threshold probability) were derived and compared to ‘colonoscopy for all’ and ‘colonoscopy for none’ in terms of net benefit and the number of saved colonoscopies (Figures 5 and 6; Table 5).

Subgroup decision-curve of RF; net benefit at threshold probabilities compared to ‘Colonoscopy for None’. RF: randomForrest.

Number of colonoscopies that can be saved per 100 patients by implementing specific colonoscopy strategies compared to ‘Colonoscopy for All’.
Subgroup analysis: Net benefit and saved colonoscopies by threshold probability.
RF: randomForrest.
Even at low threshold probabilities (1–5%) the ‘high risk’, ‘intermediate/high risk’ and ‘RF score’ strategies saved the highest number of colonoscopies and had the greatest net benefit. It is worth noting that the net benefit of the ‘colonoscopy for all’ approach drops considerably above 1% and becomes negative just beyond a threshold probability of 5%. It is also worth acknowledging that ‘colonoscopy for all’ is only superior at extremely small threshold probabilities of approximately 0.1%.
Discussion
To the best of our knowledge this is the first paper to utilise MLA in combination with decision analysis to improve risk stratification for CRC. Decision analysis is performed to account for a patient's and referrer's risk preference and this is presented as a potential unique strategy to risk stratify patients for colonoscopy. MLA in conjunction with decision analysis may help offer colonoscopy more appropriately and save colonoscopy resources. Performing colonoscopy on intermediate and high-risk patients for CRC, determined by RF, would save a significant number of colonoscopies (30–33%) even at low threshold probabilities (1–5%).
The application of DCA initially focused on the evaluation of different prostate biopsy strategies and its advantages have since become more widely recognised by high-profile journals such as JAMA, BMJ and Journal of Clinical Oncology.21–25 The unique benefit of utilising DCA is the ability to evaluate if a particular intervention is beneficial to a patient and to decide between different management strategies. In the present study, the application of DCA proves to be a useful strategy to rationalize colonoscopy for CRC to achieve a net clinical benefit compared to default strategies at reasonable threshold probabilities.21–23,26–29 A threshold probability of 1% for a patient would assume that the benefit of treating a CRC in that patient, considering improvement in prognosis, is 99 times greater than the harm of a colonoscopy in that patient. At such low threshold probabilities (1%) and even up to 20%, the proposed strategies such as ‘colonoscopy for high risk’ and ‘colonoscopy for intermediate and high risk’ would result in greater net benefit than default strategies.
Colonoscopy cannot be performed without accepting its complications (e.g. bleeding/perforation), significant discomfort and considerable cost to health boards.30,31 The risk of complications relating to colonoscopy has been reported to be between 0.32% and 0.5% in colonoscopy but 0.7% in those requiring biopsy.32,33 As such, in patients with a threshold probability of 1%, the risk of serious complication would range from 32% to 70% for every CRC found. In cases of threshold probabilities of 10%, patients would only accept a 3.2% to 7% risk of serious complication for every CRC found. More comorbid patients may not survive serious complications and as such they will have higher threshold probabilities. Through such a perspective, more appropriate decisions regarding referral for colonoscopy can be made. Although knowing the exact threshold probability is not necessary; if a patient approximate threshold probability is estimated, the optimal model (e.g. colonoscopy for intermediate/high risk) can be implemented.
In decision curves one must consider the repercussions of missing a CRC in that specific patient. For example, in a comorbid patient, the benefit of surgery and impact of missing CRC may be low and this marginal benefit may be offset by the risk of colonoscopy. This would result in a higher threshold probability and proceeding without colonoscopy should be strongly considered. Other acceptable but less invasive diagnostic options (e.g. computed tomography (CT) colonography) should be considered in these patients.34,35
Clinicians have a responsibility to not cause harm to patients, but also to select treatment pathways that optimise net benefit overall by maximising true positives and minimising false positives. As demonstrated this must be achieved by estimating the threshold probability for a patient individually and not through a ‘one size fits all’ approach. It is not surprising that a ‘colonoscopy for all’ approach causes relative harm compared to the present proposed strategies across a range of reasonable threshold probabilities and sheds lights on the clinical detriment from inappropriate risk stratification.
Whilst unidimensional guidelines aim to direct referrers towards the appropriate decisions, they do not necessarily represent overall risk. Incorporation of MLA, through a multidimensional view of demographic and clinical variables, may be a possible solution. For example, simply acknowledging the positive association between anaemia and CRC would fail to acknowledge that a negative-qFIT result is extremely reassuring in the presence of anaemia.10–13 With a multitude of relevant variables (gender, age, symptoms, anaemia and qFIT) MLA models may hold potential in improving our decision making beyond that of our own clinical judgement. 36 The present study demonstrates the strong predictive ability of qFIT in comparison to other variables. Although the models using qFIT alone offer impressive diagnostic accuracy as a single variable in all models, incorporating all variables was superior in each MLA. Further MLA models that incorporate an even more extensive set of relevant variables may help create superior models.5–7
The present study is only an initial step towards implementing full-proof and resource-efficient algorithms to improve risk stratification. Additional investigations (e.g. double qFIT-testing) in those deemed low risk may be valid methods to identify missed cancers without subjecting all patients to colonoscopy. Such additional investigations could be incorporated into future decision analyses to strengthen net clinical benefit and enroll more robust strategies. 37
There remain a number of concerns that cannot be dealt with by machine learning algorithms alone. The algorithms are only as good as the available data and are limited by the availability of large databases that contain sufficient numbers of CRC. Acquiring such datasets is the route to establishing sophisticated and full-proof patient-specific decision-making algorithms. The more detailed and accurate the dataset, the more effectively the MLA models can predict a patient's specific risk and propose appropriate management. Collaboration between centers will be required for such an endeavor and should be encouraged.
Whilst we report other less sinister yet significant bowel pathology (colonic polyps and inflammatory bowel disease) and attempt to predict these findings, the MLA models were significantly less accurate for all significant bowel pathology (accuracy 0.60–0.64; AUC, 0.58–0.67) versus CRC alone (accuracy, 0.72–0.80; AUC, 0.75–0.91). This can most likely be explained by the strong association between qFIT and specifically CRC that contributes a significant diagnostic component in each model. For this reason, in cases where the decision to perform colonoscopy is to detect alternative pathology, the DCA in the present study cannot be applied effectively. The present MLA models should be used where the detection of CRC is the primary question, and if the detection of alternative significant bowel disease is unlikely to influence management or further treatment. Further research is required to determine the optimal decision strategies for colonoscopy in the detection of CRC in symptomatic patients that also optimises the detection of other pathology. Lastly, another possible limitation to the study is the limited data available on symptoms. Although we have the type and number of symptoms, data on the frequency, severity and duration was not available for all patients. Nevertheless, symptoms are poorly associated with CRC in multiple studies and adding this information is unlikely to improve the diagnostic accuracy of the MLA.10,38,39
In conclusion, the present study proposes a novel strategy to risk stratify patients in order to maximise the net clinical benefit to patients, reduce the demand on colonoscopy resources and reduce overall harm to patients. The proposed strategies, although appear superior to current default strategies, are by no means a final model and should be considered a stepping-stone to more superior decision tools. 40 Future studies should aim to improve net clinical benefit through DCA by relying upon larger more comprehensive datasets and by incorporating alternative pathways (e.g. double qFIT strategies).
Footnotes
Acknowledgements
We acknowledge all authors for their hard work with this article.
Author statement
JL contributed to writing, method, results and discussion; EB to method and discussion; HH to data acquisition; PD to method and discussion; NC to method, results, discussion and supervision.
Data availability statement
The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.
Data sharing statement
Data can be shared upon reasonable request.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
