Abstract
Objective
The purpose of this study is to perform multiple (
Approach
In this work, a novel deep neural network called SkV-Net is developed to reconstruct multiple material density images from the ultra-sparse spectral CBCT projections acquired using the ultra-slow kV switching technique. In particular, the SkV-Net has a backbone structure of U-Net, and a multi-head axial attention module is adopted to enlarge the perceptual field. It takes the CT images reconstructed from each kV as input, and output the basis material images automatically based on their energy-dependent attenuation characteristics. Numerical simulations and experimental studies are carried out to evaluate the performance of this new approach.
Main Results
It is demonstrated that the SkV-Net is able to generate four different material density images, i.e., fat, muscle, bone and iodine, from five spans of kV switched spectral projections. Physical experiments show that the decomposition errors of iodine and CaCl
Significance
SkV-Net provides a promising multi-material decomposition approach for spectral CBCT imaging systems implemented with the ultra-slow kV switching scheme.
Introduction
Cone-beam computed tomography (CBCT) is widely used in image-guided interventions and radiation therapy. Despite its significant advantages, such as providing high isotropic resolution volumetric images at relatively low radiation doses and cost, the imaging performance of current CBCT systems still lags behind that of diagnostic multi-detector CT (MDCT), particularly in terms of image quality and the ability to provide quantitative imaging information. 1 Spectral CT imaging offers attenuation information across different X-ray spectra, enabling material decomposition and improved tissue differentiation. In clinical practice, spectral CT has demonstrated promising capabilities in image contrast enhancement, 2 artifact mitigation, 3 and material recognition.4,5
There are three major potential approaches to achieve spectral CBCT imaging. The first is dual-source technique,6,7 where two X-ray tubes with different energy settings are used simultaneously to acquire data at two different energies. This method allows for spectral imaging and material decomposition but suffers from the disadvantage of increased system complexity and higher cost due to the need for two separate X-ray tubes, as well as potential cross scattering challenges. Another approach is the use of dual-layer detectors,8,9 which consist of two layers of detector materials that capture X-ray data at different energies. This technique provides simultaneous dual-energy information, but it can lead to increased detector thickness and lower spatial resolution due to the dual-layer design. A third method is fast kV switching,10,11 where the X-ray tube alternates between high and low kV settings during the scan, rapidly switching between different energy levels to capture spectral information. It has the disadvantage of longer scan times and also requires high-cost hardware.
Additionally, researchers are also working on developing new, cost-effective solutions for spectral CT technology. For instance, Petrongolo et al. 12 proposed obtaining high-energy and low-energy data by inserting a primary beam modulator between the X-ray source and the object. Tivnan et al.13,14 introduced the use of a spatial spectral filter to generate multiple spectral channels on the detector surface. Another approach is the slow kV modulation technique, in which the X-ray tube voltage is gradually switched as the gantry rotates around the object.1516–17 Compared to fast kV switching technology, the spectral information acquired through slow kV modulation is angularly sparse, which introduces additional challenges in material-specific CT image reconstruction. Therefore, specialized algorithms are required to achieve high-quality slow kV switching spectral CT imaging. For example, Szczykutowicz et al. proposed an iterative two-material decomposition algorithm based on the prior image constrained compressed sensing (PICCS) algorithm. 18 Additionally, one-step material decomposition methods can also generate material-specific CT images by incorporating the slow kV modulation information into the forward imaging model. Notable examples include the joint statistical one-step iterative material image reconstruction algorithm proposed by Mechlem et al., 18 and the non-convex primal-dual one-step reconstruction algorithm proposed by Chen et al. 19
Recently, deep learning techniques have been employed to enhance material decomposition in spectral CT imaging.2021–22 Convolutional neural networks (CNNs) have been successfully applied to reconstruct two-material basis images in slow kV switching CT imaging.23,24 While these approaches have been validated, it is important to note that CNNs have primarily been used for two-material decomposition. To the best of our knowledge, deep learning-based multi-material (
In this study, we introduce an innovative CNN-based network, named SkV-Net, developed to achieve high-quality four-material decomposition in ultra-slow kV switching spectral CBCT imaging, where the tube voltage changes slowly from low to high voltage during a single rotation. The SkV-Net takes CT images reconstructed from each kV segment as input and automatically generates four material-specific CT images with high precision. Numerical experimental results demonstrate that the proposed SkV-Net can accurately predict multi-material images with excellent accuracy and image quality.
The rest of this paper is organized as follows: Section II introduces the design of the SkV-Net network, Section III provides details of data preparation and network training, Sections IV and V present the experimental setup and results, respectively, and Section VI provides the discussions and a brief conclusion.
Method
Imaging model
The spectral CT based on fast kV switching and ultra-slow kV switching techniques are illustrated in Figure 1. In fast kV switching technique,25 the low and high tube voltages (80 kV and 140 kV) are switched on every alternate projection view. In contrast, each rotation of the CT scan based on ultra-slow kV switch contains several equally distributed spans of different kV settings. In this study, it is assumed that five kV spans are equally distributed between 85 kV and 125 kV, see the bottom illustration in Figure 1.

The fast kV switching scheme (top), and the ultra-slow kV switching technique (bottom). In the fast kV switching scheme, it is assumed that the kV switches between the minimum and maximum tube voltages. In the ultra-slow kV switching scheme, it is assumed that the kV switches for five times between the minimum and maximum tube voltages.
To perform material decomposition, the imaging model of ultra-slow kV switching CT need to be built. Assuming the attenuation coefficient
Let
The SkV-Net
Recent research has demonstrated that deep neural networks based on the U-Net architecture exhibit exceptional performance in medical imaging tasks.
25
By integrating advanced attention mechanisms into the U-Net framework, image quality can be further enhanced.26,27 Inspired by these advancements, we have developed an innovative deep neural network, SkV-Net, designed to perform material decomposition in ultra-slow kV switching CT. The architecture of SkV-Net is illustrated in Figure 2. The CBCT volume images are processed in a slice-by-slice manner by SkV-Net. It takes five CT image slices, reconstructed from the limited-view sinograms at each kV setting, as input and outputs four material-specific density images, including water/muscle, bone/CaCl

Architecture of the proposed SkV-Net. The backbone structure is based on the U-Net. At the highest-dimensional image level, an attention mechanism is applied for spectral signal feature extraction.

Structure of the axial attention model. The input consists of deep spectral signal feature maps extracted by the U-Net network, and the output is the feature maps encoded by the multi-head axial attention mechanism. Embedding the axial attention model aids in the fusion of deep spectral signal features, thereby more effectively guiding the decomposition of base materials.
Data and network training
Data preparation
Due to the unavailability of an ultra-slow kV switching CT system, it is challenging to collect spectral CT data and their corresponding basis material images. Therefore, we generate simulated data for network training 22 and evaluate the material decomposition performance of SkV-Net using both numerical and experimental data. The labels used during network training are the basis images generated through natural image processing, as illustrated in III A.1. The input data for training consists of CT images reconstructed from the ultra-slow kV switching sinogram, which is synthesized from the basis images using the spectral CT imaging physics, see III A.2.
Basis image
The basis image generation scheme is shown in Figure 4. First, a large amount of natural images are downloaded from the ImageNet database.
29
Then the pixel values of the R-channel of the RGB image are normalized to a range between 0 and 1, and the resulting image is denoted as

Key steps of generating numerical basis images from a natural image. Specifically,
Ultra-slow kV switching sinograms
The ultra-slow kV switching projections were generated according to equation 3. The normalized spectra
Network loss
During network training, the loss function for SkV-Net is the mean squared error (MSE) between the material density image predicted by the network, denoted as
Training strategy
The training of the SkV-Net was performed in the Pytorch platform environment on a high-performance workstation equipped with an Intel Xeon® Silver 4210R CPU and an Nvidia RTX A6000 GPU. The initial learning rate was set to
Physical experiments
The experimental ultra-slow kV switching data were acquired from a benchtop CT system, which was equipped with a medical-grade X-ray tube (G-242, VAREX, UT, USA) and a flat-panel detector (4343CB, VAREX, UT, USA),as shown in Figure 5 (c). The geometrical parameters are listed in Table 1. Five full CT scans were performed with different tube voltages (from 85kV to 125kV in 10kV increments) and beam filtration of 1.5 mm Al and 0.4 mm Cu. The ultra-slow kV switching data were synthesized from these five datasets by selecting the corresponding projection data at different views. Two samples were scanned in this study, a tube phantom and a pork specimen, as shown in Figure 5 (a) and (b). In the tube phantom, seven tubes (10 mm in diameter) of different substances were immersed in water (90 mm in diameter), including one tube of vegetable oil, three tubes of calcium chloride (CaCl

(a) and (b) are the tube phantom and pork specimen used in the physical experiments, (c) shows the experimental CT imaging benchtop.
Key parameters used for experimental data acquisition.
Results
Numerical simulation results
The decomposition results of the numerical XCAT phantom are shown in Figure 6. Images in different columns correspond to four material bases: bone, iodine, fat and muscle. We can observe that the proposed SkV-Net could generate basis images that are very close to the ground truth with relatively small residuals. Different materials are well separated while fine anatomical structures are preserved, see the zoomed-in region of interest (ROI) in the fat image. The quantification results of the iodine inserts are listed in Table 2. The decomposed iodine densities are highly consistent with the true values with less than 6

Decomposition results of the XCAT phantom. From top to bottom, they correspond to the ground truth, SkV-Net results, and residual images. The display windows for bone, iodine, fat, and muscle basis images are [0, 1.4] g/cm
The measured mean values and standard deviations(unit: mg/ml) of the iodine concentrations for numerical XCAT phantom.
The measured MAE and SSIM values for numerical XCAT phantom.
Figure 7 demonstrate the CT images of different tube voltages synthesized from the decomposed material-specific images. The ground truth was generated via FBP reconstruction from the full scan projection data. We can observe that the CT images synthesized by the SkV-Net is highly consistent with the ground truth in terms of structural similarity and quantitative accuracy. These results indicates that high-quality CT images of different tube voltages can be obtained by the SkV-Net even though each voltage is only used in a limited-view angular range.

Images from (a1) to (a5) are CT images reconstructed from the full scan projections of each kV. Images from (b2) to (b5) are CT images of different tube voltages synthesized from the decomposition results of SkV. All display Windows are [0, 0.42] cm−1. The scale-bar denotes 35 mm. Plots in (c) and (d) are the measured attenuation values in two ROIs highlighted by circles in (a1) and (b1), which correspond to the iodine and muscle, respectively.
Physical experiment result
Tube phantom
Figure 8 shows the decomposition results of the water phantom. Images in the first row correspond to the decomposed CaCl

Tube phantom results. Images from (a1) to (a5) are the decomposed CaCl

Quantification results of (a) Iodine and (b) CaCl
Pork specimen
The decomposition results of the pork specimen is demonstrated in Figure 10. Images in the first row correspond to the decomposed bone, iodine, fat, muscle basis images and the colour overlay image. Images in the second and third rows correspond to the the FBP images reconstructed from the full scan projections of different tube voltages, and the CT images synthesized from the decomposed basis images. Plots in Figure 10(d) and (e) compares the attenuation values of the FBP reconstructed and synthesized CT images. It is observed that basis materials in the pork specimen can also be identified. Structural information are preserved, see the zoomed-in ROI in Figure 10(a1). The synthesized CT images at different kVs have close appearance and consistent attenuation values with reduced image noises compared to the FBP images, see Figure 10(d) and (e).

Pork specimen results. Images from (a1) to (a5) are the decomposed bone, iodine, fat, muscle images and the colour overlay image. Display windows for four basis material images are [0, 1.4] g/cm
Computation time
In the current implementation, it takes approximately 0.14 seconds for the SkV-Net to decompose four
Discussions and conclusion
In this study, a dedicated multi-material decomposition network, SkV-Net, was developed for ultra-slow kV switching spectral CBCT. Compared to the fast kV switching technique, which alternates between low and high tube voltages in every other projection view, the ultra-slow kV switching technique requires only a few (five was assumed in this study) kV changes per gantry rotation. This makes it much easier to implement using existing CBCT systems. The proposed SkV-Net features a U-Net structure integrated with a multi-head axial attention mechanism. It takes the reconstructed CT images from different tube voltages as input and outputs multiple basis material density maps. Numerical simulations of the XCAT phantom, as well as physical experiments with a tube phantom and a pork specimen, were conducted. The results demonstrated that SkV-Net can accurately predict the density distributions of multiple basis materials.
While the preliminary results are promising, there are several limitations to the current work. First, due to the unavailability of ultra-slow kV switching hardware, only numerical simulations and experimentally synthesized data were used to evaluate the performance of SkV-Net. Further validation with real ultra-slow kV switching data is necessary once such CT systems become available. Second, the network design is still relatively empirical. More complex network architectures, such as transformer-based networks, 32 generative adversarial networks, 33 conditional denoising diffusion probabilistic model (C-DDPM), 34 may provide additional performance improvements. Third, a more accurate imaging model is needed to generate high-quality numerical training data, accounting for factors such as detector responses, Compton scattering effects, and the X-ray source spectrum. Lastly, the tube voltages used in this study were limited to the range of 85 kV to 125 kV with 10 kV increments. Future studies could explore a broader range of acquisition settings with smaller increments.
In summary, a novel convolutional neural network, SkV-Net, has been developed for multi-material decomposition in ultra-slow kV switching spectral CBCT imaging. Both numerical and experimental results demonstrate the potential of this network for high-quality spectral CBCT imaging. It offers promising possibilities for the development of low-cost spectral CBCT systems in the future, which could enhance the capabilities of medical imaging in clinical practice.
Footnotes
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported in part by the Guangdong Basic and Applied Basic Research Foundation (2021TQ06Y108), the National Natural Science Foundation of China (62422123, 62201560, U23A20284).
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
