Abstract
Accurate material decomposition constitutes the foundation of Spectral Computed Tomography (Spectral CT) applications across diverse domains. Nevertheless, conventional model-based material decomposition methods face significant limitations including sparse-view sampling artifacts, slow convergence rates, noise amplification, and inherent ill-posedness—challenges that are particularly pronounced in geometrically inconsistent imaging. To overcome these constraints, we propose an unsupervised deep learning framework that synergistically optimizes virtual monochromatic images (VMIs) through the probabilistic diffusion model for direct material decomposition in sparse-view spectral CT. The proposed methodology introduces VMIs as critical differentiation enhancers for polychromatic projections, effectively addressing convergence limitations in iterative reconstruction algorithms. By incorporating probabilistic diffusion priors into the optimization process, we achieve superior refinement of material-specific representations. Our framework systematically enforces dual constraint: 1) data fidelity term ensuring measurement consistency, and 2) probabilistic regularization suppressing unwanted structures, thereby guaranteeing anatomically plausible material image reconstruction. Comprehensive validation on preclinical data demonstrates that our method achieves a 10 dB improvement in the peak-signal-to-noise ratio (PSNR) and a 4.31% increase in structural similarity (SSIM) for soft-tissue reconstructions compared to the optimal comparison algorithm with 90 projections. Experimental results confirm the algorithm's robustness under challenging conditions, maintaining reconstruction fidelity even with geometric inconsistency and sparse sampling.
Introduction
Spectral Computed Tomography (Spectral CT) leverages multi-energy X-ray spectra to enable material identification, a critical capability for applications such as bone density measurement, 1 liver fibrosis assessment, 2 abdominal imaging, 3 and breast cancer risk assessment. 4 Material decomposition algorithms in spectral CT are broadly classified into direct and indirect methods. Direct material decomposition methods (DMDM), 5 contrast with indirect methods, which encompass image-domain and projection-domain techniques. Image-domain methods involve reconstructing multi-channel spectral CT images followed by material decomposition. 6 These methods rely on simplified linear imaging model that inadequately align with actual polychromatic scanning condition. Projection-domain methods decompose projection data into material-specific projections prior to image reconstruction, 7 leveraging spectral and polychromatic models for nonlinear decomposition. 8 While continuously refined,9,10 these methods require precise spectral calibration and strict geometrical consistency across energy channels, posing practical implementation challenges.
DMDM directly reconstruct material images from projection data, accurately modeling spectral CT physics to mitigate beam-hardening artifacts and minimize cascaded information loss. In 2015, Zhao et al. proposed the extended algebraic reconstruction technique (EART) for dual energy CT, achieving promising results at the cost of low computational efficiency. 11 Subsequent improved algorithms are developed, such as the extended simultaneous algebraic reconstruction technique (ESART) 12 and virtual monochromatic images (VMIs)-guided iterative methods. 13 To improve reconstruction quality, DMDM incorporate regularization terms exploiting intrinsic image properties, such as the total variation (TV) 14 and block matching strategies. 15 Chang et al. 16 unified basis material reconstruction and spectral estimation into a TV-regularized non-convex optimization. However, traditional regularization designs remain empirically driven, limited in characterizing noise correlations and insufficient in describing the mapping relationships between spectral CT images and basis materials.
Deep convolutional neural networks (CNNs) have revolutionized spectral CT material decomposition. Zhang et al. 17 pioneered an end-to-end butterfly network architecture, later enhanced through loss functions optimization 18 and the generative adversarial network (GAN) structure for multi-material decomposition (MMD) in. 19 Moreover, the decomposition of three materials under dual-energy spectra have been realized by combining constraint condition, 20 while projection-domain decomposition and DMDM now incorporate sparse-view reconstruction and end-to-end mappings. 21 Li et al. 22 proposed interpretable decomposition iterative network for low-dose CT denoising. Despite overcoming material recognition via hierarchical features, supervised CNNs depend critically on high-quality labels—a significant limitation in medical image. Moreover, insufficient integration of physical imaging mechanisms often compromises data consistency and reconstruction fidelity.
Diffusion models (DMs) have gained prominence for their ability to learn image probability distributions and generate high-fidelity samples without label.23,24 Comprising a forward diffusion process and a backward sampling process, 25 DM excel in inverse problem solving under noisy conditions, exemplified by diffusion posterior sampling (DPS) 26 and manifold constraint gradient (MCG). 27 Liu et al. 28 applied DM to sparse spectral CT material decomposition, achieving accurate contrast agent distribution recovery. To address the slow sampling of DM, various approaches have been proposed, including jump sampling strategies, 29 denoising diffusion implicit models (DDIM), 30 and medical segmentation optimizations 31 — improve efficiency. DM-based spectral CT decomposition remains challenged by computational complexity, slow convergence and poor adaptability to material-specific nonlinearities, particularly under sparse sampling conditions.
This work proposes an unsupervised spectral CT direct material decomposition framework synergizing data-driven DM with physical imaging models. To address the slow convergence problem caused by geometric inconsistencies under fast kVp switching imaging, VMIs are employed to augment angular separation between polychromatic projection curves. Critically, DM iteratively refines VMIs via probability distribution learning and superior feature representation. To enhance the iteration efficiency, we introduce a truncated reverse diffusion method, allowing the DM to participate only in the latter half of the VMIs-guided iterative process. By bypassing early high-noise phase detrimental to CT image generation, this approach effectively mitigates the severe ill-posedness of spectral CT material decomposition, primarily caused by high correlation across different spectral channels, material attenuation curve similarities, and sparse-view protocols.
The structure of this paper is as follows: the second section introduces the background, primarily focusing on physics model for spectral CT and DM. The third section introduces the proposed method and the solution of the model. The fourth section is the experiment and analysis of the experimental results. The paper concludes with a brief discussion and conclusion.
Background
Physics model for spectral CT
This work takes into account the high and low energy projections to direct reconstruct two basis material maps. The dual-energy CT projection model can be expressed as follows:
For the specific expression of the coefficients, please refer to appendix I. Eq. (2) can be briefly written as the following linear system,
For the specific expression of
Diffusion probability model
Generative models are a class of models capable of learning data distributions and generating new samples, replicating, and even creating new samples similar to the training data.
32
Diffusion models,33,34 as one of the most popular probabilistic generative models, learn complex data distribution by estimating the gradient of the prior probability density
The score
Method
Direct material decomposition based on probability diffusion priors
The solution of the following inverse problem can be examined from a probabilistic point of view,
The problem of reconstruction material images based on measurement projections
According to Bayes’ rule, the MAP estimate can be rewritten as follows,
Our methodology integrates pre-trained probabilistic diffusion models as implicit probability regularizer within the Plug-and-Play (PnP) architecture 36 framework, combining the theoretical rigor of model-based approaches with the representational power of deep learning techniques. This unified approach enables simultaneous optimization of generative models and direct material decomposition, establishing a novel paradigm that synergistically integrates data-driven and model-driven methodologies. The diffusion-based regularizer overcomes the representational limitations inherent in conventional regularization approaches, creating a new framework for addressing inverse problems in spectral CT material decomposition.
Direct material decomposition framework integrates VMIs guidance with diffusion probability priors
The polychromatic nature of broad-spectrum X-ray sources leads to spectral overlap in acquired projection. This inherent mixing of multi-energy photons introduces high cross-correlation among attenuation measurements at different energy bins, resulting in the system severely ill-posed system. Additionally, the LAC of different materials can be very similar. For example, in breast imaging, the LAC differences between adipose (fat) and glandular/fibrous tissues across different energy are small. Sparse sampling further exacerbates the insufficiency of the acquired projection data. Collectively, these factors increase the ill-posedness of the reconstruction problem. This causes the spectral attenuation cureves to lie extremely close to one another, leading to slow convergence of the ESART algorithm. 12
To accelerate convergence, Zhang et al.
13
proposed a VMIs-guided method. This method increases the angle between polychromatic projection curves by introducing monochromatic images, and thus accelerating the iterative process. Specifically, two distinct energies
These VMIs can be represented as linear combinations of the basis materials,
Eq. (13) can be equivalently represented as the following,
The introduction of VMIs for material decomposition increases the angle between polychromatic attenuation curves, thereby accelerating the iterative convergence. More importantly, it facilitates the implementation of deep prior regularization using probability diffusion models. VMIs exhibit highly consistent distribution features with CT images, sharing similarities in anatomical structure edges and tissue contrast, and they also have the same grayscale scaling. In previous work, our research group successfully trained a probabilistic diffusion model using the publicly available LIDC-IDRI lung CT dataset. 37 Leveraging this, we can directly employ this pre-trained model without requiring retraining on the material distribution image dataset, providing substantial convenience for the smooth progress of our research. The flowchart of the proposed DMDM framework with virtual monochromatic guidance, incorporating a fused probability diffusion prior, is illustrated in Figure 1.

Visual representation of the proposed method. Under inconsistent geometry, the material residual projections of the VMIs are first updated, followed by VMIs reconstruction using the ESART as defined in Eq.(17). Subsequently, denoising is performed on the updated VMIs via the probabilistic diffusion model. Finally, material decomposition is executed based on the optimized VMIs following Eq.(14) to obtain material images.
Solution of the direct material decomposition framework
The objection function based on VMIs guidance is rewritten as follows,
The specific expression of
The iteration of the proximal gradient descent algorithm consist of two steps:
Step 1, Gradient descent step: Perform gradient updates on the fidelity term
The solution process in Step 1 can utilize the approach of ESART,
13
first updating the material residual projections, and then updating the VMIs. With the inconsistent geometry, the X-ray paths of the dual energy passing through the object to be scanned are no longer exactly the same, and the weights need to be recalculated to update the error. By the equation (25) in,
13
the material projection residuals are obtained. To enable the iteration to converge rapidly, the entire projection dataset is divided into T parts, and the index set of the rays within each projection is
Assume
Reference 38 presented this proposition, and its explanation is similar to that in, 39 but neither provided an explicit proof. This paper offers a complete proof process for the proposition, and the details can be found in appendix II.
Step 2, Proximal operator step: Apply the proximal operator to the non-smooth term
The definition of the proximal operator is provided as follows, for a function f on
According to Proposition 1, the solution in Step 2 can be approximated by the sampling process of the DM, as represented below,
After the above iterative process, the optimization
Summarizing the optimization process of the algorithm is listed in Algorithm 1.
Convergence analysis of the direct material decomposition framework
In this section, we present conditions for the convergence of DMDM to ensure the algorithm is convergent. First, under appropriate assumption, the monotonic decrease of the objective function is established.
Assume that: (1) the gradient of
The detailed proof is provided in Appendix II. This Proposition 2 ensures the convergence of the sequence of objective function values, thereby paving the way for the convergence of algorithm 1.
Now we are ready to establish the convergence analysis of Algorithm 1. By combining the result of Proposition 2 and some assumptions, we can derive the theoretical results as follows.
Assume that
Existence of solution to subproblem and boundedness of the generated sequence
Then the Algorithm.1 based on proximal gradient descent converges.
The detailed proof can be seen in Appendix II. Theorem 1 indicates that Algorithm 1 designed in this paper is convergent, and the solution deviation is controllable under certain conditions.
Assume that the generated sequence
Replace the score function
The detailed proof can be seen in Appendix II. Theorem 2 establishes that when neural networks are employed to approximate the score function with approximation errors constrained with
Experiment design
In this section, we introduce the comparison algorithms and the datasets, as well as the key details of training and evaluating the model. We describe the comparison algorithms employed, including the two-step decomposition approach (FBP-decomp), the method based on TV (ESART-TV), IRM-MI, 13 OPMT, 40 and GAN-MD. The FBP-decomp algorithm involves first reconstructing sparse images using FBP and then directly inverting them in the image domain. The GAN-MD algorithm, adapted from reference, 19 replaces the multi-material decomposition network with a dual-material decomposition network and shifts the focus from low-dose denoising to sparse artifact suppression. The GAN-MD network was trained on 1095 data samples constructed from dataset of the local hospital, partitioned into training, validation, and test sets at an 8:1:1 ratio. The trained model was then directly applied to real mouse scan data for material decomposition. The modified GAN-MD model takes dual-energy CT images as input and outputs two basis material images (bone and soft tissue). However, two key limitations affect model performance: (1) insufficient training data impedes model convergence, and (2) the lack of ground truth material labels restricts the construction of high-quality training dataset. These challenges are further compounded by the sparse-view scanning protocol employed in this study, which significantly increase the problem complexity. Consequently, the GAN-MD model exhibits constrained performance, particularly impairing its artifact suppression capability in the output material images. Furthermore, since the GAN-MD model was trained solely on simulated data, its generalization capability to real data may be limited, potentially affecting artifact suppression efficacy in practical applications.
We have optimized the implementation of each comparison algorithm to ensure the fairness of the experiments. To quantitatively assess the performance of the proposed algorithm, we use evaluation metrics such as Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM),
41
and Root Mean Square Error (RMSE). The reconstructed image exhibits lower distortion when demonstrating either reduced RMSE or increased PSNR values. The RMSE and PSNR are calculated as follows:
This model is constructed on the public PyTorch platform with the NVIDIA RTX A6000 GPU and Inter(R) Xeon(R) Silver 4216 CPU @ 2.10 GHz and 64GB RAM. We employed a pre-trained model provided by our team. Our diffusion model network is trained by Adam optimizer
42
and the learning rate is set to 1 × 10−4. The number of epochs is 300 and batch size is 1. The total training time for the diffusion model is about 300 h. The distance between the scanner rotation center and the X-ray source is set to 300 mm. The projections in each view are collected using a linear detector, which consists of 1024 bins with a size of 0.124 mm. The spectra of the simulation data under tube voltages of 80 kVp and 140 kVp were obtained by using the ImaSim software. The corresponding spectrum distributions are shown in Figure 2. Through simulating the fast kVp switching scan, the ray paths corresponding to the two spectrum were inconsistent. In the experiment, 90 views or 120 views evenly distributed among 360 views are selected to validate the performance of the suggested algorithm. In this work, parameters

The experimental data include images from the LIDC dataset slices in (a) and (b), with (c) is synthesized CT image from segmented image by local hospital with display window [0.001 0.03], and (d) is the real mouse data reconstructed from full energy spectra with display window [0 0.08]. Figure (e) denotes normalized X-ray spectra and (f) the attenuations coefficients of bone and tissue.
Data preparation
To evaluate the performance of the algorithm proposed in this paper, extensive experimental verifications were conducted using preclinical data as shown in Figure 2. This preclinical data includes two datasets: one is a thoracic dataset pre-segmented and labelled by a local hospital, covering 1591 thoracic images from 5 patients. The primary basis materials within the images include bones and soft tissues, with the segmentation tasks mainly assisted by radiologists from the local hospital. The radiologists manually corrected the labelled images for each pair of bone and tissue using the bone removal function in the Syngo. Via software, equipped with the Siemens SOMATOM Definition
Flash CT scanner. The LAC of bone and tissue materials can be obtained from the National Institute of Standards and Technology. The other is the publicly available LIDC-IDRI lung dataset, which contains lung images from 20 patients. Due to the lack of accurate basis material image segmentation within the LIDC dataset, it was unfeasible to use them to validate the algorithm performance. Instead, we selected 50 images from the pre-segmented thoracic dataset of the local hospital for verification.
The real experimental mouse data as shown in Figure 2(d) were acquired using a spectral CT imaging system at the Institute of High Energy Physics, Chinese Academy of Sciences. Post euthanasia, mouse-tail specimens injected with iodine contrast agent were scanned with fixed source-to-object and source-to-detector distances of
Experiment results
In the numerical simulation experiments, we first validated the proposed algorithm using the pre-segmented and labelled dataset by the local hospital. The reconstructed image results and numerical evaluation metrics are presented in Figure 3 and Table 1, respectively. Figure 3 illustrates the reconstructed image of the second slice under 120 sparse views. It is evident that the basis material image reconstructed by the FBP-decomp exhibits significant noise and numerous streak artifacts. Although the OPMT method improves the quality of the material decomposition image, edge blurring issues remain. The IRM-MI method reduces artifacts compared to the two-step decomposition method and provides clearer structures compared to the OPMT method; however, soft tissue parts remain difficult to discern clearly. The GAN-MD model exhibits limited space artifact suppression capability, preserving most structural details in reconstructed images but introducing noticeable streak artifacts. Compared to the aforementioned methods, the ESART-TV method shows remarkable effectiveness in noise suppression and artifact removal, yet when examining the magnified regions of interest (ROI) in soft tissue basis materials, the details appear overly smooth, and the decomposed soft tissue material image contains noise. The proposed algorithm in this study achieves superior decomposition results compared to the ESART-TV method, producing synthesized VMIs closest to the reference images with minimal noise and more realistic textures.

Reconstructed material images on the local hospital simulated clinical data from 120 projections. The 1st-7th columns stand for the reference, FBP-decomposition, ESART-TV, OPMT, IRM-MI, GAN-MD and our proposed methods. The 1st row represent bone material and the corresponding ROIs, and the 2nd row represent soft tissue material and the corresponding ROIs. The 3rd row represent virtual monochromatic images at 75 keV.The display windows are [0.0001 0.8], [0.1 0.88], and [0.001 0.035], respectively.
Statistical reconstruction PSNR/SSIM (10−3) for different algorithms on the local hospital clinical data from 120 projections (50 test slices).
The bolded data in this table represent the corresponding evaluation metrics of the algorithm proposed in this paper, highlighted solely for the purpose of emphasizing its performance indicators.
As indicated in Table 1, the numerical evaluation metrics for the bone material images significantly outperformed those of the soft tissue basis material images, primarily due to the simpler structure of bone materials leading to better decomposition results. In contrast, the complex structure of soft tissues, especially when using the FBP-decomp, results in ineffective artifact removal and poorer evaluation metrics. The proposed algorithm achieved the best evaluation metrics, consistent with the visual observation results.
The results of material image reconstruction on LIDC dataset are presented in Figure 4. The FBP-decomp results still exhibit noticeable artifacts. The OPMT produces soft tissue material images with blurred boundaries. The IRM-MI also retains some noise, which is particularly evident in the magnified ROI. The GAN-MD algorithm exhibits significant noise, resulting in degraded texture clarity. Visually, both the ESART-TV and the proposed method exhibit better performance in material image, effectively preserving tissue structures and image details. However, pixel line profiles drawn along the yellow dashed lines in Figure 4 reveal significant deviations between the pixel intensity values of ESART-TV and the ground truth, whereas the values from the proposed method closely match the true values. The proposed algorithm accurately maintains edge and texture information while removing sparse artifacts. The deviations of pixel values from the material ground truth for other comparison algorithms can also be visualized in Figure 5.

Reconstructed material images on LIDC data from 120 projections. The 1st-7th columns stand for the reference, FBP-decomposition, ESART-TV, OPMT, IRM-MI, GAN-MD and our method. The 1st row represent bone material and the corresponding ROIs, and the 2nd row represent soft tissue material and the corresponding ROIs, with the gray-scale value ranges [0.0001 0.9], and [0.1 0.99], respectively.

Line profile of different reconstruction methods corresponding to the yellow dotted line in figure 4.
As shown in Table 2, the proposed algorithm achieves the optimal values for all three evaluation metrics. Compared to the ESART-TV method, the PSNR of the reconstructed bone material images improved by 7.01%, and the RMSE decreased by 58.54%, demonstrating the superior performance and effectiveness of the proposed algorithm in material image reconstruction.
Reconstruction PSNR/SSIM of LIDC data by different methods from 120 views.
The bolded data in this table represent the corresponding evaluation metrics of the algorithm proposed in this paper, highlighted solely for the purpose of emphasizing its performance indicators.
To comprehensively validate the performance of the proposed algorithm, we conducted experiments using 90 sparse views. The basis material images under 90 sparse views on the LIDC dataset were presented in Figure 6. As shown in Figure 6, the FBP-decomp method produces poor-quality material images, making it difficult to clearly discern tissue structures and details. The OPMT and IRM-MI both reduce artifacts compared to the FBP-decomp but still exhibits significant noise. The GAN-MD algorithm yields favorable error images, yet exhibits blurred structural details that obscure texture discernibility, as visually confirmed in the corresponding ROI images. The ESART-TV algorithm provides relatively clear reconstructed images; however, it suffers from excessive smoothing, as evidenced in the magnified ROI. Additionally, noise residues are clearly visible in the error images of the second and fourth rows. In contrast, the proposed algorithm effectively eliminates artifacts and suppresses noise. The magnified ROI show that the reconstructed images well preserve structural details and remain sharp edges. Table 3 lists the quantitative evaluation metrics of different decomposition methods under 90 sparse angles, which are consistent with Figure 6.

Reconstructed material images on the LIDC data from 90 projections. The 1st-7th columns stand for the reference, FBP-decomposition, ESART-TV, OPMT, IRM-MI, GAN-MD and our methods. The 1st row and the 3rd row represent bone material and soft tissue material and the corresponding ROIs. The 2nd and 4th row represent the error maps. The gray-scale value ranges for bone and soft maps are [0.0001 0.9], and [0.1 0.99], respectively.
Reconstruction PSNR/SSIM of preclinical LIDC data by different algorithms from 90 views.
The bolded data in this table represent the corresponding evaluation metrics of the algorithm proposed in this paper, highlighted solely for the purpose of emphasizing its performance indicators.
Real data experiments
This section evaluates algorithm performance using real mouse data. Figure 2(d) displays the full-spectrum CT reconstruction obtained through FBP. Comparative material decomposition results for bone and soft tissue are shown in Figure 7, where all algorithms demonstrate effective streak artifacts suppression. Our ROI analysis provides qualitative assessment of the artifact suppression and detail preservation capabilities of the different algorithms. The FBP-decomp reconstruction exhibits the most severe noise contamination with poorly resolved structures. Both OPMT and IRM-MI reconstructions exhibit significant edge blurring, while ESART-TV introduces noticeable blocky artifacts that compromise anatomical detail. Unlike the simulation training data, real scanned data present additional challenges including pronounced noise and beam hardening artifacts. These factors contribute to substantial residual artifacts in the deep learning-based GAN-MD reconstruction. As evidenced in Figure 7, despite being trained exclusively on simulation data, our proposed method achieves significantly better performance than GAN-MD on real mouse data. While GAN-MD reconstructions contain noticeable noise and artifacts, our method demonstrates superior artifact suppression while maintaining excellent preservation of texture details and edge structures. This performance gap highlight our model's enhanced generalization capability in real-world imaging scenarios.

Reconstructed material images on the real data from 120 projections. The 1st-6th columns stand for FBP-decomposition, ESART-TV, OPMT, IRM-MI, GAN-MD and our methods. The first row and the second row represent bone material and soft tissue material and the corresponding ROIs. The gray-scale value ranges for bone and soft maps are [0.01 0.8], and [0.1 0.7], respectively.
The red boxes (A-D) in Figure 7 denote the ROIs for quantitative evaluation of reconstructed image quality. Regions A and B contain rich structural details, while regions C and D are smooth and homogeneous. Since the ground truth of real data reconstruction is unavailable, we employed two metrics for algorithm performance assessment: mean equivalent attenuation coefficient (for reconstruction accuracy) and noise standard deviation (SD, for noise suppression capability). Table 4 presents the comparative results of these metrics across different algorithms. Our method achieved mean equivalent attenuation coefficients comparable to other approaches, demonstrating its accuracy in spectral CT reconstruction of real data. Regarding noise standard deviation, while GAN-MD yielded the highest values, our algorithm maintained finer image details while achieving lower noise levels. The proposed method achieved optimal balance between reconstruction fidelity and noise reduction.
Mean values and SDs of ROI A, B, C, and D in figure 7.
The bolded data in this table represent the corresponding evaluation metrics of the algorithm proposed in this paper, highlighted solely for the purpose of emphasizing its performance indicators.
Ablation experiments
To optimize material image reconstruction fidelity, this study innovatively integrates deep priors from the probability diffusion model (DM) into the iterative reconstruction framework. The experimental results from comprehensive simulations demonstrate the algorithm's superior capability in material decomposition. Ablation studies contrasting implementations with and without DM integration reveal critical performance difference. As illustrated in Figure 8, the baseline method (without DM) exhibits pronounced artifacts in bone material decomposition and blurred soft-tissue boundaries, whereas the DM-enhanced approach effectively suppresses noise while preserving structural edges, particularly in arrow-indicated regions. Quantitative analysis in Table 5 confirms these improvements, showing an 11.76% reduction in bone material RMSE and 4.85% increase in soft-tissue PSNR compared to the algorithm without the probability diffusion model implementations, thereby robustly validating the DM's capability to enhance decomposition accuracy while maintaining projection data consistency.

Reconstruction results from the ablation experiment from 90views. From left to right are: the reference, with the diffusion model, and no diffusion model. The gray-scale value ranges are [0.0001 0.9], and [0.1 0.99], respectively.
Reconstruction PSNR/SSIM of ablation experiment.
Discussion
To address the limitations of model-based material decomposition methods, we propose an unsupervised direct material decomposition framework that leverages deep priors. This approach synergistically optimizes VMIs through a diffusion probabilistic model in sparse-view spectral CT. This deep prior mechanism facilitates structural detail preservation by exploiting a comprehensive statistical distribution analysis of data. Through precise modeling of intrinsic image statistics, our framework simultaneously accomplishes two critical objectives: (1) robust image restoration through artifact suppression, and (2) fundamental recovery of essential physical attributes with preserved edge integrity. The proposed method employs VMIs as critical discriminative enhancers for polychromatic projections, effectively overcoming convergence limitations in iterative reconstruction algorithms. By integrating probabilistic diffusion priors into the optimization process, our framework achieves superior refinement of material-specific representations. It systematically enforces a data fidelity term to ensure measurement consistency, while applying probabilistic regularization to suppress unwanted structures–thereby guaranteeing anatomically plausible material image reconstruction. The results of the experiments on preclinical data and real data demonstrate the algorithm's robustness under challenging conditions, including geometric inconsistencies and sparse sampling, while maintaining enhanced reconstruction fidelity.
However, while diffusion models (DM) demonstrate remarkable reconstruction capabilities, they inherently suffer from low sampling efficiency due to their iterative refinement mechanism. Conventional implementations typically require hundreds to thousands of sequential steps to gradually denoise Gaussian noise into high-fidelity images, with each step demanding computationally intensive neural network predictions. As evidenced in Table 6, this cumulative computational burden results in slightly higher time costs for our proposed algorithm in this work compared with ESART-TV, posing practical challenges for real-time clinical applications. To address this efficiency bottleneck, we implement a truncation diffusion model activation strategy within the iterative reconstruction workflow. This approach strategically restricts diffusion model-based optimization program to the final stages of the reconstruction process. Empirical analysis in Table 7 reveals that material-specific RMSE values for both bone and soft tissue exhibit minimal fluctuations and insignificant changes during these final rounds. With a reduction in the number of truncated iterations, computational efficiency increases, but the evaluation metrics decline. Specifically, the PSNR and SSIM values peak when the iteration count is set to 50, indicating optimal reconstruction performance at this point. Therefore, this study determines to introduce the optimization strategy of the DM in the last 50 rounds of iterations.
Iteration time of different methods (s).
Reconstruction PSNR/SSIM of different truncation iterations.
This paper proposes a novel framework for material image reconstruction that leverages direct material decomposition in the projection domain. This approach unifies the nonlinear observation model of polychromatic projections with the fundamental physics of material attenuation. By integrating spectral sensitivity modeling and decomposition priors into a single optimization paradigm, the proposed methodology eliminates cascading approximations inherent to traditional two-step approaches (e.g., projection-domain decomposition followed by image reconstruction). Consequently, it mitigates cumulative information loss while enhancing decomposition accuracy. Notably, the current DMDM implementation assumes precise a priori spectral calibration—a condition rarely satisfied in practical scenarios due to spectral drift caused by detector imperfections, source variability, or object scatter. Experimental validation on real-world datasets reveals that spectral estimation errors induce non-negligible biases in decomposition outcomes. This spectral dependency underscores a critical research frontier: advancing robust spectral self-calibration techniques or hybrid learning-physics architectures to relax stringent spectral knowledge requirements while preserving decomposition fidelity.
Conclusion
In this study, we propose a novel spectral CT direct material decomposition framework enhanced by a diffusion probability model to address the challenges of sparse-view spectral CT imaging, particularly in scenarios involving fast kVp switching protocols. To overcome the inherent limitations of slow convergence in iterative spectral CT reconstruction, we innovatively integrate virtual monochromatic images as spectral discriminators, which expand the angular coverage of the spectral surface equations and dynamically refine the reconstruction manifold. By synergistically optimizing virtual monochromatic images through the probability diffusion model-guided denoising, our framework achieves improvement in material decomposition accuracy compared to conventional two-step approaches. Comprehensive validation with preclinical data across diverse sparse-view configurations (90–120 projections) demonstrates consistent performance for both bone and soft-tissue materials. These results not only validate the framework's robustness against spectral inconsistency and noise amplification but also establish a paradigm for physics-informed diffusion models in spectral imaging. Future work will focus on relaxing spectral calibration dependencies through self-supervised spectral estimation, further bridging the gap between simulation-driven innovation and clinical applicability.
Footnotes
Acknowledgements
This work was supported by the National Natural Science Foundation of China (Grant No. 62271504, 62101596) and Natural Science Foundation of Henan (Grant No.252300420395). This work was also supported by Technology Innovation Leading Talent Project of Zhongyuan (Grant No. 244200510015).
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Research and development of early defect insitu detection technology and intelligent warning equipment for key components of ultra-high/extra high voltage electrical equipment Key Research and Development Special Project of Henan Province, China under Grant number 251111220600.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Appendix I
In this section, we provide the specific expressions for the coefficients in the first-order Taylor expansion, and the linearized representation of the DECT imaging model is as follow,
For the sake of simplicity, Eq. (24) is rewritten as a linear system in the following form,
The VMIs-guided method increase the angles between polychromatic projection curves by introducing monochromatic images, and thus accelerating the iterative process. Specifically, two distinct energies
Eq. (13) can be equivalently represented as the following,
Applying the projection matric A to the above Eq.(28), and denote
Similarly, if the n -th iteration result of the VMIs is known as
Eq. (30) can be equivalently represented as the following linear system,
