Abstract
Compared with solid-colored fabrics, the textures in yarn-dyed fabric images are more complex, making the task of defect detection more challenging. To achieve efficient detection, this study proposes an automatic detection framework for dyed fabric defects. The proposed framework consists of a hardware system and a detection algorithm. For efficient and high-quality acquisition of fabric images, an image acquisition assembly equipped with three sets of light sources and a mirror was developed. In addition, a defect detection algorithm based on Fourier convolution and a convolutional autoencoder is proposed. Abandoning the common way of adding noise, this paper proposes to generate image pairs for training using a random masking method in the training phase. In the autoencoder, some traditional convolutional layers are replaced with Fourier convolutional layers. Ablation experiments verify the effectiveness of the mask generation method and Fourier convolution. Compared with other defect detection methods, the proposed method achieves the best performance, which verifies the superiority of the method. The maximum detection speed of the developed system can reach 41 meters per minute, which can meet real-time requirements.
Keywords
The occurrence of fabric defects will seriously impair the quality, appearance and performance of fabric, thereby reducing the profit of the manufacturer, especially for yarn-dyed fabrics. Therefore, defect detection is an important step in the quality control of fabric products. At present, the most common defect detection method in textile enterprises is manual inspection, which is mainly based on the subjective experience of cloth inspectors. Manual inspection suffers from low efficiency and high labor cost, and is already unsuitable for modern textile manufacture. Replacing manual detection with computer-vision-based detection is an inevitable trend of automation and information for the textile industry. Thus, fabric defect detection methods based on computer vision 1 have garnered much attention,2–4 and have become a research hotspot in both industry and academia. Yarn-dyed fabrics, which are woven from dyed yarns according to specific rules and combinations, are very popular in the fields of clothing and home textile. It is precisely because of the richness of colors that the task of defect defection of yarn-dyed fabrics is very difficult and challenging.
To achieve automatic fabric detection, many methods 5 have been proposed. These methods can be mainly divided into statistical, spectral, model-based, learning-based and hybrid methods. The statistical methods6,7 employ various statistical properties of texture and defects to estimate defects. However, the diversity of fabric textures and defect shapes seriously affects the detection accuracy of such methods. In particular, it is very expensive to design different statistical indicators for defects of different complexity. Therefore, statistical methods have great limitations in actual fabric defect detection, that is, these methods are difficult to directly apply to yarn-dyed fabric defect detection due to the complexity of color and texture. The spectral methods8,9 convert an image in the spatial domain to the frequency domain, and achieve the detection of defects in the fabric by using the strong periodicity in the fabric image. However, such methods do not work well when the contrast between defect areas and defect-free areas is low or when the defects are small. The model-based methods10,11 represent fabric texture as a stochastic process and assume that texture images can be viewed as samples generated by stochastic processes in the image space. Defect detection is treated as a hypothesis testing problem with statistics from the model. Such methods usually have a large computational overhead, and thus cannot meet the real-time requirements of detection. However, if a model-based algorithm is introduced into the defect detection of yarn-dyed fabric, a specific model for each texture is required, and the cost of each model is prohibitive.
From the above analyses, it is evident that most fabric defection methods are based on conventional feature engineering. These methods are heavily dependent on the feature extraction algorithm, which makes the methods less robust. Consequently, many researchers have start to combine feature engineering with machine learning algorithms to improve detection efficiency. Even though these methods12,13 can achieve good performance on some solid-color fabric defect detection, it is difficult to achieve good performance on yarn-dyed fabrics with complex textures and colors.
Recently, significant progress14–16 has been made on image analysis by moving low feature-based algorithms to deep learning-based end-to-end frameworks. Many researchers17,18 use deep learning technology to solve the problem of fabric defect detection. Compared with earlier combined methods, deep learning-based methods can extract higher-level features of images. However, most deep learning-based methods are supervised, that is, the learning of the model needs to be driven by a large amount of labeled data. Li et al. 19 first introduced deep learning technology in the field of fabric defect detection by proposing an autoencoder model. By treating the defect detection task as an object detection task, Jing et al. 20 used the improved YOLOv3 software to achieve efficient detection of six classical defects. These supervised deep learning-based methods required a large number of precisely labeled fabric samples. In practical applications, it is difficult to obtain a large number of accurately labeled yarn-dyed fabric samples, which makes it difficult for supervised methods to be applied to yarn-dyed fabric defect detection.
Many recent researches have shown that the unsupervised deep learning-based methods can achieve good performance without labeled data. Liu et al. 17 proposed a multistage generative adversarial network for fabric defect detection. Mei et al. 21 used an autoencoder network to recognize fabric defects. These methods can achieve good performance on plain fabrics without a large amount of training data, but their recognition rates for yarn-dyed fabrics with complex textures are relatively low. Therefore, there is still a great deal of room for improvement in the defect detection of yarn-dyed fabrics with complex textures.
Compared with solid fabrics, yarn-dyed fabric images contain more complex texture features. Generally, the appearance of defects in yarn-dyed fabrics will destroy their periodicity to varying degrees. Using this rule, a novel defect detection method based on Fourier convolution 22 and a convolution autoencoder (CAE)23,24 is proposed in this paper, which is abbreviated to FCAE. The motivation of the proposed method can be summarized as follows: (1) Fourier convolution can effectively extract periodic features of fabric images; (2) the CAE is trained in an unsupervised manner, so it does not require a large amount of annotation information. Specifically, to achieve the automatic detection of yarn-dyed fabric defects, a defect detection framework is developed, which mainly includes a hardware system and a detection algorithm. The following sections will introduce the key components in both systems separately.
Hardware system
In this section, the key components of the hardware system are introduced in detail. Figure 1 shows the overall diagram of the developed equipment, which consist of the unwinding mechanism, traction mechanism, winding mechanism, image acquisition component and computer. The frequency conversion motor realizes the unwinding, pulling and winding of the cloth by controlling the rotation of the roller. When the cloth passes through the image acquisition area, the camera automatically captures the fabric image and sends it to the software system in the computer for detection, as shown in Figure 2. Apart from the image acquisition component, the developed equipment is similar to other automatic defect inspection equipment. Therefore, this section focuses on the introduction of the image acquisition component.

Hardware system: (a) overall diagram of the equipment; (b) front view of the equipment and (c) rear view of the equipment.

The internal structure diagram of the developed automatic cloth inspection equipment.
In real-time inspection, the choice of camera is an important factor to obtain high-quality fabric images. There are two types of industrial cameras commonly used in defect detection: line-scan cameras and area-scan cameras. This paper studies yarn-dyed defect detection technology on the basis of surface images, so the area-scan camera is selected as the image acquisition device. In the developed equipment, eight industrial cameras (MER-502-79U3M) are arranged linearly, which can realize the rapid acquisition of fabric images, as shown in Figure 3(a). To ensure stability, the lighting system shown in Figure 3(b) is designed, which contains three light sources and a reflector. Light sources 1 and 2 enhance the texture of the fabric surface from two angles, respectively, while light source 3 acts as a transmitted light source to enhance the outline of the fabric. With the cooperation of the three sets of light sources, the reflector can obtain a complete fabric image and transmit it to the camera.

Image acquisition components: (a) field of view with eight cameras and (b) light source configuration.
In the experiment, the size of the fabric image captured by each camera is 2048 pixel × 2048 pixel, which corresponds to the actual size of the fabric of 28.85 cm × 28.85 cm (71 pixels/cm, 0.141 mm/pixel). The width of the overlapping area between the images captured by adjacent cameras is about 1.6 cm. The equipment can realize defect detection of fabrics with a maximum width of 2.2 m.
Detection algorithm
In this section, the procedures of the proposed detection algorithm are introduced in detail. Figure 4 shows the overall architecture of the proposed framework in the training phase and testing phase. Procedures in the training phase mainly aim to optimize the parameters in the FCAE model. In the testing phase, the following procedures are undertaken: (1) use the trained FCAE to reconstruct the input fabric image; (2) calculate the residual between the reconstructed image and the original image; (3) judge whether there are defects in the image. Specific illustrations are presented as follows.

The overall architecture of the proposed framework. The proposed framework consists of two phases, the training phase and the testing phase.
Training phase
Procedures in the training phase aim to optimize the parameters in the FCAE so that the model can better reconstruct the fabric images. Specifically, the training phase mainly includes mask generation, patch generation and model training.
Mask generation
The appearance of defects will destroy the original texture structure of the image to varying degrees. Many related works add noise, such as salt and pepper noise and Gaussian noise, to fabric images to simulate the damage effect of defects on the image texture. However, defects only appear in local areas of the image, and it is difficult to simulate the effect of defects by adding noise globally. To simulate more realistic defects, a simple mask generation method is proposed. Specifically, (1) firstly, randomly select two points in the image area as the start and end points; (2) use a brush whose color and width can change randomly to connect the start and end points. It is stated here that the direction can be randomly changed between zero and three times during the connection process. As shown in Figures 5(a)–(d), the generated masks have a certain similarity to the defects in morphology. The generated mask will be added to the original image by adding the corresponding values. Several examples are shown in Figure 2, from which it can be seen that the texture damage of the fabric image using the proposed mask is similar to that of real defects.

Four examples of mask generation: (a)–(d) randomly generated masks; (e)-(h) the effect of the mask on the original fabric.
Patch generation
The purpose of this phase is to create training data for the proposed FCAE model. In Figure 4, “Create Patches” is patch generation, which is introduced here. For pixel-wise prediction, characterizing pixels based on local neighborhood information may be more robust than using just a single pixel. Nevertheless, the size of the patch greatly affects the training efficiency. A small patch size will lead to a large computational overhead. In the proposed framework, the input fabric image is equally divided into four patches, which are then concatenated together for input to the model.
Model training
The purpose of model training is to optimize the parameters in the proposed FCAE. As shown in Figure 4, the architecture of the FCAE model is based on an encoder–decoder paradigm. Referring to the CAE, the FCAE contains three downsampling and upsampling operations. Three downsampling layers and several convolutional layers constitute the encoder, and three upsampling layers and several convolutional layers constitute the decoder. As we all know, the appearance of defects will destroy the original texture structure of the fabric image, which is more obvious in the frequency domain space. To enhance the learning ability of the model for fabric image texture, the Fourier convolutional layer (FCL) is designed, the structure of which is presented in Figure 6. Conceptually, the FCL is comprised of two inter-connected paths: the left-hand path that conducts ordinary 1 × 1 convolutions on the input feature channels, and the right-hand path that operates in the spectral domain. Each path can capture complementary information with different space. Information fusion between the two paths is performed through subsequent 3 × 3 convolutions. The Fourier transformation can be formulated as

The structure of the proposed Fourier convolutional layer.
During training, for an input defect-free fabric image A, a randomly generated mask is firstly attached in a point-to-point manner to obtain an image B whose texture is damaged. The target of the FCAE is to remove the mask and reconstruct the image (Afake represents the reconstructed fabric image). The input of the FCAE is a defect-free fabric image with the generated mask, and the output is the reconstructed image without the mask. The objective function is defined as follows
Testing phase
The goal of the testing phase is to judge the presence or absence of defects in the input fabric image to be tested. The unknown fabric image is first reconstructed by the trained FCAE to repair abnormalities. Then the residual map can be computed by subtracting the grayscale reconstructed image from the original grayscale image. Inevitably, there will be many isolated noises in the residual map. The solution proposed in this paper is to first binarize the residual image with Otsu thresholding, and then use the open operation to remove isolated noises. The procedures in the testing phase are shown in Figure 7. Finally, whether there is a defect is predicted by calculating the area of the salient region in the binary image.

The procedures in the testing phase. FCAE: Fourier convolution autoencoder.
In this study, the size of the input fabric image is set to 512 × 512 × 3. As mentioned above, the size of the image captured by each image is the equipment is 2048 × 2048 × 3. The real-time detection strategy proposed in this paper is as follows: (1) divide each collected image into four sub-images with a size of 512 × 512 × 3, and record the positions of all sub-images; (2) input all 32 sub-images as a mini-batch to the model for reconstruction.
Results and discussion
In this section, some experiments are carried out to demonstrate the effectiveness of the proposed method. Firstly, the used dataset, evaluation criteria and implementation details are introduced. Then, several sets of experiments are presented to evaluate the performance of the proposed method. Detailed descriptions are introduced below.
Dataset and implementation
Learning-based methods summarize and generalize regularities from data, so the dataset available for training is essential. Furthermore, for a fair comparison, a public dataset named YDFID-1, 27 which contains 17 different classes of yarn-dyed fabrics, is used in the experiments. Specifically, 6378 defect-free images and 312 defect images were collected in YDFID-1, with a resolution of 512 × 512 × 3. In this study, 3189 defect-free images are used as the training set, and the other images are used as the testing set. According to the pattern of yarn-dyed fabrics, the fabrics in the dataset are divided into three categories, namely simple lattices (SLs), stripe patterns (SPs) and complex lattices (CLs). Figure 8 presents some samples of defective yarn-dyed fabric images, where the defects are marked with red arrows. Compared with solid-colored fabrics, yarn-dyed fabrics contain more complex backgrounds, making the task of defect detection more difficult.

Some samples of defective yarn-dyed fabric images in YDFID. (Color online only.)
In this study, the proposed method is implemented by using Pytorch toolkit 1.9.0 + CUDA11.4 +cuDNN8.2.1. The hardware environment is as follows: CPU = E5 2623V4@2.60 GHz, RAM = DDR4 32G, GPU = GeForce RTX 3090(24G) × 2. As shown in Figure 9, the proposed model starts to converge after 100 epochs of training on the YDFID dataset.

The change curve of loss during training.
Evaluation criteria
The evaluation criteria utilized in our experiments include two aspects: pixel-level and image-level performance metrics. The former is used to quantitatively evaluate the segmentation and localization effects of different methods for defects. As shown in Figure 10, TPp refers to the number of pixels in the defective area that are correctly predicted and FPp refers to the number of pixels in the defect-free region that are incorrectly predicted; TNp and FNp have similar meanings. Then, the three criteria can be computed by

Definitions of TNp, FNp, TPp and FPp.
It can be found that the F1-score is a comprehensive indicator that combines recall and precision.
The image-level criteria measure the accuracy of predicting a fabric image as defective or defect-free. Here, three widely used evaluation metrics are employed, namely the detection rate (DR), false alarm rate (FR) and detection accuracy (DACC). The three metrics are computed as follows
Definitions of TP, FN, FP and TN in fabric defection
Ablation study
There are two main innovations in the FCAE model: mask generation and the FCL. To demonstrate their effectiveness, this section includes the results of a conducted ablation study.
The purpose of mask generation is to construct image pairs for training. At present, the commonly used method for generating image pairs is to randomly add noise, such as salt and pepper noise and Gaussian noise, to the image. In this experiment, the pairwise image generation method is replaced by adding different noises for the FCAE while other configurations of the FCAE remain unchanged. The comparison results are shown in Figure 11 and Table 2. It is clearly observed that defects in yarn-dyed fabric images can be completely segmented using the proposed mask generation method. Moreover, the method of mask generation outperforms the others on quantitative criteria. The main reason for these results is that the randomly generated mask can effectively simulate the shape of defects in yarn-dyed fabric images, so that the model can effectively reconstruct the image of defects and detect defects.

Detection results using different pairwise image generation methods.
Quantization results corresponding to different image pair generation methods
The best results in the table are marked in bold.
The FCL is designed to incorporate texture distribution information of fabrics into the model training process. To demonstrate its effectiveness, a comparative experiment was conducted with the CAE, where the CAE is the FCL in the FCAE replaced by the ordinary convolutional layer. The experimental results are shown in Table 3, from which it can be found that the performance of the FCAE is significantly better than that of the CAE, indicating that the proposed Fourier convolution can effectively improve the reconstruction performance of fabric images with strong periodicity.
Comparative experimental results of the Fourier convolution autoencoder (FCAE) and the convolution autoencoder (CAE)
The best results in the table are marked in bold.
In this study, the image size input to the model is also a very important factor. To verify the rationality of the size selection, we compare the performance of models with different input sizes. The experimental results are presented in Table 4. It can be observed that a smaller input size can achieve better detection performance. The reason for this result is that the model has a better reconstruction effect for images of smaller sizes. Although a smaller input size means fewer model parameters, multiple batches are required to process a fabric image of size 2048 × 2048 × 3, so it takes more time to detect an image. Considering the time complexity and accuracy comprehensively, this study chooses 512 as the input size of the model.
Comparative experimental results of different input sizes
The best results in the table are marked in bold.
Comparisons
To further validate the performance of the proposed FCAE method, this section qualitatively and quantitatively compares the performance of the proposed method with five other methods, namely sparse dictionary learning (SDL), 28 frequency domain saliency (FDS), 29 the denoising convolutional autoencoder (DCAE), 30 the multi-scale denoising convolutional autoencoder (MSDCAE) 21 and the U-shaped denoising convolutional autoencoder (UDCAE). 31 SDL is an adaptive fabric defect detection method based on SDL and a grid search for a different texture. When implementing SDL, the grid size is configured as 32 × 32 pixels. FDS is an unsupervised fabric defect detection algorithm based on the human visual attention mechanism. The DCAE, MSDCAE and UDCAE are all defect detection methods for fabrics based on the CAE. It is stated here that all the above methods are based on deep convolutional neural networks to learn fabric image reconstruction methods in an unsupervised manner. In this experiment, all methods are implemented in the same software and hardware environment.
Some qualitative results are shown in Figure 12. SDL detects defects in fabric images by a grid search, in which the size of the grid seriously affects the detection results. It can be found that many false detections appear in the results of SDL. FDS can accurately detect defects in simple yarn-dyed fabrics, while it is not ideal for yarn-dyed fabrics with complex backgrounds. The reason for this result is that the saliency of the background is higher than that of the defect. The detection performance of the four variational auto encoder (VAE)-based methods is better than that of the previous two methods. However, the DCAE, MSDCAE and UDCAE have different degrees of over-detection and missed detection. The experimental results show that the FCAE achieves the best performance. Even for the fourth sample, which is difficult to segment, the FCAE successfully locates the defect area, which is highly similar to the ground-truth.

Comparison of the detection results for yarn-dyed fabrics. SDL: sparse dictionary learning; FDS: frequency domain saliency; DCAE: denoising convolutional autoencoder; MSDCAE: multi-scale denoising convolutional autoencoder; UDCAE: U-shaped denoising convolutional autoencoder; FCAE: Fourier convolutional autoencoder.
Table 5 presents the quantitative comparison results of the six methods. The FCAE achieves the best performance in all indicators. It is obvious that CAE-based methods achieve better performance than other methods, especially for image-level metrics, which verifies that the CAE has a strong learning ability. Mask generation and Fourier convolution are introduced in the FCAE, which makes the model have stronger reconstruction ability for dyed fabrics with more complex textures, thus achieving better detection performance. It is worth noting that the false alarm rate of the proposed method is only 4.2%, which can meet the requirements of industrial detection, while those of the other methods are all higher than 10%. In summary, the proposed method has good performance for the defect detection of dyed fabric images, which verifies the superiority of the FCAE.
Comparative experimental results
The best results in the table are marked in bold.SDL: sparse dictionary learning; FDS: frequency domain saliency; DCAE: denoising convolutional autoencoder; MSDCAE: multi-scale denoising convolutional autoencoder; UDCAE: U-shaped denoising convolutional autoencoder; FCAE: Fourier convolutional autoencoder.
Time complexity analysis
In industrial applications, detection algorithms are expected not only to have high accuracy, but also to be fast enough to meet real-time requirements. The number of parameters determines the computational cost of the model. In general, the larger the amount of data, the longer the model needs to be trained, and the results in Table 6 are highly consistent with this rule. Although the training time is quite different, the time-consumption of the four methods in the testing phase is very close, which is slightly higher than 0.1 s. In the testing phase, the strategy proposed in this paper is to treat all the images collected by the eight cameras as a mini-batch and feed them into the FCAE for detection at the same time. It is actually measured that the FCAE takes 0.42 s to recognize 32 images at the same time. Combined with the FCAE, the maximum detection speed of the developed equipment can reach 41 m/min, which can meet the real-time requirements of detection.
Time complexity comparison of different methods
The best results in the table are marked in bold.DCAE: denoising convolutional autoencoder; MSDCAE: multi-scale denoising convolutional autoencoder; UDCAE: U-shaped denoising convolutional autoencoder; FCAE: Fourier convolutional autoencoder.
On-machine testing
The light sources at three different positions in Figure 2 have certain differences. In this section, we test the impact of different light source configurations on machine detection performance. The three light sources in Figure 3 are abbreviated as LT1, LT2 and LT3. In this experiment, we tested five rolls of defective yarn-dyed fabric. The six metrics mentioned in the Evaluation criteria section are still used in the experiments to evaluate the algorithm performance.
The experimental results are presented in Table 7. Observing the first three rows of data in the table, it can be found that the performance of the model is not significantly different when only LT1 or only LT3 is turned on, while the performance of the model is poor when only LT2 is turned on. The data in the fourth to sixth rows show that turning off LT2 has a great impact on the model performance. Overall, the closing or absence of any light source will cause the performance of the model to decrease, which verifies the rationality of the light source configuration of the developed equipment.
Experimental results of on-machine testing under different light source configurations
The best results in the table are marked in bold.
Combined with the proposed detection algorithm and developed equipment, the detection speed can reach 39.8 m/min, which can meet the real-time requirement of fabric defect detection.
Conclusion
In this paper, a novel automatic detection system for fabric defects was developed, which includes hardware systems and detection algorithms. In the hardware system, three light sources and one mirror are configured to achieve efficient and high-quality acquisition of yarn-dyed fabric images. By comprehensively analyzing the characteristics of yarn-dyed fabrics, a defect detection model, the FCAE, based on the CAE and Fourier convolution, is proposed. In the training phase, a random mask generation method is designed to generate paired data for FACE training, so the learning of the FACE is in an unsupervised manner. Four Fourier convolution layers are configured in the FCAE, which enhances its ability to learn fabric textures. Ablation experiments demonstrate that the training strategy and configuration of the FCAE can effectively improve its performance. Compared with other methods, the FCAE achieves the best performance in all indicators, which proves the superiority of the FCAE. Combined with the FCAE, the maximum detection speed of the developed equipment can reach 41 m/min, which can meet the real-time requirements of detection.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship and/or publication of this article: This work was supported by the National Natural Science Foundation of China (grant 61976105), in part by the National Key R&D Program of China (grant 2017YFB0309200) in part by Applied Research Project of Public Welfare Technology of Zhejiang Province(No. LGG21F030007), in part by China Postdoctoral Science Foundation (No. 2020M681736), in part by Scientific Research Start-up Project (No. 20195026), and in part by Scientific Research Project of University (No. 2019LG1006).
