Abstract
The inspection of underwater structural health is crucial for comprehensive bridge health assessments. In underwater structure imaging, traditional methods include single-camera and binocular camera inspection. However, due to water turbidity and long working distances with a small field-of-view, obtaining clear and high-quality detection images with these methods takes much work. To address this problem, this paper presents a method for planar array image stitching based on Harris corner point extraction, utilizing the advantages of planar array cameras characterized by short working distances and wide field-of-view. The core contribution of this paper is the introduction of an innovative image sequence stitching algorithm utilizing Harris corner point extraction and the combination of the first proposed planar array cameras with the image sequence stitching algorithm, which solves the problem of long distance and small field-of-view during the underwater inspection. The image stitching method involves calibrating camera parameters with a checkerboard and stitching underwater images from planar array cameras to reveal underwater structural features. Furthermore, five quantitative evaluation metrics and the method for calculating the field-of-view loss rate are presented to evaluate and analyze the stitched images. A series of experiments were performed on concrete surfaces, aquatic and underwater, with a total field-of-view of the underwater image after stitching of 358.86 mm × 319.24 mm at a working distance of 160 mm. Five evaluation methods were used to quantitatively evaluate the quality of the stitched images and calculate the field-of-view loss rate of the images. The results indicate that the proposed method improves the ability to inspect underwater. The stitched images achieve notable metrics: an entropy of approximately 6.7, an average gradient of about 1.7, a spatial frequency of around 3.5, an edge strength of about 17, mutual information of approximately 1.2, and a field-of-view loss rate of <0.1, facilitating more effective underwater structure inspection.
Keywords
Highlight
(1) An efficient stitching and inspection method based on Harris corner point extraction for planar array underwater images to detect surface features of underwater structures. (2) A novel underwater planar array camera system was proposed and the image stitching algorithm exhibits superior fusion performance and stitching quality. (3) Five quantitative metrics for evaluating the stitched image quality and a method to calculate the loss ratio of the image field-of-view. (4) At a working distance of 160 mm, a total field-of-view spanning 358.86 mm × 319.24 mm with a field-of-view loss rate below 0.1.
Introduction
After long-term service, underwater structures are inevitably affected by a combination of objective and subjective factors, resulting in varying degrees of damage (Medina et al., 2023; Nasr et al., 2020; Wang et al., 2016). Assessing these structures is paramount to ensuring their safety. The traditional method of assessing the condition of underwater structures primarily relies on visual inspection, typically conducted by engineering divers who rely on their naked eyes to observe the structure and assess its integrity (Lu et al., 2004). However, this method is hindered by underwater lighting conditions and the limitations of human divers, making it challenging to obtain clear images of the target.
Currently, inspection methods for underwater structures can be achieved through a variety of methods, primarily divided into acoustic inspection methods (Chen et al., 2021; Guerneve et al., 2018; Hou et al., 2022; Li et al., 2023; Shi et al., 2022), visual inspection methods (Bi et al., 2023; Hong and Kim, 2020; Wu et al., 2023; Zhuang et al., 2021) and inspection methods based on fiber optic sensors (Min et al., 2021). Acoustic inspection methods typically involve the use of remotely operated vehicles (ROVs) (Choi et al., 2018; Ludvigsen and Sorensen, 2016) or autonomous underwater vehicles (AUVs) (Cho et al., 2018; Wang et al., 2023) equipped with sonar imaging systems for identification and detection. In the inspection methods based on optical fiber sensors, Tan et al. (2021) presented a method to measure and visualize strains and cracks in high-performance fiber-reinforced concrete using distributed fiber optic sensors based on optical frequency domain reflectometry. Tan et al. (2023) investigated the interfacial mechanics for distributed fiber optic sensors undergoing debonding through mechanical analysis and metaheuristic-based inverse analysis. However, acoustic inspection methods encounter constraints such as the underwater environment, limited detection range, background noise interference, the existence of blind zones, and difficulties in data interpretation. Similarly, methods based on optical fiber sensors necessitate high installation and maintenance efforts for fiber optic sensors and are constrained in their applicability.
In contrast, visual inspection methods offer a more intuitive approach to underwater identification, with higher image resolution, detailed imagery, multi-modal perception, real-time feedback and control, and visualization capabilities (Wu et al., 2021). Therefore, methodologies based on visual imaging for underwater identification are rapidly advancing, offering effective solutions and application pathways for assessing underwater structures. Wang et al. (2023a) developed an automatic identification method for identifying underwater damage areas accurately and efficiently. Chen et al. (2019) proposed a novel monocular underwater calibration method for a simple underwater vision measurement system, while Chen et al. (2020) introduced a new underwater salient object inspection method combining 2D and 3D visual features.
Visual inspection methods can be categorized into single-camera and camera array methods based on the number of cameras utilized. The single-camera methods involve rotating and moving the camera to capture images of the entire structure, typically suitable for static scenes (Shum and Szeliski, 2000). On the other hand, camera array methods can simultaneously capture images from multiple cameras, applicable to both static and dynamic scenes, offering a more comprehensive and intuitive identification of underwater structures through image stitching algorithms (Qu et al., 2022). Hu et al. (2022) proposed a camera array calibration method to obtain a complete camera model and global mapping coordinates, while An et al. (2020) introduced a camera array calibration method based on a feature-based camera calibration method for spherical camera arrays. Li et al. (2024) presented a novel calibration method for multi-camera systems utilizing a ruler equipped with double-layer circularly encoded landmarks. Quevedo et al. (2017) proposed a super-resolution fusion algorithm based on a multi-camera environment, enhancing the quality of underwater video sequences without significantly increasing computation.
Utilizing planar array cameras can significantly improve the quality and clarity of underwater images by capturing them from various viewpoints simultaneously and then fusing and processing these images. This innovative approach to the underwater identification method offers an effective solution and an application pathway for underwater identification. Leveraging the characteristics of planar array cameras, such as the short distance and wide field-of-view, this paper proposes a rapid stitching and inspection method for planar array underwater images to detect surface features of underwater structures. This paper innovatively introduces an image sequence stitching algorithm grounded in the extraction of Harris corner points. The core contribution lies in the integration of this algorithm with planar array cameras, enabling the inspection of underwater structure surfaces in a novel manner. Quantitative evaluation metrics and a calculated field-of-view loss rate are employed to evaluate the stitched images. Experimental procedures, including image acquisition, distortion correction, and stitching, were conducted both aquatically and underwater on concrete surfaces. Evaluation metrics were utilized to evaluate the quality of the underwater stitched images, ultimately yielding a calculated field-of-view loss rate. At a working distance of 160 mm, this method produces a stitched image with a camera field-of-view of 245.5 mm × 185 mm, contributing to a total field-of-view of 358.86 mm × 319.24 mm, meeting requirements for long working distances and wide field-of-view. The stitched image demonstrates impressive metrics, including entropy, average gradient, spatial frequency, edge strength, and mutual information, indicating excellent image fusion and stitching quality. Additionally, the high resolution of the planar array cameras, with individual image pixels measuring 2592 × 1944, enables precise analysis with minimal field-of-view loss.
This paper is roughly divided into six parts. The overview of the planar array camera system hardware and the image stitching methodology can be obtained in Section 2. The performance of the planar array cameras is discussed, with an emphasis on the ability to detect underwater structural surfaces in Section 3. The image stitching process and the evaluation metrics used are explained in Section 4. Subsequently, the effectiveness and feasibility of the inspection methods in various underwater scenarios are assessed in Section 5. The limitations of the proposed method and suggested future work are presented in Section 6. Finally, the objectives, methodologies, findings, and implications of the study are summarized in Section 7.
System hardware and methodology
System hardware overview
The measurement of underwater structures presents challenges due to the turbidity of the underwater transmission medium, low light intensity, and resulting low brightness and poor clarity in image acquisition (Chen et al., 2023). Consequently, the physical design of the planar array cameras must meet these demanding measurement requirements. Following the pinhole imaging principle, as the object distance of the camera decreases, the field-of-view also diminishes. The conventional approach to addressing this is to employ a camera with a large field-of-view for image acquisition. In contrast, the typical method involves a single camera that rotates and moves to capture images of the structural surface integrity. However, this approach extends the shooting duration and escalates the cost of image acquisition. This paper proposes a solution utilizing planar array cameras to collect images from the surface of the underwater structure. The underwater portion comprises an array imaging system, illustrated in Figure 1, employing a skeleton design akin to a butterfly truss. This design minimizes the overall device weight and mitigates the impact of water currents. The aquatic component includes a power supply system, an umbilical cable and recovery unit, and a laptop computer with specialized software. This comprehensive setup ensures efficient image acquisition while addressing visibility and structural integrity issues in challenging underwater conditions. (a) Model diagram of the proposed planar array cameras and (b) physical drawing of planar array cameras.
The underwater array imaging system consists of six imaging units. Each imaging unit is equipped with a channel state information (CSI) interface, a complementary metal oxide semiconductor (CMOS), an image sensor, and a set of LED array light sources; the resolution of a single camera is greater than or equal to five megapixels, the focal length is 3.9 mm, the field of view is 91° × 75° × 60°, the lens aberration is <0.38%, the focus mode is manual, the aperture is fixed, the sensor chip is imx335. The light sources are arranged in a diamond-shaped row and feature adjustable brightness to ensure optimal lighting conditions. This design aims to reduce motion blur, enhance image quality, and shorten exposure times. All imaging units are securely affixed to a bracket, forming a stable imaging system. Aquatic and underwater systems are connected via a floating umbilical cable with built-in Power over Ethernet (POE) Gigabit Ethernet for rapid communication. This integrated setup facilitates efficient data transmission and communication between the underwater array imaging system and the aquatic components.
Methodology
As illustrated in Figure 2, the underwater inspection methodology based on planar array cameras can be systematically divided into four key components. Firstly, the camera calibration involves using a chessboard to calibrate camera parameters, which are then saved for future use. Secondly, the camera acquisition phase utilizes camera acquisition software to capture underwater images of the concrete surface. The acquired data is transmitted and grouped based on the sequence number of the shots, with each group named according to the corresponding camera number. Following calibration and acquisition, distortion parameters obtained are applied to correct distortions in the original underwater image data, resulting in processed images. The final step involves stitching these processed images to generate a comprehensive image representing the underwater detection data. This multi-step process ensures accurate calibration, acquisition, and correction of underwater images, ultimately providing a clear and complete representation of the underwater structure feature. Methodology for underwater inspection with planar array cameras.
This paper gives priority to the image stitching algorithm based on Harris corner point extraction for underwater structure inspection, owing to its computational efficiency and lower sensitivity to parameters compared to other feature extraction algorithms. This choice is motivated by the high pixel value of a single image, reaching up to 2592 × 1944, and the requirement for timely underwater structure inspection. Consequently, the algorithm can swiftly extract and stitch image features together, even when dealing with high-resolution images or smooth concrete surfaces.
Planar array cameras performance
This paper focuses on detecting underwater structural surfaces, aiming to achieve high detection accuracy, equipment stability, and accurate analysis of results from the underwater structural surface feature detection system based on array imaging. The system is designed to operate in water environments with a maximum current velocity of 2-3 knots, a depth of 20 m or more, and structures with a width or diameter of 2 m or less. After testing and analysis, the system meets the requirements for detecting structural surface features underwater.
Theoretical and actual field-of-view
Theoretical field-of-view
As depicted in Figure 3(a), the schematic illustrates the apparent structure distance of the camera array at 160 mm. The theoretical calculations indicate that with an object distance of 160 mm, the field-of-view of the camera module is 245.5 mm × 185 mm, as shown in Figure 3(b). The overlapping area between two images must exceed 30% to facilitate image stitching, so the object distance is adjusted to meet the field-of-view requirements and subsequent image stitching. Calculate the number of camera modules required and compare the actual measured values with the theoretical values to confirm that the measurement requirements are met, thereby achieving large-scale measurements of underwater structure surfaces. (a) Camera working distance of 160 mm and (b) field-of-view at working distance of 160 mm for all lenses.
Actual field-of-view
In utilizing planar array cameras to capture underwater images, it is crucial to determine the measured field-of-view of the cameras and adjust the overlap area between images to meet the requirements of image stitching. Therefore, before capturing images of underwater structures, it is necessary to confirm the relationship between the image capture distance of planar array cameras and the field-of-view statistics. The design working distance is 160 mm, with a test working distance of approximately 160 ± 5 mm.
Field-of-view Range Statistics.
Sufficient and uniform light
Due to the low visibility of the underwater environment, it is necessary to increase light intensity by adding light sources around the planar array cameras. Considering the demand for light in the underwater environment, the LED array light source is selected to ensure proper illumination and improve imaging quality. In this paper, the planar array cameras utilize a total of six camera modules. The LED array light source is arranged in a diamond-shaped row with adjustable brightness to ensure that each camera module receives the maximum light area while ensuring uniformity. Each camera module requires 4 LED patches arranged in a skeleton hollow design to form a sufficient brightness and uniform illumination system. The design layout for the lighting system around the six camera modules ensures sufficient additional light fields to guarantee appropriate illumination, shorten the exposure time, reduce motion blur, and improve imaging quality.
Uniformity Statistics of the Grayscale Distribution of additional Light Fields.
Image stitching and evaluation methods
The planar array cameras capture six sets of image data, which are then stitched together using a feature extraction-based stitching algorithm (Joshi et al., 2020). This process involves several steps, including image preprocessing, the extraction of feature points (Harris and Stephens, 1988), image alignment, and fusion to create a single stitched image. The stitching process is outlined in Figure 4. Planar array cameras image stitching process.
Image stitching
Image pre-processing
In image acquisition, various factors, such as environmental conditions and sensor characteristics, may contribute to image degradation, blurring, and uneven brightness. To mitigate these issues, the initial step involves converting the original images to grayscale to reduce computational complexity and noise. Following this, median filtering and gamma correction techniques enhance image quality, particularly in scenes involving beam bottom image acquisition.
Harris corner point based feature extraction algorithm
Based on Harris corner points, the feature extraction algorithm begins by ensuring that the images to be stitched exhibit more than 30% overlap. Subsequently, the two photos are converted to grayscale images and applied Gaussian blur with σ = 1. An image pyramid is then constructed to identify Harris key points at different scales. The gradient of image brightness in the x and y directions is computed using the Sobel operator, followed by smoothing the gradient with a Gaussian function having σ = 1.5 to mitigate the impact of noise on brightness. The change in luminance can be determined using equation (1). This calculation can be approximated further by equation (2).
Adaptive non-maximal suppression is employed to select a specific number of critical points. Initially, a radius r with an initial value of infinity is set. When r decreases, critical points within the radius r where the R values of all other key points are smaller than the R-value of the central point are retained and added to the queue. The search terminates when the queue’s critical points reach a preset value. A total of 500 key points are extracted from each image. Initially, the critical point Rmax with the most considerable R-value in the entire image is found and added to the queue, and Rmax×0.9 is obtained. Iterate through all the key points; if the critical point Xi has Ri> Rmax×0.9, the radius of the point is set to infinity. If the critical point Xi has Ri< Rmax×0.9, compute the distance ri to the nearest point Xi that has Rj>0.9 R (equation (3)). Finally, sort all ri and select the 500 points with the most extensive r.
A moderate Gaussian blur is applied to the image, followed by extracting a region of 40x40 pixels centered on the critical point. The region is then downsampled to a size of 8 × 8, resulting in a 64-dimensional vector. Normalization is performed on this vector. Consequently, a 64-dimensional vector represents each key point, and a 500 × 64 feature matrix is obtained for each image separately.
Feature point matching
Feature point matching employs the RANSAC (Random Sample Consensus) algorithm. In the first image, a certain number of feature points are extracted each time, and corresponding feature points are identified in the second image through the feature descriptor. The relative monoclinic matrix between the second and first images is then computed, and the number of point pairs is statistically analyzed. This step is repeated N times to obtain the perspective transformation matrix with the highest number of consistent point pairs, which serves as the final transformation matrix. Additionally, a local matching criterion is established by utilizing the grayscale information near the feature points and employing the feature point descriptor. The feature points detected in the two images are divided into one-to-one matching pairs.
After identifying their feature points for the two images undergoing stitching, a correlation window of size (2N + 1) × (2N + 1) is selected in both images, with each feature point as the center. Subsequently, each feature point in the reference image is used as a reference point to locate the corresponding feature point in the image to be stitched. Image feature alignment is accomplished by computing the correlation coefficient K between the correlation windows of feature points, as depicted in equation (4).
Here, I and I′ represent the grayscale values of the two images, while
Image fusion
Utilizing the optimal perspective transformation matrix, the overlapping regions are aligned, and linear fusion is employed to merge these areas seamlessly. Linear fusion primarily involves direct averaging and weighted averaging. Direct averaging fusion sums average the grayscale values of corresponding pixels in the overlapped area. The result is then calculated as the grayscale value of the fused pixels, as demonstrated in the following equation (5).
Evaluation methods
Upon completing the image stitching process with the plane array cameras, the stitched image must undergo evaluation based on corresponding metrics (Solh and AlRegib, 2012). This evaluation primarily focuses on the precision of image alignment, the fusion situation in overlapping parts of images, and the overall visual effect the human eye perceives. These three evaluation metrics are utilized for quantitative assessment of image stitching quality, enhancing the intuitive understanding and accuracy of underwater inspection.
Human eye visual effects metrics
As there is no standard reference image, the human eye’s visual effect entails a global analysis of the stitched image. Let G (i, j) represent the gray value of the stitched image, where the image size is M × N. The objective assessment of the stitched image includes the following three kinds of global metrics: (1) Information Entropy (E)
The information entropy of an image reflects the richness of its information and its ability to convey details. For a grayscale range of {0, 1, …}, the information entropy of the image is defined by equation (6). (2) Average Gradient (
The average gradient quantifies the speed of contrast changes in small details within an image. It is computed using equations (7), and a higher average gradient implies that more pronounced seam artifacts may be generated after image stitching. (3) Spatial frequency (SF)
Spatial frequency describes the overall activity of an image space and comprises spatial row frequency (RF) and spatial column frequency (CF), as formulated in Equation (8) and Equation (9). The overall spatial frequency value is taken from the root mean square of RF and CF with equation (10).
Image quality: edge intensity
An image’s edges' clarity indicates its overall detail and sharpness. Edge intensity is a measure of image clarity: the higher the value, the more precise the image, whereas a lower value suggests blurriness. Initially, the edges of the image are extracted using the Sobel operator, and subsequently, the edge strength is calculated using equation (11).
Mutual information
Let the overlapping regions of two neighboring images after alignment be A (i, j) and B (i, j), respectively, and the corresponding overlapping region of the stitched image be G (i, j). The amount of mutual information, a crucial concept in information theory, is utilized to measure the correlation between two variables. Applied here, it quantifies the mutual information between the stitched image and the original image. A more significant amount of mutual information signifies richer information obtained by the stitched image from the original, resulting in better fusion and stitching effects. The interaction information between G and A is expressed by equation (12), and the interaction information between G and B is expressed by equation (13).
Image field-of-view loss ratio
Measurements are taken to assess the magnitude of field-of-view loss due to image stitching at both macro and pixel levels. At the macro level, a scale is drawn to the concrete slab’s surface, providing each camera with a known scale length in the field-of-view at a specified object distance. After collecting data and completing stitching, the lost scale is used to calculate the loss of field-of-view, as demonstrated in equation (15).
Here,
At the pixel level, a scaled image with a specific field-of-view is captured using multiple cameras, and individual pixel lengths are calculated based on known scale measurements. The pixel difference between the image before and after stitching is then utilized to calculate the rate of field-of-view loss, as described in Equation (16) and Equation (17).
Here,
Experiments and results
After determining the focusing distance of plane array cameras as 160 mm, comprehensive experiments were conducted to assess its performance, encompassing aquatic and underwater tests. The collected data were stitched, compared, and rigorously analyzed to validate the system’s accuracy and robustness in detecting underwater concrete surface conditions.
Experimental setup
The experimental setup comprised two scenarios: an aquatic and an underwater environment. In the aquatic setup, a 100 × 100 concrete sewer plate (Figure 5) was utilized along with an underwater array camera system, a cable, and a host computer. The underwater scenario further included a chessboard and a test tank, as depicted in Figure 6. The array camera system boasts excellent waterproofing and seamless integration with the host computer for efficient image transmission and control. The underwater cameras were calibrated using a chessboard to ensure accurate image stitching. Concrete collection process and collection results. Underwater test scenarios: (a) concrete slab, (b) array cameras, and (c) water tank.

Experimental process
The experimental process commenced with a meticulous aquatic test, wherein the array camera system was mounted onto a support device, and the focal length was precisely calibrated at 160 mm, utilizing a reliable straightedge for accuracy. To maintain consistency in the shooting distance, the support device traversed the surface of the sewer concrete slab, systematically capturing images. A total of 10 groups were recorded, each comprising six array camera photographs capturing various sections of the concrete slab. The overlapping areas between each image were carefully planned to facilitate feature point identification and seamless stitching. The host computer directed the array cameras to capture these images under optimal lighting conditions, ensuring the clarity and detail of each shot. The captured data was systematically organized into 10 distinct groups, with each image labeled for ease of reference during the subsequent image stitching process.
Subsequently, the experiment progressed to underwater testing, where the array camera system was mounted onto the support device and submerged in a water tank. Real-time monitoring of the underwater images was conducted to determine the most effective shooting distance, ensuring that the captured images were both clear and brightly illuminated. Prior to capturing the test images, a rigorous calibration process was undertaken using a checkerboard grid placed on the concrete surface. This calibration ensured the accurate alignment and distortion correction of the six underwater cameras. The calibration images were then carefully saved for reference (Figure 7). Once the calibration was complete, the support device was maneuvered to capture images of the concrete plate’s surface. Similar to the aquatic test, 10 groups of six array camera photos were taken, covering different sections of the concrete slab. The overlapping areas between each image were identified to facilitate the stitching of feature points. The plane array cameras were instructed to capture these images under sufficient lighting conditions, ensuring that each shot captured the necessary details. The captured data was then saved and labeled, ready for the subsequent image stitching process. To ensure the reliability of the results, the entire process was repeated with another concrete slab, yielding an additional 10 groups of valuable test data. Calibration of internal and external parameters of the cameras.
Post-acquisition, the underwater array camera images underwent calibration utilizing the precise camera calibration parameters. Figure 8 presents a test image of the concrete slab captured by the plane array cameras, showcasing the successful removal of underwater aberrations, laying the foundation for successful image stitching outcomes. Aberration correction using camera parameters.
Experimental results
The aquatic test scene images were directly stitched according to the stitching process, resulting in six groups of images being stitched into a single stitched image of the concrete slab surface. Conversely, the underwater test scene involved camera calibration to obtain calibration parameters. These parameters were then used in a distortion correction algorithm to correct image distortion. Subsequently, the corrected images were stitched using the stitching algorithm to produce a stitched image of the concrete slab surface. Figure 9 shows the results after image processing. Figure 10 illustrates a comparison of the stitched images from the aquatic and underwater tests, demonstrating the effectiveness of the method in restoring the underwater concrete surface features. Results related to image processing. Detailed view comparing the results of the aquatic and underwater scenes (red: aquatic, and blue: underwater).

Evaluation result
A sample of 120 images was obtained by photographing the underwater concrete slab surface, and the images were stitched using the Harris algorithm and the Speeded-Up Robust Features (SURF) algorithm (Bay et al., 2008). Since there is no standard reference image, the information entropy, average gradient, spatial frequency, and edge strength of the images stitched based on the Harris algorithm and the SURF algorithm, respectively, were evaluated, fitted, and analyzed. Mutual information metrics were also evaluated, and comparisons were made between the original images and the stitched images with the fitted data analyzed (Figure 11). The information entropy of the image stitched by the Harris algorithm is slightly lower than that stitched by the SURF algorithm, with values around 6.6 and 6.7, respectively. Both stitching algorithms exhibit richness in image information and detail conveyance, with the SURF algorithm performing marginally better. Furthermore, the average gradient of the Harris algorithm is also slightly lower, approximately 1.7, compared to the SURF algorithm’s value of about 3.5. This suggests that the grayscale transition near the image boundary or shadow line is smoother with the Harris algorithm, indicating uniform grayscale changes in the stitched image. Moreover, the spatial frequency of the image stitched by the Harris algorithm is notably lower, around 3.5, compared to the SURF algorithm’s value of approximately 10.1, implying minimal grayscale changes in the image pixels, resulting in more natural stitching and fusion. Additionally, the edge intensity of the Harris algorithm is lower, about 17, in contrast to the SURF algorithm’s value of approximately 30, suggesting smaller gradients at edge points and thus a more effective image stitching effect. However, the mutual information of the Harris algorithm is higher, around 1.2, compared to the SURF algorithm’s value of about 0.9, indicating better image fusion capability and richer information extraction from the original image. Finally, Table 3 presents a comparison of the results of the five different evaluation metrics for the two stitching methods. The proposed Harris-based stitching algorithm exhibits superior performance compared to the SURF-based stitching algorithm, as shown in the table. Results of image stitching quality metrics: (a) entropy, (b) average gradient, (c) spatial frequency, (d) edge intensity, and (e) mutual information. A Comparison of the results of the evaluation Metrics for the two Stitching Methods.
Calculation of field-of-view loss
Macroscopic level
A total of 40 sets of data were compiled from two concrete slabs, one with scales underwater and the other without, and were stitched together. As depicted in Figure 12, at the macroscopic level, the loss of concrete area was estimated by marking scales of 5 cm for the larger scale and 2.5 cm for the smaller scale on the concrete slab surfaces. The field-of-view loss due to stitching was then determined by measuring the scale loss after the stitching process was completed during data collection. The calculated field-of-view loss for each difference in scale between two adjacent image data sets was approximately one large scale value of 5 cm. Similarly, the scale loss for the second longitudinal stitch was calculated to be a small scale value of 2.5 cm. Consequently, the total field-of-view loss for the two longitudinal stitches amounted to a large scale value of 5 cm. Concrete surface scale loss analysis: (a) Concrete Surface Scale and (b) a set of image stitching analysis results.
Pixel level
As for the pixel level depicted in Figure 13, underwater scale images were captured both vertically and horizontally using six cameras. The scale factor between pixels and physical length was computed based on the scale present in the images, and the field-of-view loss was calculated using the pixel difference and scale factor before and after stitching. The captured images had dimensions of 2592 × 1944 pixels. The mean physical length of the horizontal scale was 171 mm, while that of the vertical scale was 218 mm. Consequently, calculated by equation (16), the horizontal scale factor was approximately 0.0880 mm, and the vertical scale factor was approximately 0.0841 mm. Planar array cameras capture underwater scale images horizontally and vertically.
Multiplying the horizontal and vertical scale factors with the stitched image pixel value of 4078 × 3796 yields a total field of view of the stitched image measuring 358.86 mm × 319.24 mm. The field-of-view loss rate of the stitched image was then calculated and statistically analyzed using equation (17), with the results presented in the statistical analysis in Figure 14. The statistical analysis indicated that the field-of-view loss rate could be less than 0.1. Statistical analysis chart of field-of-view loss rate.
Discussion
Detecting underwater concrete surfaces at depths of 15 to 20 m faces challenges like water flow stability and transparency. The current image stitching method is reliant on feature extraction, lacks efficiency, and hinders inspection. This discussion examines the practical use of plane array cameras and the proposed algorithm in engineering applications. (1) Plane array cameras address the challenges of unstable water currents by employing a butterfly truss structure for better stability and resistance. This design helps mitigate the impact of water currents during underwater structure inspection. (2) Despite reduced water transparency, array cameras operate effectively at close distances with ample lighting, ensuring the consistent capture of clear inspection images. Future plans include enhancing light intensity and implementing algorithms to reduce turbidity, aiming to improve image quality and detection accuracy further. (3) The proposed method, utilizing Harris corner point extraction, offers a notable efficiency advantage over other algorithms for image stitching tasks with pixel sizes up to 2592 × 1944. While this efficiency is sufficient for practical inspection, further enhancements are needed. Moreover, achieving precise feature point matching remains challenging, especially on extremely smooth concrete surfaces. (4) Future research on array camera image stitching will focus on algorithm enhancements and exploring stitching methods independent of feature point matching, aiming to significantly improve stitching efficiency.
Conclusion
This paper presents a method for rapidly stitching and recognizing underwater concrete surface images using planar array cameras, exploiting their short distance and wide field of view. It introduces five quantitative evaluation metrics for image quality and outlines the calculation method for the field-of-view loss rate. Additionally, it provides a brief description of the performance of array cameras. The experiments consist of two phases. In the first phase, aquatic and underwater data were collected. The Harris algorithm was utilized for image stitching to compare and evaluate both sets of images qualitatively. Subsequently, underwater images underwent quantitative evaluation and analysis using the five evaluation metrics. The second phase focused on collecting data from underwater concrete surfaces, emphasizing images containing scales. The loss of field-of-view after stitching was then calculated and analyzed at both macro and pixel levels. The conclusions drawn from these two phases of analysis are as follows: (1) The detection data obtained by the underwater planar array cameras, after aberration processing correction, closely resembles the data obtained aquatically, effectively restoring the structure and condition of the concrete surface. (2) The individual camera field-of-view of the planar array cameras measures 218 mm × 171 mm, with a total field-of-view reaching 358.86 mm × 319.24 mm. This adequately meets the requirements for both close-range and wide-field imaging. Moreover, the overlap rate of the images can reach 38%, fulfilling the requirements for image stitching. (3) At a working distance of 160 mm, the stitching algorithm of the planar array cameras demonstrates stability, producing high-quality images. The underwater image stitching yields superior results, with the Harris-based stitching algorithm achieving an entropy of approximately 6.7, an average gradient of about 1.7, a spatial frequency of around 3.5, an edge strength of about 17, and a mutual information of approximately 1.2. These results indicate good image fusion and high stitching quality. (4) Multiple sets of concrete surface images with scale were stitched and analyzed for field-of-view loss. The difference in the field-of-view loss scale for each pair of horizontally stitched image data was approximately one large-scale value or about 5 cm. Similarly, the total field-of-view loss for the two vertically stitched sets equated to one large-scale value of approximately 5 cm. (5) The planar array camera system boasts a high resolution, with individual image pixels measuring 2592 × 1944. The physical size represented by a single pixel in the image is approximately 0.0841 mm × 0.0880 mm. Utilizing pixel differences to analyze and statistically assess the field-of-view loss rate indicates a rate of less than 0.1.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Natural Science Foundation of Jiangsu Province (BK20220849), the National Natural Science Foundation of China (No. 52208306, No. 52127813), the Jiangsu Provincial Key R&D Program (Social Development) (BE2022820), the Start-up Research Fund of Southeast University (RF1028623296), the Foundation of the Science and Technology on Near-Surface Detection Laboratory (6142414210804).
Data Availability Statement
The data used to support the findings of this study shall be provided upon receiving reasonable request.
