Abstract
Objective:
Sandia National Laboratories conducted an experiment for the National Nuclear Security Administration to determine the reliability of visual inspection of precision manufactured parts used in nuclear weapons.
Background:
Visual inspection has been extensively researched since the early 20th century; however, the reliability of visual inspection for nuclear weapons parts has not been addressed. In addition, the efficacy of using inspector confidence ratings to guide multiple inspections in an effort to improve overall performance accuracy is unknown. Further, the workload associated with inspection has not been documented, and newer measures of stress have not been applied.
Method:
Eighty-two inspectors in the U.S. Nuclear Security Enterprise inspected 140 parts for eight different defects.
Results:
Inspectors correctly rejected 85% of defective items and incorrectly rejected 35% of acceptable parts. Use of a phased inspection approach based on inspector confidence ratings was not an effective or efficient technique to improve the overall accuracy of the process. Results did verify that inspection is a workload-intensive task, dominated by mental demand and effort.
Conclusion:
Hits for Nuclear Security Enterprise inspection were not vastly superior to the industry average of 80%, and they were achieved at the expense of a high scrap rate not typically observed during visual inspection tasks.
Application:
This study provides the first empirical data to address the reliability of visual inspection for precision manufactured parts used in nuclear weapons. Results enhance current understanding of the process of visual inspection and can be applied to improve reliability for precision manufactured parts.
Keywords
Introduction
Visual inspection has been extensively researched since the early 20th century in order to understand the factors that impact performance accuracy (See, 2012). The most frequent and consistent observation in this vast body of research is the imperfection of human inspectors. For most inspection tasks, missed defects typically range from 20% to 30% (Drury & Fox, 1975). Despite this rich history of research involving visual inspection, three factors, discussed next, have not been thoroughly investigated.
Visual Inspection of Precision Manufactured Parts for Nuclear Weapons
Visual inspection has encompassed a wide variety of products, including piston rings, acoustical tiles, aircraft components, pipelines, highway bridges, television panels, contact lenses, printed circuit board assemblies, airport baggage, pharmaceuticals, and food products. Research has demonstrated that the specific type of product under inspection does not alter the level of 70% to 80% accuracy typically observed (See, 2012).
To date, however, accuracy of visual inspection in the nuclear weapons arena has never been empirically established. Visual inspection is widely used within the Nuclear Security Enterprise, the entity responsible for maintaining the safety, security, and reliability of the U.S. nuclear weapons stockpile and all associated components. Specifically, visual inspection is one of several techniques constituting a comprehensive quality control and assurance program designed to provide defense in depth for nuclear weapon components. Visual inspection may be used to supplement objective tests of functionality, or it may form the last line of defense to assure an item is defect-free before installation in a nuclear weapon.
On the surface, logic dictates that inspection of nuclear weapon parts should not differ from other products that have been investigated. On another level, however, inspection of nuclear weapon parts differs from inspection for many other types of products due to the consequences of missed defects. For nuclear weapons, missed defects may result in injury, fatalities, property damage, or release of radioactive material. For most ordinary consumer products, the consequences are much less severe—missed defects may lead to customer dissatisfaction or a loss in repeat business. In the absence of empirical data, the common assumption within the Nuclear Security Enterprise is that accuracy must be better than the industry standard, given the high-consequence nature of the parts involved. Thus, an experimental investigation of inspection accuracy for nuclear weapon parts is warranted to address the validity of this assumption.
Multiple Inspections and Inspector Confidence Ratings
Procedural modifications represent a proven technique to improve inspection accuracy (See, 2012). One procedural approach involves using a team of inspectors wherein multiple personnel reinspect a product and provide individual accept/reject decisions (Drury, Karwan, & Vanderwarker, 1986; Harris, 1969; See, 2012). With two or more inspectors, repeated inspections can significantly increase accuracy for critical defects; however, the magnitude of the benefits decreases as the number of inspections increases—very little increase in accuracy occurs after six independent inspections (Harris, 1969). In multiple-inspector situations, the optimal approach to maximize accuracy efficiently consists of having two inspectors inspect every item, and both must reject a product for it to be classified as defective (Drury et al., 1986).
Comparable results regarding a team approach and the necessity for independence were obtained in a similar study conducted in the field of vigilance, wherein observers must maintain their attention and focus on a display for a prolonged period in order to detect the occurrence of critical signals or targets (Wiener, 1964). In Wiener’s (1964) study, a response was scored as correct if at least one person on the team responded to a critical signal for detection during the monitoring task. Although the two-person team detected more signals than a single individual working in isolation, performance fell short of the level predicted by a probability model for independent events. This outcome suggests that independence was not achieved when observers worked as a team in the same location.
While research has shown that multiple inspections may improve accuracy, the efficacy of using inspector confidence ratings to guide multiple inspections has never been investigated. The only previous study incorporating confidence ratings was conducted for a different purpose, and that methodology was not successful. Namely, Ainsworth (1982) asked inspectors to provide confidence ratings for each accept/reject decision in order to apply a signal detection theory (SDT) analysis but had to abandon this approach because inspectors were not inclined to assign low confidence ratings to their decisions. Incorporating inspector confidence ratings as part of a reinspection process might prove useful to improve both overall performance and process efficiency. Specifically, even though the team approach has advantages for performance improvement, it degrades process efficiency since every part must be inspected twice. One approach to reduce the total number of parts requiring reinspection, thereby enhancing overall process efficiency, might be to use inspector confidence ratings to eliminate parts associated with high confidence from further inspection. This possibility has not been explored in the field of inspection.
Workload and Stress of Visual Inspection
The inspection literature is replete with allusions to the demanding and stressful nature of visual inspection tasks (Drury, 1985; Drury, 1992; Gallwey, 1982). Gallwey (1982), for example, indicates that inspection requires considerable mental processing, concentration, and information transmission, along with extensive use of both short-term and long-term memory. Further, inspection must typically be completed quickly so as not to delay production, and critical defects that may require rework must be identified early. In addition, multiple defects at various severity levels and locations may be present, all of which can create a stressful task for the inspector. As Drury (1992) pointed out, the stress associated with inspection is only magnified when inspectors are exhorted to “try harder” to be perfect. In his 1985 review, however, Drury noted that very few studies have actually been specifically designed to examine how the job itself affects inspector stress.
Although it appears to be tacit knowledge that inspection is demanding and stressful, empirical data are either lacking or predominantly confined to early inspection studies. Standardized subjective, performance-based, and physiological techniques to assess workload have never been reported in studies of visual inspection, and stress has not been addressed since Drury’s (1985) review. With respect to workload, the National Aeronautics and Space Administration (NASA)–Task Load Index (TLX) was used in at least two different studies (Drury, Ghylin, & Holness, 2006; Drury, Green, Chen, & Henry, 2006). However, for various reasons, participant workload ratings per se were not reported. Thus, empirical data to estimate the workload associated with inspection are not currently available, and newer measures of stress, such as the Short Stress State Questionnaire (SSSQ), have not been applied at all.
Objectives of the Present Study
The present study was designed to fill these three gaps in the research literature. The first goal was to empirically establish the accuracy of visual inspection for precision manufactured parts used in nuclear weapons by closely simulating an operational Nuclear Security Enterprise inspection process. It was hypothesized that performance accuracy would not differ significantly from the 70%-to-80% estimates commonly quoted in the literature. The second goal was to investigate the efficacy of using inspector confidence ratings to guide reinspection in an effort to improve overall performance accuracy while limiting the number of items for reinspection to those items rated in the lower ranges of the confidence rating scale. It was hypothesized that this approach would improve performance accuracy over that of a single inspector alone. The third goal was to empirically establish the workload associated with inspection and to measure stress using the recently developed SSSQ technique for subjective estimates of stress. It was hypothesized that subjective ratings of workload and stress would fall into the high ranges of the respective scales.
Methodology
Participants
Participants were 82 qualified inspectors (70 males) who currently compose the inspection workforce at six of the eight sites in the Nuclear Security Enterprise (Table 1). Participants had performed visual inspection on the job for 6 months to 35 years (M = 10, SD = 9.5).
Inspector Demographics
Note. ASQ = American Society for Quality.
Remaining inspectors reported either better or worse than 20/20 vision or did not know their visual acuity. No inspector had been disqualified from on-the-job inspection due to poor vision. Further, even severe myopia would not have been problematic in the current task due to the short viewing distances involved.
Test Articles
Test articles for inspection were precision parts manufactured for Sandia National Laboratories (SNL) and informally referred to as flags. The parts consisted of a thin, flexible metallic material (the banner of the flag) welded onto a rigid stake (the flagpole). The banner of the flag was less than 3 inches long, making the part easy to pick up and inspect by hand. Because the manufacturing process is automated, numerous defects can occur until errors in the automated setup are identified and corrected. In addition, all parts, regardless of quality, are shipped to SNL for inspection and disposition. As a result, the level of defects for a given batch of parts can reach as high as 60%, with an overall average across batches of approximately 30% (this defect rate should not be interpreted as indicative of the level for other parts used in nuclear weapons). Consequently, in actual operations at SNL, the parts undergo 100% receipt inspection to eliminate obviously defective parts, while accepting as many viable items as possible since good items may become damaged during subsequent processing. Defective items may be detected later in the process via objective tests of component functionality; however, the ideal time to detect defects occurs during initial receipt inspection. This process was simulated in the current study. For the purposes of the study, eight defects (see Results) were selected from the population of potential defects critical for part acceptance.
Two individuals at SNL who routinely inspect these parts provided the ground truth. They determined whether each item was defective or nondefective and identified all defects present in any rejectable part. The SNL inspectors reviewed parts multiple times until agreement was achieved. A total of 200 items was catalogued for possible inclusion in the study.
From this initial set of items, 100 test articles were selected for Part 1 of the study. The intent was to establish a sample large enough to derive reliable estimates of inspection accuracy but small enough to confine each experimental session to a single day in order to minimize the impact to productivity at each operational site. The number of defective items in Part 1 was 30, a level chosen specifically to represent the typical operational defect rate for this part in order to maximize generalizability of results to Nuclear Security Enterprise inspection. The total number of defects in each rejectable item was either one (54%), two (40%), three (3%), or four (3%).
Part 2 included 40 items, 20 of which were defective (with either one [85%] or two [15%] defects per item). Items for Part 2 were based on the responses of a single qualified SNL inspector who had no previous experience inspecting the parts. This individual evaluated the remaining 100 parts from the initial set of 200 items and provided confidence ratings for each accept/reject decision. These data were used to provide the sample of test articles for Part 2 that had previously been rated with only low or medium confidence.
Procedure
Data collection at each site was set up to closely simulate the current SNL inspection process. As such, inspections occurred in a standard conference room at each site (average lighting of 750 lux), with no magnification aids or supplemental lighting. Participants completed the inspection task individually during an approximate 8-hr session, depending upon the pace at which the inspector worked. Each session included short breaks about every 30 min to avoid vigilance effects. This objective was achieved, as evidenced by similar percentages of correct responses for the first 20 trials in Part 1 (M = 73, SD = 12) and the last 20 trials (M = 72, SD = 16), t(81) = 0.333, p = .740, 95% confidence interval (CI) of the difference [–3, 4], d = 0.04.
All 82 inspectors completed approximately 2 hr of training (Figure 1) before completing both Parts 1 and 2 of the inspection task. For both Parts 1 and 2, items were presented individually in a different random order for each participant. To avoid damaging parts during the study, each item was contained in a separate 3-by-5-inch antistatic envelope, and only the experimenter handled the parts. To present a part to the participant, the experimenter used tweezers to transfer the item from the envelope to a holding fixture that held the flag securely in an upright position by the base of the flagpole. The participant was free to rotate the part as needed to complete the inspection but touched only the holding fixture and never the part itself. Once the participant was finished inspecting an item, the experimenter removed it from the holding fixture with tweezers, replaced it in the antistatic envelope, and prepared the next item.

Inspector training. Training included four different activities.
Inspectors used a work instruction adapted directly from the current SNL procedure to determine the presence or absence of the eight predefined defects. Inspectors decided whether the part was defective, rated their confidence in that decision (low, medium, or high), and then identified any and all defects. The work instruction advised inspectors to use the full range of confidence ratings and to assign a high rating only if absolutely certain of the decision. In addition, the work instruction identified the search sequence for the eight defects, based on the current SNL approach. As a result, search occurred defect by defect as opposed to section by section of the part itself, and the search was exhaustive rather than self-terminating. Deviations from the sequence in the work instruction were permitted if inspectors immediately detected rejectable defects. Throughout the inspection process, participants were instructed to think aloud, while the experimenter captured their real-time responses in an electronic database. In addition to the work instruction, the orientation briefing and one acceptable defect-free item were available throughout the task. Instructions and procedures were identical for Part 2, except participants were informed they would reinspect only those items a previous inspector had rated with low or medium confidence in order to provide a second opinion.
Participants were informed that both accuracy and time were critical in Parts 1 and 2, but no time limit was imposed. Each trial began and ended with “start” and “done” signals from the participant. Time stamps were recorded in the electronic database for each trial and used to compute the total time per item—the interval between the inspector’s “start” and “done” signals.
Inspectors completed the NASA-TLX and SSSQ rating scales in a random order immediately after Part 1 and again after Part 2. The rating scales were the same as those used during training, with the exception of one additional question on the NASA-TLX: Inspectors were asked to estimate how the typical workload in their normal job compares to the ratings provided for the current task.
Data Reduction and Analysis
Descriptions of the performance, workload, and stress measures computed for Parts 1 and 2 of the inspection task are provided in Table 2. The measures in Table 2 applicable to Part 1 were used in independent-samples t tests, median tests, and correlational analyses to identify impacts of personal and environmental variables on inspection performance, workload, and stress.
Performance, Workload, and Stress Measures
Note. SDT = signal detection theory; NASA-TLX = National Aeronautics and Space Administration–Task Load Index; SSSQ = Short Stress State Questionnaire.
A defective part that was rejected was scored as a hit even if the reasons for the rejection were incorrect in order to simulate operational processes, wherein decisions are considered correct if nondefective parts are accepted and defective parts are rejected.
A′ ranges from .50 (chance performance) to 1.00 (perfect performance) (Grier, 1971; Macmillan & Creelman, 1990).
B′′ D ranges from −1.00 (bias to reject items) to +1.00 (bias to accept items), with 0 representing a neutral bias (Donaldson, 1992).
Corresponds to Wiener’s (1964) approach.
Inspector 2, Category 1, and Category 2 accuracy data were compared with the accuracy data for Inspector 1 alone via independent-samples t tests to determine if reinspection improved overall performance.
Corresponds to Drury, Karwan, and Vanderwarker’s (1986) approach.
NASA-TLX provides an overall workload score based on a weighted average of participant ratings on six different subscales (Mental Demand, Physical Demand, Temporal Demand, Performance, Effort, and Frustration; Hart & Staveland, 1988). Weightings are achieved by presenting 15 pairwise comparisons and asking participants to choose which subscale was more important to task workload. Subscales deemed most important in creating task workload are given more weight in computing overall workload. Both subscale scores and overall workload scores range from 0 to 100, with higher scores representing greater workload.
Correlational analyses were used to determine if workload and stress were related to performance accuracy and response time.
SSSQ assesses three broad dimensions of psychological state: task engagement, distress, and worry (Helton, 2004; Matthews, Emo, & Funke, 2005). Participants provide a rating from 0 to 4 for each of 30 statements to indicate whether the statement was definitely false (0), somewhat false (1), neither true nor false (2), somewhat true (3), or definitely true (4) regarding how they felt during the task. Responses are later combined to derive scores for each of the three dimensions. Scores range from 0 to 32, with higher scores representing greater task engagement, distress, or worry.
Results
General Inspection Observations
The most widely used inspection technique, as revealed by inspector “think-aloud” comments, began with a general assessment of the overall item to detect any salient defects, followed by a serial assessment of each individual defect in accordance with the work instruction. Given that the task was self-paced, few comments were made regarding time pressure or stress. For some items, inspectors did frequently comment they would feel more certain if they had had more inspection experience with this item to better understand the differences between acceptable and rejectable items. Inspectors also commonly stated they would consult another inspector when they were uncertain (though they might still rate that decision with high confidence).
Examination of inspector responses by individual item revealed that correct decisions for each acceptable part (or correct accepts) ranged from 17% to 95%, whereas hits for defective items ranged from 37% to 100%. None of the acceptable items was associated with 100%-correct accepts, and only two parts were associated with correct accepts below 30% (i.e., most inspectors found some reason to incorrectly reject the part). Acceptable parts with low correct accepts tended to be rejected because inspectors thought they detected rejectable creases. Five defective parts had hits of 100%, whereas none was associated with hits below 30%. Four of the five parts with 100% hits had large defects that were difficult to miss.
Visual Inspection Accuracy and Response Time
Inspection accuracy
Inspectors detected an average of 85% (SD = 10%) of defective items in Part 1, 95% CI [82, 87], and incorrectly rejected 35% (SD = 15%) of acceptable parts, 95% CI [31, 38]. A one-sample t test indicated that hits for Nuclear Security Enterprise inspectors were significantly higher than the population value of 80% commonly reported in industry, t(81) = 4.29, p < .001, 95% CI of the difference [3, 7], d = .47. Notably, however, the CI does not include accuracy scores of 90% or higher. In addition, comparable detection percentages were 77% for valid hits and 35% for exact hits, indicating that correct rejections of defective items were frequently achieved for the wrong reasons. In fact, although individual inspectors typically identified only one defect per item, the defect type was most frequently spread across five different categories across inspectors. SDT measures of perceptual sensitivity and response bias derived from hits and false alarms indicated that the task was moderately difficult, with a mean A′ score of .84 (SD = .06). Inspectors were slightly biased toward rejecting items, with an average B″ D of –.43 (SD = .45).
A breakdown of frequencies for each confidence rating and response type during Part 1 indicated that 57% of inspector confidence ratings were high, whereas only 7% were low. Table 3 identifies additional trends associated with inspector confidence ratings.
Confidence Rating Results
Note. The majority of inspector decisions (71%) were correct (either hits or correct accepts), and 29% were incorrect (false alarms or misses).
Examination of accuracy by defect type revealed that hits varied widely, whereas false alarms were approximately 5% or less for each defect (with the exception of creases; Table 4). Inspector “think-aloud” comments provided evidence that fewer than 25% of inspectors who missed a defect contemplated its potential presence in an item. In addition, in multidefect items, one of the defects tended to be much more frequently detected. As just one example, 78 inspectors detected defective welds, but only two inspectors detected the misalignment, in one multidefect item. This type of observation pertained to 12 of 14 multidefect items.
Performance Accuracy by Defect Type (in percentages)
Response time
Inspectors were slightly faster when they made correct decisions (hits and correct accepts) versus incorrect decisions (false alarms and misses) (Table 5). A repeated-measures analysis of variance with a Greenhouse-Geisser correction for violation of sphericity indicated that mean response time did vary significantly by response type, F(1.93, 152.30) = 13.01, p < .001, partial η2 = .14. Post hoc tests using the Bonferroni adjustment for multiple comparisons revealed that the time for correct responses was significantly faster than the time for incorrect responses (p < .05). No other differences were significant.
Mean Inspection Time by Response Type
Note. All times are reported in minutes:seconds. Average time per part was just under one minute (M = 00:56, SD = 00:17).
Given that response time was measured for an entire item, and almost half of the items in Part 1 contained more than one defect, response times for each separate defect could not be determined. However, to probe whether the pattern in response times described previously was relevant regardless of defect type, seven items that contained only a crease and five items that had only fingerprints were analyzed separately. As with the overall analysis of variance for all response times, mean inspection time did vary significantly by response type for creases, F(1.91, 89.95) = 5.46, p = .006, partial η2 = .10, and for fingerprints, F(1.93, 115.71) = 5.67, p = .005, partial η2 = .09. For both creases and fingerprints, post hoc tests using the Bonferroni adjustment revealed that inspection times for hits and correct accepts were significantly faster than the time for false alarms (p < .05) but not misses.
Performance measure correlations
Pearson product-moment correlation coefficients indicated that all performance measures except A′ were significantly correlated with the average time spent inspecting each part (Table 6). Inspectors who spent longer examining each part had both higher hits and false alarms, which resulted in lower B″ D scores (a bias to reject parts).
Correlations Between Performance and Time
Note. CI = confidence interval.
p < .01.
Additional Spearman rank-order correlational analyses were conducted to determine whether inspector confidence ratings were related to the time spent inspecting each part or the accuracy of the response. Inspector confidence tended to decrease as the time spent inspecting a part increased, rs(8198) = –.31, p < .001, 95% CI [–.34, –.30]. Further, confidence increased in accordance with the accuracy of the response, rs(8198) = .20, p < .001, 95% CI [.18, .22].
Reinspection Based on Confidence Ratings
Performance accuracy for Part 2 (Table 7) was degraded, as compared to Part 1, because Part 2 contained only low- or medium-confidence items. Parts that were obviously acceptable or that had very salient defects were not represented because they had been rated with high confidence. As shown in Table 7, the second inspector had higher average hits and false alarms than the first inspector; however, the overall perceptual sensitivity to differentiate defective parts from nondefective parts was similar, as reflected in the A′ scores. Independent-samples t tests comparing Inspectors 1 and 2 revealed no significant differences for any of the performance measures (p > .05).
Part 2 Performance Accuracy Results
Inspector 1 was a single individual; therefore, no standard deviations are available.
At least one inspector has to reject an item for it be classified as defective.
Both inspectors must reject an item for it to be classified as defective.
Category 1 and Category 2 data in Table 7 represent the two different methods used to combine inspector decision responses. The Category 1 method resulted in both higher hits and false alarms as compared to Inspector 1, whereas the Category 2 method was associated with both lower hits and false alarms. In terms of the critical performance measure of A′, however, independent-samples t tests indicated that neither method improved performance over that of Inspector 1 alone (p > .05).
Workload and Stress
Workload
The average global weighted NASA-TLX score for Part 1 of the inspection task was 42 (SD = 15). The primary contributors to workload were mental demand and effort (Figure 2). When asked to compare the workload of their normal daily inspection job with the study task, 68% of inspectors indicated their normal job is “somewhat” or “much more” demanding. Their normal jobs require them to inspect more than one type of part, more complex parts, and more than eight different defects at a time. Most inspectors also perform additional tasks besides inspection, and their normal jobs involve time constraints to meet deadlines that were not present during the task used in this study.

Mean National Aeronautics and Space Administration–Task Load Index subscale scores. Error bars represent standard errors.
Pearson product-moment correlations between the average global weighted NASA-TLX score for Part 1 and the performance measures of hits, false alarms, A′, B″ D , and average time per part revealed a statistically significant correlation only for false alarms, r(80) = .23, p = .038, 95% CI [.01, .42]. Inspectors who incorrectly rejected a larger number of acceptable parts reported higher overall workload. No other correlations between workload and performance were significant (p > .05).
Stress
SSSQ scores indicate that inspectors were fairly highly engaged during task completion (Mdn = 24). Distress during the task was relatively low (Mdn = 8), though worry was slightly elevated by comparison (Mdn = 14). All three SSSQ subscales were significantly correlated with false alarms (Table 8). Inspectors who committed more false alarms tended to be less engaged in the task and to report greater distress and worry. The Distress subscale was significantly correlated with the response bias measure of B″ D , such that lower B″ D scores (a bias to reject parts) were associated with greater reported distress. Finally, the significant correlation between distress and the average time per part indicated that the more time inspectors spent examining parts, the more distressed they felt.
SSSQ and Performance Correlations
Note. SSSQ = Short Stress State Questionnaire; CI = confidence interval.
p < .05. **p < .01.
Personal and Environmental Factors
Most of the personal factors were not significantly related to performance, workload, or stress. These included experience (number of years performing inspection work on the job), visual acuity (comparison of inspectors with uncorrected and corrected 20/20 vision), education, and professional inspection certification. Some differences due to gender and age did emerge. First, perceptual sensitivity was higher for males (M = .85, SD = .05) as compared to females (M = .80, SD = .08), t(80) = 2.20, p = .048, 95% CI of the difference [0.001, 0.108], d = .99. This sensitivity difference stemmed primarily from elevated false alarms for females (M = 42%, SD = 20%) as compared to males (M = 33%, SD = 14%). Second, older inspectors had fewer hits, lower perceptual sensitivity, and a greater tendency to accept parts, as compared to younger inspectors (Table 9).
Inspector Age and Performance Correlations
Note. CI = confidence interval.
p < .05. **p < .01.
The only environmental factor that could be addressed was lighting, which affected only perceptual sensitivity, r(80) = .22, p = .047, 95% CI [.003, .417]. Up to 5% of the variability in inspector perceptual sensitivity could be explained by variations in lighting, with a tendency for performance to improve as lighting increased.
Discussion
Reliability of Visual Inspection for Precision Manufactured Parts
The present results indicate that important differences do exist when inspection involves precision manufactured parts for nuclear weapons. Namely, although Nuclear Security Enterprise inspectors were not vastly superior to industry counterparts in terms of missed defects (85% vs. 80%), they committed many false alarms. In terms of hits, it is questionable whether the observed difference is practically significant. First, average hits were not in the 90% range at all. Second, inspector performance became progressively worse as the criteria for categorizing a response as a hit became more stringent (valid and exact hits), indicating that the rejectable defect in a part was often missed.
More importantly, however, the relatively superior hits for nuclear weapons inspectors were achieved at the expense of a high false-alarm rate. Incorrectly rejecting 35% of acceptable items is unusual. In most inspection tasks, errors are omissions (missed defects) rather than commissions (false alarms); when false alarms do occur, they are typically 1% to 10% (Drury, 1985; Wiener, 1984).
The high false alarms in the present study therefore represent a critical departure from the general inspection literature. This departure can be attributed to two features of Nuclear Security Enterprise inspection. First, when search is conducted defect by defect, inspectors have many opportunities to correctly detect a defect and commit false alarms. In the present study, participants inspected a given part eight different times as they searched sequentially for possible defects. Second, there is a known and openly acknowledged bias in the Nuclear Security Enterprise due to the high-consequence nature of the items under inspection. Namely, “when in doubt, toss it out” is a popular refrain throughout the Nuclear Security Enterprise—inspectors are trained to reject parts with potential defects if they have any uncertainty.
The high false-reject rate is concerning because it represents considerable waste. Specifically, the product engineer must conduct additional evaluations for any part associated with uncertain inspection results, using inspector input and other available information (e.g., inclusion of subsequent functional testing) to decide whether to accept the part and return it to stock or scrap the part altogether. Regardless of the final decision, such engineering evaluations incur considerable time, money, and effort, all of which can increase as the number of false reject decisions increases.
The extremely cautious approach adopted by nuclear weapons inspectors may in fact differ from approaches in comparable high-consequence fields (aircraft, medical, and food industries) that also face relatively severe consequences. For example, if tainted meat is missed during final inspection, people who consume the meat may fall seriously ill or die. Similarly, if cracks are not detected during aircraft inspection, passengers may die and expensive equipment may be irreparably damaged. However, despite similar severe consequences, these fields have not reported false alarms in the range observed in the present study (See, 2012).
Finally, with respect to the overall level of reliability in the present study, the question of representativeness must be addressed since the observed reliability may represent an upper or lower performance bound (Table 10). Given these two opposing views, a logical approach is to use the observed reliability as a starting point for expected performance until additional relevant information, such as defect rates, is known.
Rationale for Upper and Lower Performance Bounds
Isolated test is not representative of demands, time pressures, and interruptions associated with normal working conditions, all of which can degrade performance (Sinclair, 1979).
Typical defect rate for inspection is <10%. Multiple studies have demonstrated that inspection accuracy improves as defect rate increases (See, 2012).
Process of searching an item serially by defect rather than by area is associated with higher performance accuracy (See, 2012).
Reliability may be similar to that of a new part after initial training, such that the inspector has minimal experience observing differences between acceptable and rejectable parts. This view assumes that more experience will necessarily lead to better performance, a supposition that is not wholly supported by the literature (McCornack, 1961). In fact, the present study did not identify any correlations between experience and performance.
Research has demonstrated that performance is degraded as the number of defects that must be included in search increases (See, 2012).
Understanding the Process of Visual Inspection
Several findings in the present research enhance current understanding of the process of visual inspection. First, previous research has demonstrated that a team approach to inspection can be used to improve overall performance. However, the team approach in the current study guided by inspector confidence ratings was neither effective nor efficient, at least not in the manner implemented. This outcome has practical significance for attempts to improve inspection performance, particularly since procedural and process modifications represent one of the most effective approaches to improve performance (See, 2012).
Implementing reinspection based on confidence ratings in the present study did verify the supposition that parts rated with higher confidence are associated with greater accuracy. This outcome implies that reinspection guided by inspector confidence ratings could potentially prove effective, provided that a better method for gauging inspector confidence is developed. In accordance with observations from previous studies, inspectors in the current study rarely used low confidence ratings, despite periodic reminders to use the full range of the rating scale. Inspector comments suggested that confidence ratings were strongly related to their qualifications as inspectors, such that a low or medium rating was equated with an inability to perform their assigned work.
More importantly, however, low and medium confidence ratings were not avoided altogether, and there was a tendency to assign such ratings disproportionately to acceptable items and incorrect decisions. Further, while thinking aloud during inspection, inspectors often expressed a desire for more experience with the item being judged before engaging in the inspection task. These results imply that inspectors are able to differentiate various levels of confidence in their decisions. Inspectors also voluntarily commented they would seek a second opinion for some questionable items. Such a process is widely used throughout the Nuclear Security Enterprise and suggests a better approach to measure inspector confidence (e.g., with a rating scale that gauges likelihood of seeking a second opinion). In essence, using the word confidence in the rating scale may be the issue, not the rating in and of itself.
A second contribution of the current results is the empirical data to support claims in the literature that inspection is a demanding task. The nature of the results suggests that workload may play a larger role in inspection than stress. There was less support for the notion that inspection is stressful, which may have been due to administering the test in isolation, separate from other normal work demands and stressors. Indeed, Drury (1985) distinguishes between the tasks of inspection (locating defects, deciding whether defects are rejectable, and documenting decisions) and the job of inspection (which encompasses the individual tasks embedded in the overall organizational and social environment at the work site). Stress ratings in the present study largely represent task ratings and might well be higher if inspectors had also been asked to indicate how their ratings compare with the stress of their normal daily inspection job. Such a comparison was relatively easy to make for the NASA-TLX, with only six workload subscales, but not for the SSSQ, with its 30 different items.
A third contribution of the present results stems from multiple analyses suggesting that the more problematic element of the inspection process involves decision making for acceptable items. For example, no acceptable item was associated with 100% correct decisions across all inspectors, but five defective parts had 100% hits. Further, the minimum percentage correct for acceptable parts (17%) was lower than the comparable minimum for defective parts (37%). As just described, acceptable items were also more likely to receive lower confidence ratings. In addition, inspectors who committed more false alarms on acceptable parts reported higher subjective workload, lower task engagement, greater distress, and increased worry. These outcomes suggest that inspectors may have realized they were committing errors on some level, especially for acceptable parts.
A fourth contribution is derived from the analysis of inspector misses. Namely, when misses occurred, the defect was rarely considered at all by the inspector. There were few instances wherein the inspector noticed the defect but actively decided that it fell within the criteria for acceptance. Similarly, there appeared to be a tendency for reduced attentiveness once an inspector had already detected a rejectable defect, leading to an increased likelihood of missing other defects in multidefect items. These outcomes indicate that inspectors may not have seen the defect at all during search, or they may have thought the defect did not meet the criteria for rejection and was therefore not worth mentioning. Given that search was exhaustive and serial, the latter possibility appears more probable, suggesting that training to improve decision making might be more useful for this task than training to improve search.
Finally, the finding that inspection times were slightly faster for correct as opposed to incorrect responses lends further insight into inspector behavior during the inspection process. Similar findings have been reported in previous detection studies and can be interpreted in terms of the two-component model of inspection, which combines visual search and decision theory to describe inspection in terms of separate search and nonsearch processes (decision making; Spitz & Drury, 1978). The two-component model is most directly applicable when each item contains only one defect and search is self-terminating rather than exhaustive. For these reasons, the present data did not permit a formal evaluation of the applicability of the two-component model.
Nevertheless, the pattern of results that was observed with respect to response type and response times corresponds well with the predictions from the two-component model of inspection, in particular, the decision theory component of the model. Under the decision theory component, response latency and distance from the criterion are assumed to be inversely related (Parasuraman & Davies, 1976; Pike, 1968; Sekuler, 1965). In essence, observations close to the criterion provide less evidence to determine if the observation represents a signal (defective part) and will therefore require longer response times. For observations farther from the criterion, the operator has stronger evidence to form a decision and can respond more quickly.
Overall, the nature of the response time and accuracy results implies that inspectors can discriminate at some level between reject decisions in which the part actually is rejectable and those in which no defect is present. Along these lines, inspectors who took longer on each part reported lower confidence and higher distress on the SSSQ. Such outcomes suggest there is an optimal time for inspection that maximizes hits but prevents inordinate false alarms. Finding this “sweet spot” is the challenge, particularly since current practice in the Nuclear Security Enterprise stresses accuracy over speed. Having virtually unlimited time for inspection may actually open the door for more errors.
Improving Visual Inspection
The present results can be addressed in terms of recommended guidance to improve inspection. Current findings demonstrate that despite considerable research on visual inspection, it continues to be difficult to design an optimal process, even for high-consequence products. See’s (2012) review indicated that potential solutions to improve inspection performance can be classified into three major categories: training, inspection procedures, and apparatus. In terms of the high false-reject rates observed in the present study, the best approach to circumvent additional costs due to this issue is to reduce the gray area of doubt that leads to false alarms by providing clear work instructions and tailored training. Such an approach can yield fewer false alarms (less scrap) as well as higher hits (better product going forward in the development process).
Work instruction modifications and tailored training might be based in part on variations in performance accuracy due to defect type, such as those observed in the present study, in order to focus on the most problematic defects. For instance, defects with high hits and low false alarms, such as the welds in the current study, do not require intervention. On the other hand, defects with exceptionally low hits (e.g., extra welds and puffiness) or with both high hits and high false alarms (e.g., creases) are top candidates for additional training or improved work instructions to help inspectors better differentiate defective and acceptable parts. Intervention may include practice with additional limit samples for defects associated with poor performance to provide more examples of acceptable and rejectable parts and better delineate the dividing line between the “worst” accept and the “best” reject. Intervention might also involve examining additional examples that illustrate consequences of the defect during subsequent processing.
Results from significant correlations between false alarms on the one hand and the workload and SSSQ scores on the other hand also suggest an approach for tailored training. Such analyses indicated that inspectors who committed more false alarms reported higher subjective workload, lower task engagement, greater distress, and increased worry. As with similar results concerning inspector response times, this outcome suggests that inspectors may have realized they were committing errors on some level. Such information may prove useful to optimize performance by training inspectors how to recognize the signs of high workload and stress associated with high false alarms.
In accordance with the significant correlation between lighting and perceptual sensitivity, one relatively simple apparatus modification is to introduce improved lighting for the inspection process. Although the illumination in the present study exceeded typical office lighting of 500 lux, it fell short of the 1,000 lux generally recommended for difficult inspection work (Megaw, 1979). Thus, increasing room illumination or providing supplemental portable lighting might improve inspection performance. However, care must be taken to avoid glare and other issues associated with increased lighting. For example, successful detection of fingerprints on the flags depended on angling the banner away from the light. Thus, increased ambient lighting might actually interfere with detection of this particular defect.
Finally, the present study largely confirms previous research suggesting that only minimal improvements may be achieved through personnel selection based on inspector demographics. For example, no significant associations with performance were identified for experience, vision, education, or professional certification in inspection. The present study did show reduced perceptual sensitivity for females and for older inspectors. However, such results do not conform to the majority of previous research, which has not provided consistent evidence to support gender or age differences in performance (See, 2012).
Directions for Future Research
The current results suggest numerous directions for future research to replicate aspects of the present study and further enhance understanding of visual inspection. First, additional research on the use of confidence ratings during inspection is warranted to identify a more effective confidence rating scale that better distinguishes among levels of inspector confidence. Second, future work might involve a direct experimental manipulation of inspector instructions (low-consequence vs. high-consequence instructions) to probe the impacts of the consequences of missed defects. Finally, research designed to train inspectors how to recognize the signs of high workload and stress associated with high false alarms could support efforts to determine whether such training can improve inspection performance.
Key Points
Eighty-two qualified inspectors in the U.S. Nuclear Security Enterprise correctly rejected 85% of defective items, a level of performance that falls short of 90% accuracy and is not vastly superior to the industry average of 80%.
Inspectors incorrectly rejected 35% of acceptable items, a level that is much higher than typically seen in industry, in an attempt to avoid the severe consequences associated with incorrectly passing defective items.
Use of a phased inspection approach based on inspector confidence ratings did not prove to be an effective or efficient technique to improve the overall accuracy of the process, but this outcome may have been due to the manner in which inspectors were asked to rate their confidence.
Inspector workload ratings verified that inspection is a workload-intensive task, dominated by mental demand and effort.
As inspectors took longer to inspect parts, hits, false alarms, and distress increased whereas confidence decreased. Inspectors who committed more false alarms reported higher subjective workload, lower task engagement, greater distress, and increased worry.
Footnotes
Acknowledgements
This work was sponsored and funded by the National Nuclear Security Administration. The author would like to thank the Human Factors Department staff at Sandia National Laboratories as well as Caren Wenner, Katherine Curry, and Glenda Mathes for their contributions.
Author(s) Note:
The author(s) of this article are U.S. government employees and created the article within the scope of their employment. As a work of the U.S. federal government, the content of the article is in the public domain.
Judi E. See is a systems analyst at Sandia National Laboratories in Albuquerque, New Mexico. She obtained a PhD in experimental psychology from the University of Cincinnati in 1994 and certification in human factors from the Board of Certification in Professional Ergonomics in 2009.
