PARTSIXTY6 HFAA • SECTION 1 • LESSON 1.3

Why Maintenance Errors Occur

FAA Human Factors for Aviation Mechanics

Learning Objective

By the end of this lesson, you should be able to explain the principles of human error; distinguish slips, lapses, mistakes, omissions, commissions, and extraneous actions; identify conditions and preconditions that increase the probability of an unsafe act; trace an error through active failures, latent conditions, and failed defenses; and select practical controls that prevent, detect, or contain maintenance error.

FAA ACS alignment: AM.I.L.K2 — Human error principles; AM.I.L.K10 — Conditions/preconditions for unsafe acts; AM.I.L.K11 — Types of human errors.

FOUNDING IDEA

A maintenance error is rarely explained by the final action alone. To prevent recurrence, examine what the person intended, what conditions shaped the action, what information and resources were available, and why the system did not prevent or detect the unwanted result.

1. Error is an outcome, not a complete explanation

Begin with the difference between intention and result

Aircraft maintenance is exacting work performed by people inside a technical and organizational system. The work may involve hundreds of individually correct actions: identifying the aircraft, opening access, isolating energy, selecting data, removing hardware, protecting open systems, inspecting parts, installing components, applying adjustments, completing tests, restoring configuration, accounting for tools, and recording the result. A maintenance error occurs when an action or omission produces a result that differs from the intended or required state. The definition describes the outcome; it does not yet explain why the difference occurred.

That distinction is essential. Saying that a mechanic “made an error” identifies neither the error mechanism nor the control that should change. The action may have been a slip during a familiar task, a forgotten step after an interruption, a mistaken diagnosis based on incomplete evidence, an intentional shortcut that had become normal in the workplace, or reckless disregard of a known and substantial risk. These events require different responses. More training may help a knowledge deficiency, but it will not correct an unreadable work card, an impossible schedule, a misleading connector design, or a norm that rewards undocumented work.

Human error is not random in the sense of being beyond analysis. Human performance varies, but errors tend to become more likely under identifiable conditions. Weak lighting reduces visual information. Similar parts make selection errors easier. A long task with hidden dependencies increases memory demand. A poorly controlled interruption can separate intention from action. Fatigue slows detection and weakens attention. Informal communication can strip a message of aircraft identity, configuration, limits, or completion status. Once these mechanisms are understood, the maintenance system can add defenses that support the correct action and expose an incorrect condition.

The purpose of Human Factors is not to remove personal responsibility. A certificated mechanic remains responsible for exercising privileges within applicable requirements, technical data, authorization, and professional standards. Human Factors strengthens that responsibility by replacing vague judgment with disciplined analysis. It asks what was required, what was intended, what actually occurred, what influenced the difference, and which controls should prevent, detect, or contain a recurrence. This approach protects both safety and fairness because it bases the response on evidence rather than assumption.

A useful investigation therefore avoids stopping at the nearest person. The technician’s action may be the last movable part in a longer chain that includes engineering, planning, supervision, stores, tooling, documentation, staffing, training, software, workplace design, and organizational priorities. The final action still matters, but it must be placed in context. If the analysis ends with “be more careful,” the system has learned almost nothing. If it identifies the error type, contributing conditions, missing defenses, and required verification, it can produce a measurable improvement.

Professional principle: Describe the event precisely before judging it. Separate the required state, the person’s intention, the action taken, the actual result, and the defenses that should have controlled the result.

2. The anatomy of a maintenance event

Requirement, plan, execution, result, and detection

Every controlled maintenance action begins with a required state. The requirement may come from a regulation, an airworthiness directive, approved or acceptable technical data, a manufacturer’s instruction, an operator’s program, an engineering authorization, or an organization’s procedure. The requirement defines what must be accomplished and what condition must exist afterward. Human performance converts that requirement into a plan, the plan into physical action, and the action into an aircraft or component state. Inspection, testing, and records provide evidence that the achieved state agrees with the requirement.

An error can enter at any transition. The wrong requirement may be selected because effectivity was misunderstood. The correct requirement may be interpreted incorrectly. A correct plan may be executed incorrectly because the technician operates the wrong control. A correct action may be incomplete because one item is omitted. The physical work may be correct while the record identifies the wrong component. An incorrect condition may be created and then missed by an inspection that repeats the same assumption. The event is not only the moment at which a wrench turns; it includes the information path before the action and the verification path afterward.

Consider a torque requirement. The controlled path includes identifying the correct joint, selecting current and applicable data, understanding the specified units and condition, selecting suitable calibrated equipment, setting the value, applying the tool with correct technique, confirming completion, and recording or inspecting the work when required. “Incorrect torque” could result from selecting the wrong fastener, reading the adjacent table row, confusing inch-pounds and foot-pounds, failing to compensate for an adapter, setting the tool incorrectly, losing place after interruption, using an out-of-calibration tool, or signing for work not completed. The same visible outcome can have several different causal pathways.

Detection is a separate part of the event. Some errors are self-revealing: the component will not fit, the test fails immediately, or the system produces an unmistakable warning. Others remain latent because the installation looks normal, the verification is weak, or the incorrect condition appears only under a specific load or operational mode. A robust process does not assume that absence of an immediate problem proves correctness. It identifies the error modes that could survive the task and assigns independent, meaningful defenses against them.

Recovery also matters. If an error is detected, the aircraft or component must be placed in a controlled state. The organization must establish the actual configuration, protect affected systems, determine the scope, use authorized technical support where necessary, accomplish corrective work, repeat disturbed inspections or tests, and make accurate records. Hiding, minimizing, or informally correcting a discrepancy can destroy evidence and allow related conditions to remain. Timely reporting is therefore part of technical control, not an admission of incompetence.

HFAA 1.3: Error Anatomy
HFAA 1.3 — Error Anatomy

3. Errors of execution: slips and lapses

When the plan is correct but the action does not follow it

A slip is an execution error. The person intends the correct action but performs another action. Selecting an adjacent switch, transposing digits while entering a value, reaching for the wrong identical container, or moving a control in the wrong direction can be slips. Slips are more likely when actions are highly practiced and automatic, when controls or items are similar, when attention is divided, or when the interface makes the incorrect action easy. Expertise does not eliminate slips; in some circumstances, strong habits allow familiar actions to run with little conscious monitoring.

A lapse is commonly associated with memory. The intended action is not completed because the intention is lost or the task state is forgotten. A mechanic may intend to reinstall a safety device after a test, then be interrupted and resume at the next visible step. A cap may be placed temporarily near an opening and forgotten during restoration. A measurement may be taken correctly but not entered before the display changes. A required reinspection may be deferred until access is available and then disappear from the active work picture. The plan was not necessarily wrong; continuity between intention and completion failed.

Interruptions are a powerful mechanism for both slips and lapses. The incoming demand competes with the current task for attention and working memory. When the technician returns, the physical scene may look familiar enough to suggest a point farther ahead than the last positively completed step. Simply telling people not to be distracted is not sufficient because many interruptions are operationally necessary: an urgent radio call, a safety warning, a request for assistance, a change in aircraft status, or a supervisor’s instruction. The task needs a controlled pause and restart method.

Effective interruption control externalizes task status. Before turning away, make the aircraft safe, protect loose or open items, mark the exact step, record any value not yet transferred, and state what remains incomplete. On return, reestablish aircraft identity and configuration, review the governing instruction, inspect the work area, and restart from the last positively verified point rather than the point that feels familiar. For critical sequences, repeating an earlier step may be safer than guessing. A visible status marker is stronger than relying on memory alone.

Design can reduce slips and lapses. Distinct labels, physical separation, connector keying, color used with text rather than alone, tool and part segregation, checklists at the point of use, positive feedback, step signoffs that reflect real completion, and independent inspection can all help. The defense should match the mechanism. If two controls are confused, make identity and selection unmistakable. If an intention is lost after interruption, preserve task state. If a repetitive action is skipped, make completion visible and testable.

HFAA 1.3: Slip Lapse Mistake
HFAA 1.3 — Slip Lapse Mistake

4. Errors of planning and judgment: mistakes

When the action follows the plan but the plan is wrong

A mistake occurs when the person carries out an intended plan, but the plan itself is unsuitable for the situation. The execution may be smooth and deliberate. A technician may troubleshoot the wrong system because the symptom was classified incorrectly, apply a familiar procedure to a different configuration, accept a measurement as normal because an incorrect limit was recalled, or replace a component based on confirmation of the first plausible diagnosis. Unlike a slip, the action matches the intention; the problem lies in the selected rule, interpretation, or reasoning.

Some mistakes are rule-based. A known rule or procedure is applied in the wrong circumstances, or a correct rule is applied incorrectly. The technician may remember a standard value from one model and use it on another, select a familiar maintenance practice without recognizing a configuration change, or continue with a troubleshooting branch after a prerequisite has not been satisfied. Strong habits and experience can make the incorrect rule feel convincing. Verification of identity, effectivity, preconditions, units, and limits is therefore essential even for familiar work.

Other mistakes are knowledge-based. The situation is unfamiliar, no obvious rule applies, and the person must reason from available evidence. Troubleshooting intermittent faults, evaluating unexpected damage, interpreting contradictory indications, and working with a newly modified system can create this demand. Knowledge-based work consumes attention and working memory. It is vulnerable to incomplete information, premature conclusions, confirmation bias, time pressure, and overconfidence. The safest response may be to slow the process, obtain additional data, consult authorized technical support, and test competing explanations.

Confirmation bias is the tendency to notice or value evidence that supports an existing belief while discounting evidence that conflicts with it. In maintenance, an early diagnosis can shape every later observation. A replaced unit that temporarily clears a fault may be treated as proof even though a disturbed connector caused the change. A leak found near a stain may appear to explain the entire condition while another source remains. A good troubleshooting method asks what evidence would disprove the favored diagnosis and whether the final test reproduces the relevant operational conditions.

Training helps prevent mistakes when it builds usable mental models rather than isolated facts. The technician should understand how the system normally behaves, what indications depend on one another, what configuration assumptions govern the test, and how faults can propagate. Current data, clear decision points, diagnostic aids, peer consultation, engineering support, and independent review strengthen judgment. However, consultation is only useful when uncertainty and contradictory evidence are communicated honestly. Asking for help after presenting a conclusion as certain can simply recruit another person into the same assumption.

5. FAA types of human error: omission, commission, and extraneous action

Classifying what happened in the task

The FAA maintenance training framework commonly identifies three practical types of human error: omission, commission, and extraneous action. These categories describe how the performed work differed from the intended task. They can coexist with the psychological classifications of slip, lapse, and mistake. For example, an omission may be caused by a lapse, while an incorrect installation may be a commission caused by a mistaken interpretation. Using both views produces a more useful analysis: one describes the task deviation, and the other helps explain the human mechanism.

An error of omission means a required action was not performed. Examples include failing to install a cotter pin, leaving a connector unsecured, omitting a lubrication point, not removing protective material, failing to complete a required test, or leaving a discrepancy out of the turnover. Omissions are particularly dangerous when the missing action leaves little visible evidence. A completed-looking assembly can conceal an absent internal item. Defenses include sequence control, positive completion indicators, independent inspection where appropriate, functional testing, configuration checks, and records that identify open work precisely.

An error of commission means an action was performed, but it was performed incorrectly or should not have been performed in that manner. Installing an incorrect part, connecting lines to the wrong ports, applying the wrong torque, cutting material beyond a limit, using an incorrect fluid, or entering an incorrect configuration value are commissions. The action may be deliberate in the sense that the person intended to act, while the result remains unintended. Controls focus on positive identity, applicability, differentiation, correct data, clear units, tooling, feedback, and verification of the achieved state.

An extraneous error is an action not required by the task that introduces an unwanted change. A technician may disturb an adjacent connector unnecessarily, remove additional hardware and fail to restore it, change a setting while navigating a test menu, damage a nearby wire with a tool, or introduce foreign material into an open system. Extraneous actions remind us that maintenance affects more than the target component. Work boundaries, access planning, protection of adjacent systems, configuration control, tool discipline, and post-work inspection help detect unintended disturbance.

The classification must remain factual. “Omission” does not automatically mean laziness, and “commission” does not automatically mean recklessness. Ask why the required step was absent or the incorrect action appeared reasonable at the time. Was the step visible? Was the part distinguishable? Was the work interrupted? Did the procedure reflect the configuration? Was the technician working from memory? Did the test provide clear feedback? Classification should guide the search for controls, not become a label attached to the person.

Task-level typeMaintenance meaningTypical exampleControl direction
OmissionA required action is not completed.A locking device or required test is omitted.Make sequence, open status, completion, and verification visible.
CommissionAn action is completed incorrectly.The wrong part, port, value, fluid, or method is used.Strengthen identity, applicability, differentiation, feedback, and inspection.
ExtraneousAn unnecessary action changes the system.An adjacent item is disturbed or an unintended setting is changed.Define work boundaries, protect nearby systems, and inspect for disturbance.
HFAA 1.3: Task Error Types
HFAA 1.3 — Task Error Types

6. Unsafe acts, conditions, and preconditions

Why the context surrounding the action matters

An unsafe act is an action or omission that increases risk or creates an unwanted condition. Conditions and preconditions are the circumstances that make that act more likely, harder to detect, or more consequential. They may exist in the person, task, equipment, workplace, team, supervision, or organization. The distinction matters because correcting only the act can leave the same conditions in place for the next person. Human Factors therefore examines both the immediate action and the environment that shaped it.

Personal preconditions include fatigue, stress, illness, reduced alertness, unfamiliarity, high workload, distraction, physical limitation, and loss of situational awareness. These conditions do not guarantee an error, and their presence should not be used as an automatic excuse. They change the margin available for reliable performance. A professional technician monitors fitness for duty, recognizes degraded capability, uses available controls, and reports when the task cannot be performed safely under the actual condition.

Task and equipment preconditions include poor access, weak lighting, excessive noise, ambiguous instructions, hidden markings, similar components, unsuitable tooling, unclear feedback, repeated interruptions, long or repetitive sequences, and work that depends heavily on memory. Time pressure can amplify each of these. A difficult task may still be completed correctly, but the probability of error rises as the person must compensate for more mismatches. Controls should remove or reduce the mismatch rather than assume skill will permanently overcome it.

Team and supervisory preconditions include unclear responsibility, ineffective turnover, authority gradients, incomplete coordination, inadequate staffing, conflicting priorities, and failure to act on reported hazards. A supervisor may never say “take a shortcut,” yet a schedule, staffing decision, or repeated reward for rapid completion can communicate that production has priority over verification. Conversely, a strong supervisor clarifies the standard, provides resources, protects a justified stop, and treats early reporting as a chance to control risk before it reaches the aircraft.

Organizational preconditions develop over time. Recurrent procedure difficulties may be tolerated. Deferred improvements may accumulate. Workarounds may become normal because they appear efficient and have not yet produced a visible event. Purchasing decisions may create look-alike parts or incompatible systems. Training may lag behind modification. Performance measures may reward output without measuring rework, reporting, or verification quality. These conditions are often distant from the technician’s action, but they influence what is possible at the worksite.

7. The FAA Dirty Dozen as an operational scan

Twelve common maintenance error-producing conditions

The FAA’s maintenance Human Factors material uses the Dirty Dozen as a practical list of common contributors: lack of communication, complacency, lack of knowledge, distraction, lack of teamwork, fatigue, lack of resources, pressure, lack of assertiveness, stress, lack of awareness, and norms. The list is not a complete accident model and does not prove causation merely because one item is present. It is a prompt for recognizing conditions that can combine and reduce performance margin.

Lack of communication can remove identity, status, limits, or intent from the information path. A statement such as “it is ready” is inadequate if the receiver does not know which aircraft, which system, what work was completed, what remains open, and what verification occurred. Closed-loop communication requires a clear message, confirmation of critical elements, and resolution of disagreement. Written status should support rather than contradict the verbal turnover.

Complacency can arise from repeated exposure to familiar tasks and normal results. Expectation begins to replace observation. The technician may see what is usually present rather than what is actually present, or treat an inspection as confirmation that nothing changed. Countermeasures include deliberate variation of search, use of objective criteria, attention to critical items, independent inspection, and asking what would be different if an abnormal condition existed.

Lack of knowledge exists when the person does not possess or cannot apply the information needed for the task. It may involve a new system, modification, material, software version, inspection method, or unusual damage. The safe response is to recognize the boundary, stop, use current technical data, and obtain qualified assistance. Guessing, copying an adjacent aircraft, or silently experimenting on an airworthy product is not a substitute for authorized information and competence.

Distraction breaks the connection between the current intention and the next action. It may come from another person, a radio, a mobile device, an alarm, a nearby task, or the technician’s own concern. Because some interruptions cannot be avoided, the process should preserve task state. Secure the work, mark the interruption point, record untransferred values, and verify the last positively completed step before resuming.

Lack of teamwork appears when people share a task or aircraft state without shared understanding and coordinated responsibility. Teams require clear roles, timely information, cross-checks, and willingness to speak. One person’s assumption that another person restored, inspected, or recorded an item can leave work incomplete. The defense is explicit assignment and acceptance: who performs, who verifies, who controls configuration, and who closes the record.

Fatigue can reduce alertness, attention, working memory, reaction, judgment, and error detection. It is influenced by sleep opportunity and quality, time awake, time of day, workload, health, commuting, and repeated schedule disruption. Caffeine and motivation may temporarily change how alert a person feels without fully restoring performance. Fitness-for-duty decisions should be made early, and tasks should use controls appropriate to the actual fatigue risk.

Lack of resources includes more than missing tools. It can involve insufficient time, staffing, access, lighting, data, parts, test equipment, training, protective equipment, engineering support, or workspace. Improvisation sometimes appears productive but can transfer hidden risk into the aircraft. The mechanic should identify the missing resource precisely and use the organization’s process to obtain it or reschedule the work rather than silently lowering the standard.

Pressure may come from schedules, customers, supervisors, peers, cost, weather, aircraft availability, or the individual’s desire to finish. Pressure narrows attention toward the immediate goal and can make uncertainty feel like delay. A professional response separates urgency from technical acceptability. State the unresolved requirement, explain the consequence, propose a controlled path, and obtain the needed authority or resource. No delivery promise changes the required aircraft condition.

Lack of assertiveness is a failure to express a safety-relevant concern clearly enough for action. Assertiveness is not aggression. It uses specific facts: identify the aircraft or part, state the observed condition, connect it to the applicable requirement or uncertainty, describe the risk, and request a decision or stop. If the concern is not resolved, use the escalation path. Respect for experience or rank must not require silence about an unsafe condition.

Stress can be created by workload, conflict, personal circumstances, uncertainty, environmental conditions, or perceived consequences. Moderate activation may focus performance, while excessive or sustained stress can narrow attention, accelerate decisions, and impair working memory. The control depends on the source: clarify priorities, reduce simultaneous demands, pause, seek support, improve information, or reassign the task. Concealing stress removes the opportunity to manage it.

Lack of awareness means an incomplete understanding of the current situation, the changes produced by the work, or the possible consequences. A technician focused on one component may not recognize that another team energized the system, that an adjacent connector was disturbed, or that a configuration change invalidated the next test. Frequent status checks should ask: What is the aircraft state now? What changed? What remains open? Who else is affected? What could happen next?

Norms are informal ways of working that a group accepts. Some norms support safety, such as challenging uncertain part identity. Others normalize deviation: skipping a difficult step, using memory instead of current data, signing before completion, or treating recurring discrepancies as harmless. The phrase “we always do it this way” is not technical authority. Compare the norm with the governing requirement, report conflicts, and correct the system that made the workaround attractive.

Dirty Dozen application: Do not choose one label and stop. Scan for combinations, identify the evidence, and connect every contributor to a specific control. Two or more weak conditions may interact even when none would have produced an error alone.
HFAA 1.3: Dirty Dozen
HFAA 1.3 — Dirty Dozen

8. Active failures, latent conditions, and failed defenses

How weaknesses align across the maintenance system

An active failure is an unsafe act whose effects are closely connected in time and place to the operation. Incorrectly connecting a line, omitting a required item, entering the wrong configuration, or releasing incomplete work can be active failures. They are visible near the event and are therefore easy to treat as the whole cause. However, the action may have been shaped by conditions created earlier and elsewhere.

Latent conditions are weaknesses embedded in design, planning, procedures, training, staffing, supervision, purchasing, software, workplace layout, or organizational decisions. They may remain unnoticed until they combine with local circumstances. Examples include a procedure with ambiguous effectivity, repeated acceptance of incomplete turnover, similar parts stored together, a test that cannot detect a critical installation error, or a schedule that routinely depends on overtime. Latent does not mean harmless; it means the condition can wait for an opportunity.

Defenses are barriers intended to prevent an error, detect it before release, or contain its effect. Physical keying can prevent an incorrect connection. Positive part identification can prevent an ineligible installation. A required independent inspection can detect a critical assembly error. A functional test can reveal an incorrect result. Tool control can prevent foreign-object release. Accurate records can protect the next person from acting on a false aircraft state. Each defense has a purpose and a failure mode.

Several defenses can appear independent while sharing the same weakness. A technician installs a component from memory; an inspector checks it from the same memory; the functional test confirms only that the unit powers up, not that its configuration is correct; and the record copies the planned part number rather than the installed identification. Four activities occurred, but none independently verified applicability. Defense count is less important than defense quality, independence, and connection to the specific error mode.

When weaknesses align, an error path opens. The useful response is not to search for one “root cause” and ignore the rest. Identify which immediate action created the condition, which preconditions influenced that action, which latent conditions made the situation possible, and which defenses failed to interrupt the path. Then strengthen multiple levels. Correct the immediate discrepancy, improve the local task, and address systemic recurrence. Layered controls are valuable when one imperfect defense does not make every other defense ineffective.

9. Error, deviation, and accountability

Not every unwanted act belongs in the same category

An unintentional error differs from an intentional deviation. In an error, the person did not intend the unwanted result and may not have intended the incorrect action. In a deviation, the person knowingly departs from a procedure, rule, or expected method. The result may still be unintended. A deviation can be motivated by a desire to complete the work, compensate for a poor procedure, help a colleague, or meet a schedule. Good intent does not make the departure safe or authorized, but understanding the reason is necessary to prevent recurrence.

At-risk behavior occurs when a person chooses a method that increases risk but the risk is not fully appreciated, or the behavior has become accepted because previous outcomes appeared successful. Repeatedly using an unofficial shortcut without an immediate event can create false evidence that the shortcut is safe. This is normalization of deviance: the absence of a bad outcome is interpreted as proof that the departure is acceptable. The system must restore the standard and correct the conditions encouraging the behavior.

Reckless behavior is different. It involves conscious disregard of a substantial and unjustifiable risk. Deliberately falsifying work, knowingly concealing a serious discrepancy, or intentionally releasing an aircraft while aware that a required safety condition is absent cannot be managed as an ordinary slip. Fair accountability does not mean that all conduct receives the same response. It means that the response considers intention, knowledge, choices, organizational influence, history, and the quality of the applicable rules.

A just culture seeks both learning and accountability. If every report leads automatically to punishment, people may hide errors and the organization loses safety information. If every action is excused as “human error,” standards and professional responsibility collapse. A balanced process protects honest reporting, distinguishes inadvertent error from deliberate misconduct, examines whether a reasonable peer could have made the same decision under the same conditions, and applies corrective or disciplinary action consistently.

Mechanics support fair accountability by reporting facts promptly and preserving evidence. State what was required, what was understood, what action occurred, what condition was found, and what remains uncertain. Avoid speculation about motive. Do not alter records to make the event look cleaner, and do not perform an undocumented correction that removes the opportunity to assess scope. Accurate early reporting allows the aircraft to be protected and the organization to learn before the same condition reaches another task.

10. Traps in error investigation

Hindsight, outcome bias, and the search for a single cause

Hindsight makes the correct path appear more obvious after the outcome is known. Investigators can see the complete sequence, compare all records, and focus on the clue that eventually proved important. The technician at the time had only the information available at each moment, mixed with normal signals, interruptions, assumptions, and uncertainty. A fair analysis reconstructs that changing information state. Ask what the person could see, what the procedure displayed, what others communicated, and why the chosen action appeared reasonable or acceptable then.

Outcome bias judges the quality of a decision mainly by its result. The same shortcut may be praised when the aircraft leaves on time and condemned when a discrepancy appears. Conversely, a careful decision may produce an unfavorable result because information was incomplete. Safety analysis should evaluate the process and evidence available at the time, not only whether luck produced a good or bad outcome. Near misses and successfully detected errors are valuable because they expose weak controls without requiring an accident.

The search for a single root cause can oversimplify a system event. One factor may be necessary, but prevention often depends on several contributors. If a wrong component was installed, the analysis should not stop at selection. How was the part ordered, received, identified, segregated, issued, compared with effectivity, physically differentiated, inspected, tested, and recorded? Which opportunities existed to interrupt the error? Addressing several points creates resilience even if one control fails again.

“Failure to follow procedure” is also a starting fact, not always a sufficient explanation. Determine whether the procedure was current, applicable, available, understandable, physically usable, consistent with the aircraft, and supported by required resources. Determine whether deviations were known, tolerated, or encouraged. If the procedure was suitable and the choice was deliberate, accountability remains relevant. If the task could not be accomplished as written, the organization must correct the procedural or engineering mismatch rather than repeatedly blame workers for adapting.

Training is another common default action. Training is appropriate when the event reveals missing knowledge, an incorrect mental model, or a skill deficiency that instruction and practice can change. It is weak when the person already knew the requirement and the real problem was missing equipment, misleading design, production pressure, ambiguous responsibility, or a normalized shortcut. Corrective action should name the mechanism it changes and define how the organization will verify effectiveness.

11. Prevention, detection, and containment

Design controls around the error mechanism

Prevention makes the incorrect action less likely or impossible. Examples include physical keying, positive component segregation, controlled data selection, configuration-specific work cards, adequate lighting and access, suitable tooling, clear units, standardized handover, and realistic staffing. Prevention should be placed as close as practical to the point at which the error could enter. A warning in a distant manual is weaker than an interface that makes the correct identity visible at the worksite.

Detection assumes prevention may fail and seeks independent evidence before release. Inspection, operational checks, leak tests, duplicate inspections, measurements, witness marks, software verification, inventory, and record reconciliation can serve this purpose when they target the specific error. “Look over the work” is too general. The detection method should define what condition must be observed, under what configuration, using which acceptance criteria, and without relying on the same assumption that produced the work.

Containment limits consequence if an error escapes. Isolation, protective devices, system monitoring, staged testing, controlled engine runs, warning systems, and conservative release conditions can prevent a local discrepancy from becoming a larger event. Recovery procedures then establish how to stop the test, protect personnel and equipment, capture evidence, return the aircraft to a known state, and escalate technical support. A mature system expects that imperfect humans and imperfect defenses require safe recovery paths.

Checklists and signoffs are effective only when connected to real task state. A checklist should support attention and memory, not become a substitute for understanding. A signoff should represent completed and verified work, not expected future completion. If boxes are routinely marked in groups after the task, the record no longer controls sequence. If an independent inspection is performed by someone who watched and accepted the original assumption, independence may be weak. Control quality depends on how the defense functions in practice.

The mechanic contributes to all three layers. Before work, identify critical steps, likely confusion points, required resources, and final tests. During work, maintain configuration, protect open systems, preserve task position, challenge unexpected conditions, and communicate changes. After work, inspect for intended and unintended effects, confirm restoration, complete required testing, account for tools and materials, and record the actual result. If evidence does not support release, stop; confidence is not a substitute for verification.

12. Maintenance scenario: a repeated connector discrepancy

Operational scenario

An aircraft returns with an intermittent indication after a recent avionics replacement. The fault clears during ground troubleshooting. A connector in a restricted location appears seated, and the locking feature is difficult to see. The technician has performed similar replacements many times. The shift is near its end, another aircraft is waiting for the same test equipment, and the work card uses a diagram that does not clearly show the viewing orientation. The technician reseats the connector, the indication remains normal, and the aircraft is prepared for return to service.

The visible maintenance action is reseating the connector. A narrow analysis might say that the technician should have checked it more carefully. A systems analysis first establishes the required condition: correct connector identity, full engagement, positive locking, proper support, acceptable pin and shell condition, and completion of the applicable test. It then asks what the technician intended, what evidence was available, whether the diagram and access supported verification, and what defense should have detected an incompletely locked connector.

Several error-producing conditions are present. Familiarity can support complacency. Restricted access reduces visual and tactile feedback. The ambiguous diagram increases interpretation demand. Shift-end fatigue and schedule pressure reduce margin. Shared test equipment creates a production incentive. The intermittent fault encourages a quick explanation. None proves that an error occurred, but together they create a credible path by which an incomplete connection could be accepted.

The event could involve a mistake if the technician concluded that disappearance of the indication proved the connector was secure. It could involve a slip if the hand contacted or locked an adjacent connector instead. It could involve an omission if the required locking verification or follow-on test was not completed. It could involve a commission if the connector was forced or secured by an unauthorized method. Classification depends on evidence, not the investigator’s first impression.

Strong controls would improve access and lighting, confirm connector and channel identity, provide an orientation-correct illustration, define positive locking criteria, inspect the connector and pins, protect harness routing and strain relief, and perform the test under conditions capable of reproducing the fault. If the work crosses shifts, the turnover must identify the exact fault history, actions taken, actual connector state, tests completed, and unresolved uncertainty. The aircraft should not be described as ready merely because the indication is absent at that moment.

The corrective action should also examine recurrence. Have other technicians reported difficulty seeing the lock? Are similar connectors present? Does the work card require an inspection that cannot physically be performed from the normal position? Are intermittent faults repeatedly closed after reseating without identifying the mechanism? Does the schedule encourage abbreviated tests? Local correction protects this aircraft; systemic correction protects future aircraft and technicians.

13. A mechanic’s practical error-control routine

Before, during, after, and when something goes wrong

Before the task: establish aircraft and component identity, authority, data effectivity, required configuration, critical sequence, tools, parts, access, environmental conditions, inspection points, tests, and responsibilities. Ask where an omission, incorrect selection, or unintended disturbance could occur. Identify any Dirty Dozen conditions already present. If a critical resource or requirement is missing, resolve it before the aircraft is opened or the system state becomes more complex.

During the task: keep the instruction and aircraft state synchronized. Positively identify items before acting. Protect removed parts, open systems, and adjacent equipment. Treat unexpected resistance, configuration, damage, indication, or procedural conflict as information. Control interruptions by preserving the last completed step. Communicate any change affecting another person. Do not let familiarity replace the current data, and do not let schedule pressure redefine the acceptance standard.

After the task: verify both intended work and possible collateral effects. Confirm installation, security, locking, routing, clearances, configuration, servicing, restoration of protective devices and panels, tool and material accountability, required inspections, and applicable operational or functional tests. Compare the evidence with the requirement. Complete records from the actual work and actual component identity, not from the planned task or memory.

If an error or uncertainty is found: stop the affected activity, protect people and the product, preserve evidence, identify the aircraft and system state, and report through the authorized process. Do not hide the condition or perform an informal correction that makes the event impossible to reconstruct. Determine the extent of disturbance and whether other aircraft, parts, tasks, records, or shifts could be affected. Corrective work must use the same discipline as the original task.

When reviewing recurrence: name the error type, contributing conditions, latent weaknesses, and failed defenses. Select controls at more than one level. Assign ownership and a completion date. Then verify effectiveness using evidence: observation, audit, repeat discrepancy rate, quality findings, employee feedback, or successful detection tests. A corrective action is not complete because a memo was issued; it is complete when the relevant risk is demonstrably better controlled.

Release rule: The aircraft state must be supported by applicable data, observable evidence, completed verification, and accurate records. When those elements disagree, stop and resolve the disagreement before release.

What We Learned

A maintenance error is a difference between the intended or required result and the result actually produced. The word “error” describes an outcome but does not, by itself, explain intention, mechanism, contributing conditions, or accountability. Effective analysis separates the requirement, plan, action, result, detection, and recovery.

Slips and lapses occur when a suitable intention is not executed as planned; mistakes occur when the selected plan or judgment is unsuitable. At the task level, errors may be classified as omission, commission, or extraneous action. Conditions and preconditions—such as fatigue, distraction, pressure, weak communication, poor access, ambiguous data, or workplace norms—can increase probability and weaken detection.

The FAA Dirty Dozen provides a practical scan for common maintenance contributors. Active failures occur close to the event, while latent conditions may have been created much earlier through design or organizational decisions. Defenses must prevent, detect, or contain specific error modes and should not all depend on the same assumption.

Fair accountability distinguishes inadvertent error, at-risk behavior, intentional deviation, and reckless conduct. Learning and responsibility are not opposites. Mechanics protect safety by using current data, preserving task state, verifying the actual result, reporting discrepancies promptly, and refusing to substitute confidence, routine, or urgency for evidence.

FAA alignment and controlled sources

Lesson Check — 15 Questions

Select the best answer, then check your work.

1. Which statement best defines a maintenance error?

2. A technician intends to select the correct adjacent switch but operates the wrong one. This is primarily a:

3. A required safety device is forgotten after an interruption. The task-level error is:

4. A mechanic follows a familiar procedure that does not apply to the modified configuration. This is best described as:

5. Which is an error of commission?

6. Which is an extraneous action?

7. What is the strongest restart after an interruption?

8. The FAA Dirty Dozen should primarily be used to:

9. Which is a latent condition?

10. Why can several apparent checks still provide weak defense?

11. Which response best reflects fair accountability?

12. What is the main weakness in “be more careful” as corrective action?

13. What does physical fit prove about a replacement part?

14. Which defense is most meaningfully independent?

15. When an error or unresolved discrepancy is found, the first professional priority is to:

← GO BACK