Andon / Incident Flow: Operator Steps, Escalation & Reset

A clear, operator-focused playbook that standardizes the immediate actions, required data, escalation timing, reset criteria, and follow-up documentation after an Andon trigger — designed to cut time-to-response, contain problems at the frontline, and create reliable inputs for root-cause work.

Purpose

This playbook defines a concise, repeatable operator-driven response when an Andon is pulled. It focuses on fast containment, consistent data capture for effective escalation, and clear reset and follow-up steps so incidents stop repeating.

Scope

Applies to all frontline operators, shift technicians, maintenance, and engineering teams on the production floor. Use for machine, process, tooling, material, and quality triggers that require immediate attention.

Key definitions

  • Andon: Any signal (pull cord, button, dashboard alert) indicating abnormal condition requiring assistance.
  • Containment: Immediate operator actions that prevent product flow or process from producing more bad output or causing damage.
  • Owner: The role assigned to perform the action (operator, shift tech, maintenance).
  • Data snapshot: Small, standardized set of information captured at escalation to help troubleshooting.

Flow (compact)

  1. Operator signals Andon and selects a cause code from the standard taxonomy.
  2. Operator performs immediate containment and quick fixes (owner: operator) within 5 minutes.
  3. If unresolved after containment attempt, escalate to shift technician with required data snapshot (owner: operator → shift tech) within 10–30 minutes depending on severity.
  4. If still unresolved, shift technician escalates to maintenance/engineering according to the escalation rule-set and priority.
  5. When normal operation resumes, the incident is reset and documented: root cause, countermeasure, verification plan, and learning notes.

Operator immediate checklist (first 5 minutes)

  1. Pull Andon (or acknowledge the alarm on the dashboard).
  2. Select one primary cause code from the standard taxonomy and, if needed, one secondary code (examples below).
  3. Stop the line for affected unit(s) if required by safety/quality rules.
  4. Perform immediate containment steps (examples below).
  5. Record a short description (one sentence) of what you saw and what you did on the operator log/dashboard.

Containment examples (operator-owned)

  • Remove defective parts to a quarantine bin and tag them.
  • Hold production for the affected SKU/lot and mark work centers.
  • Perform a quick tool change or adjustment if trained and authorized.
  • Perform a safety stop and notify neighbors if hazard present.

Required data snapshot for escalation

When escalating, include the following minimum dataset so the shift tech or maintenance can triage quickly:

  • Timestamp of Andon pull
  • Station/line ID and machine ID
  • Operator name and shift
  • Primary cause code and short description (1–2 sentences)
  • Parts affected (SKU/lot) and quantity at risk
  • Recent process readings or alarm values (temperature, pressure, speed, OEE status snapshots)
  • Photos or short video clip if helpful

Escalation matrix and timing

  • Operator → Shift Technician: 10–30 minutes from Andon if containment didn’t restore normal operation. Shift tech must arrive, gather snapshots, and attempt diagnostic or temporary fix.
  • Shift Technician → Maintenance: Escalate immediately if safety, electrical, or mechanical failure suspected, or if the shift tech cannot restore operation within agreed time window.
  • Maintenance → Engineering: Escalate for repeat failures, complex root causes, design changes, or when a permanent countermeasure is required.
  • Times and priorities should be tuned per line criticality and customer commitments; publish a concise escalation table by line on the dashboard.

Reset criteria (how to know when to reset Andon)

Only reset the Andon when:

  • Normal process parameters are restored and verified (operator or tech confirmation).
  • Products in process are either quarantined or reworked per quality rules.
  • A temporary countermeasure is in place if permanent fix is pending, and the risk is acceptable.
  • Required documentation and snapshots are recorded in the log/dashboard.

Documentation & follow-up (within 24–48 hours)

  1. Owner records root cause hypothesis, countermeasure, verification plan, and who will own the permanent fix.
  2. Assign a follow-up ticket if engineering or cross-functional work is required.
  3. Capture lessons in the team huddle (use the Andon incident as a learning prompt) and update the standard work or poka-yoke when appropriate.

Training, placement & visibility

  • Place the flow as a short visual on operator dashboards, next to Andon pull points, and in the quick reference binder at the station.
  • Practice the flow in Gemba walks and routine drills so it becomes muscle memory.
  • Include the taxonomy and snapshot checklist in operator training; make it easy to choose the right cause code under stress.

Suggested metrics to track

  • Time-to-first-response (Andon → operator containment action)
  • Time-to-escalation (Andon → shift tech arrival)
  • Time-to-reset (Andon → normal operation)
  • % of Andons resolved by operator without escalation
  • Repeat incidents by machine/component (to catch chronic problems)

Sample cause taxonomy (starter)

  • Material: wrong material, contamination, shortage
  • Machine: tool wear, sensor fault, mechanical jam
  • Process: setup error, parameter drift, inconsistent feed
  • Human: operator error, missing instruction
  • Quality: dimensional out-of-spec, surface defect
  • Safety: hazard present

Quick operator-to-tech message template

"Andon on line L2, machine M4. Cause code: Machine – mechanical jam. Containment: stopped line, quarantined 12 parts. Snapshot: RPM dropped to 0 at 09:23, vibration alarm active. Operator: J. Doe."

Common pitfalls to avoid

  • Vague descriptions — always use the taxonomy and a one-sentence factual observation.
  • Skipping the snapshot — wastes technicians’ time and delays correct diagnosis.
  • Resetting prematurely — leads to repeated Andons and hidden root causes.

Next steps for teams

Adopt this playbook at the line level, run a one-week pilot, and track the suggested metrics. Use learnings from the pilot to tune times, cause codes, and escalation responsibilities to local realities.


Discussion

Comments and conversation will live here.