Business Continuity Incident Runbook (Plant Outage)

A concise, actionable runbook to guide initial actions, communications, triage, containment, alternate capacity options, and recovery milestones during a major plant outage. Includes checklists, role assignments, sample communication templates, prioritization guidance, and post-incident steps so teams respond consistently and restore customer service as quickly as possible.

Purpose and Scope

This runbook guides the immediate response to a major plant outage that threatens safety, production, or customer commitments. It is focused on the first 72 hours but includes steps for recovery, customer communications, and post-incident review. Use it to respond consistently, protect people and product, triage critical needs, and restore production reliably.

Activation Criteria

  • Loss of production affecting > X% of daily capacity for more than Y hours.
  • Major equipment failure that cannot be corrected within a single shift.
  • Utility outage (power, compressed air, steam) that halts critical lines.
  • Supplier failure that interrupts lines producing customer-critical SKUs.
  • Event that creates regulatory, environmental, or safety exposure.

Immediate Safety Check (First 15 minutes)

  1. Ensure all personnel are safe; account for everyone on-site. (Supervisor & Safety Lead)
  2. Isolate hazards and secure energy sources if required (LOTO). Do not restart equipment until cleared. (Maintenance & Safety)
  3. Treat any environmental release or injury as the highest priority; call emergency services if needed. (Plant Manager)
  4. Establish an Incident Command location (physical or virtual) and identify Incident Lead. (Plant Manager)

Roles & Responsibilities (assign immediately)

  • Incident Lead: Owns the response, decisions, and communications.
  • Operations Lead: Coordinates production triage, containment, and alternate capacity.
  • Maintenance Lead: Diagnoses equipment causes, estimates repair timeline, and executes repairs.
  • Safety/Environmental Lead: Confirms site safety and regulatory reporting needs.
  • Supply Chain/Materials Lead: Manages parts, suppliers, and material reroutes.
  • Customer & Commercial Lead: Manages customer notifications, prioritization, and commitments.
  • Communications Lead: Issues internal and external messages and maintains a single source of truth.
  • Logistics/Shipping Lead: Adjusts shipments, carriers, and dock schedules.

Initial Incident Actions (First 30–60 minutes)

  1. Confirm activation and notify core Incident Team by phone/SMS/Teams using the emergency contact list.
  2. Run the Immediate Safety Check and secure the scene.
  3. Collect the facts: what failed, when, affected lines/SKUs, current output, estimated immediate loss (hours), and any injuries or environmental impacts.
  4. Record the incident start time and create an incident log (who, what, when). Keep one person assigned to capture updates.
  5. Decide whether to stop, hold, or continue unaffected lines with documented reasoning.
  6. Identify critical customers and critical SKUs using the Critical Products & Customer Escalation list (see template below).

Containment & Triage (First 1–4 hours)

Use the triage matrix: Safety > Regulatory > Critical Customer SKUs > High Volume SKUs > Low-volume, low-priority items.

  • Isolate affected components and preserve evidence for root cause analysis where required.
  • Stabilize partial lines to resume limited production if safe and practical.
  • Segregate salvageable product and tag clearly (Hold/Quarantine). Track lot numbers and timestamps.
  • If product quality is in doubt, sample and test before release. Follow Quality Lead guidance.

Alternate Capacity Playbook (Quick decision checklist)

  1. Can sister plant or nearby facility run the SKU? If yes, check spare capacity and lead time to transfer work.
  2. Can a contract manufacturer subcontract production? Contact prequalified partners from supplier contingency list.
  3. Is expedited shipping from alternate sites viable to meet customer needs? Evaluate cost vs. customer criticality.
  4. Can product be rerouted from inventory at distributor or warehouse? Pull from safety stock where appropriate.
  5. Prioritize SKUs: map customer commitments and rank shipments to preserve top customers and contractual obligations.

Communication Templates (use and adapt)

Initial Internal Alert (short):

Incident Alert: [Plant] experienced [brief description]. Incident started at [time]. Incident Lead: [name]. Safety: [status]. Production impact: [high/medium/low], affected lines: [list]. Incident Command: [location]. Next update: [time].

Customer Notification (use Commercial Lead):

Dear [Customer Name],
We experienced an unplanned outage at [Plant] on [date/time] affecting production of [SKU(s)]. Our incident team has been activated and we are working to restore service. Current estimate for partial/ full recovery: [initial estimate]. We are prioritizing [critical SKUs/customers] and will update you at [times]. Contact: [name, phone, email].

Supplier Request Template:

Urgent Support Request: We need immediate assistance for [part/service] for outage at [Plant]. Required by: [time/date]. Potential impact: [brief]. Contact: [name]. Please confirm availability and earliest delivery.

Recovery Milestones (example timeline)

  1. Hour 0–2: Incident declared, safety secured, incident log opened, initial customer notifications sent.
  2. Hour 2–8: Containment, triage, establish temporary workarounds, secure replacement parts if available.
  3. Day 1: Partial production resumed (if feasible) or alternate capacity confirmed. Update customers with revised ETAs.
  4. Day 2–3: Full production recovery actions underway; validate product quality and ramp plan.
  5. Day 3+: Plan for sustained operations, full ramp to target throughput, and handback to normal operations with clear acceptance criteria.

Decision Points & Escalation

  • Escalate to Regional/Enterprise Continuity when outage > 24 hours or impacts multiple sites/customers.
  • Escalate to Legal/Regulatory when there is a potential reportable event or customer/agency notification requirement.
  • If repairs will exceed X hours (preset threshold), trigger alternate capacity options immediately.

Post-Incident Steps (once production stabilizes)

  • Verify product quality for all affected lots and release only after Quality Lead approval.
  • Perform formal root cause analysis (5-Why, fishbone, or equivalent) and document corrective actions with owners and due dates.
  • Conduct a customer-facing summary where appropriate: what happened, customer impact, remediation taken, prevention plan.
  • Update runbooks, supplier lists, emergency contacts, and spare parts lists based on gaps discovered.
  • Schedule a Lessons Learned session within 7 days and track improvement items in the action register.

Quick Reference Checklist (printable)

  • Notify Incident Team and establish Incident Command.
  • Immediate Safety Check completed.
  • Record incident start time and open incident log.
  • Identify affected SKUs and critical customers.
  • Decide line status (stop/hold/run) and document rationale.
  • Initiate containment and quarantine suspect product.
  • Contact suppliers for emergency parts and capacity partners if needed.
  • Send initial customer notification for impacted orders.
  • Define short-term recovery milestones and next update times.
  • Schedule Post-Incident Review and assign RCA owner.

Templates & Attachments You Should Maintain

  • Current Emergency Contact List (names, roles, mobile/pager, alternates).
  • Critical Products & Customer Escalation matrix (SKU → customers → priority → recovery options).
  • Prequalified alternate manufacturers and contract terms for expedited work.
  • Spare parts inventory and lead times for critical equipment.
  • Sample incident log spreadsheet and RCA template.

Practice & Testing

Run tabletop exercises at least twice per year and a full-scale exercise annually. Validate contact lists, supplier response times, and alternate capacity agreements. Update this runbook after each exercise.

Final Notes

Keep the runbook concise and accessible at the plant floor, in the site intranet, and to the enterprise continuity team. During an incident, maintain a single source of truth and keep customer communications honest, timely, and focused on commitments.


Discussion

Comments and conversation will live here.