Business Continuity & Crisis Response Checklist
A practical, plant-level playbook with clear incident classification, immediate containment actions, team roles and phone-tree templates, communication scripts, alternate-sourcing and recovery-priority checklists, and a post-incident review template so teams can respond faster, reduce downtime, and restore production with confidence.
Purpose and scope
This playbook helps plant teams respond to major disruptions—plant outages, critical equipment failures, supplier collapse, severe quality events, utility loss, cyber incidents, and other events that threaten production or customer commitments. It gives a concise, repeatable runbook: how to classify an incident, who does what first, what must be communicated, how to find alternate supply, how to sequence recovery, and how to capture lessons afterward.
How to use this checklist
- Quickly classify the incident using the Incident Classification guide below.
- Assemble the Crisis Response Team (CRT) and run the Immediate Containment Actions checklist.
- Follow communication templates to notify internal stakeholders and customers.
- Execute alternate-sourcing and recovery-sequencing checklists to restart production with prioritized SKUs/customers.
- Run the Post-Incident Review to capture root causes and permanent fixes.
Includes
- Incident classification
- Immediate containment actions
- Communication templates (internal / external)
- Alternate sourcing checklist
- Team roles & phone tree template
- Recovery sequence priority list
- Post-incident review template
Incident classification (use this immediately)
Decide severity quickly — this sets the level of escalation, who must be notified, and the cadence of updates.
- Severity 1 — Business-critical: Plant-wide outage or event causing complete stoppage for multiple lines, imminent customer loss, regulatory exposure, or safety-critical condition. Escalate to senior leadership and trigger full Crisis Response Team.
- Severity 2 — Major: One or more key lines down, large volume impact, or multi-shift disruption requiring cross-functional response and alternate sourcing within days.
- Severity 3 — Moderate: Single-line or localized equipment failure, quality hold affecting customers but manageable with temporary controls or schedule adjustments.
- Severity 4 — Minor / localized: Low-impact issues handled by on-shift responders and routine maintenance processes.
Immediate containment actions (first 60–120 minutes)
Assign a named owner for each action. Do not assume someone else will act.
- Ensure safety: secure area, treat injuries, lockout/tagout (LOTO) as required — Safety Lead.
- Stabilize production: stop affected line(s) or materials to prevent further damage — Shift Lead / Operations.
- Preserve evidence: hold suspect product, serial numbers, process data, and shift logs — Quality.
- Contain quality exposure: implement quarantine, hold-and-inspect, or stop-ship decisions — Quality Manager.
- Isolate systems if cyber or contamination suspected — IT / Engineering.
- Start incident log with timestamped actions and decisions — Incident Scribe (Operations Coordinator).
Team roles & phone tree (template)
Maintain a printed and digital phone tree with primary and secondary contacts. Example essential roles:
- Incident Commander / Plant Manager — overall decision authority
- Operations Lead / Shift Lead — production control
- Maintenance Lead — equipment diagnosis & repair
- Quality Lead — hold decisions, root-cause testing
- Supply Chain / Purchasing — supplier contact & alternate sourcing
- Safety / EHS — personnel safety and regulatory reporting
- Communications / Customer Relations — external messages and customer updates
- IT / OT — systems isolation and recovery (if applicable)
- Incident Scribe — maintains timeline, notes, and action register
Phone tree tip: list two contacts per role and include preferred contact times and escalation rules (e.g., if unreachable in 15 minutes, call the alternate).
Communication templates
Use short, consistent messages. Update cadence depends on severity (e.g., hourly for S1, every 4 hours for S2, EOD for S3).
Internal alert (sample)
Subject: Plant Alert — [Severity X] — [Short description]
Body: We have a [brief description: e.g., line 2 electrical failure] classified as Severity [X]. Safety is confirmed. Production impact: [expected downtime / affected SKUs]. Incident Commander: [name]. Next update: [time]. Key immediate actions: [bullets].
Customer notice (short sample)
Subject: Service update — potential delay for [SKU/Order]
Body: We are experiencing a disruption at our [plant]. We are working to contain and recover production. At this time estimated impact to your order(s): [if known]. We will provide updates at [cadence]. Primary contact: [name, email, phone].
Alternate sourcing checklist
- Identify critical components/materials by SKU and lead time.
- Pull approved supplier list and contact primaries and known alternates.
- Check safety/quality approvals and required certifications for alternates.
- Confirm logistics capacity and expedite options (air, split shipments, cross-dock).
- Estimate cost and lead time trade-offs; route approvals if cost exceed thresholds.
- Record PO changes and any contractual customer notifications required.
Recovery sequence priority list (decision guide)
Prioritize recovery using a simple scoring matrix: customer criticality, revenue impact, regulatory requirements, ease/time to restart, and safety/quality risk. Example priorities:
- Safety-critical systems and compliance (always first)
- High-priority customers / contractual commitments
- High-volume SKUs with short lead times
- Low-risk product families that can bring capacity back quickly
- Full line requalification and restart
Document who authorized priority choices and any customer agreements about partial shipments or late deliveries.
Post-incident review template (use within 72 hours)
- Timeline: create a minute-by-minute timeline from detection to full recovery.
- Root cause: immediate cause, contributing factors, system failures.
- Corrective actions: short-term fixes (who, due date) and long-term CAPAs (owner, target date).
- Communication effectiveness: what worked, gaps, recommended script changes.
- Supplier & logistics assessment: how alternate sourcing performed and gaps.
- Training/drill needs: who needs refreshers and when.
- Follow-up verification: how and when each corrective action will be verified and closed.
Testing, training & maintenance
Keep this playbook alive by scheduling recurring activities:
- Tabletop exercises: quarterly for high-severity scenarios.
- Full-scale drills: annual or after major process changes.
- Phone-tree validation: monthly check-ins to verify contacts and alternates.
- Supplier contingency review: biannual review of alternate suppliers and lead times.
- Runbook review: after any incident or significant process change.
Quick printable checklist (one-page)
Keep a laminated one-page checklist at the supervisor station and in the control room. Sections to include at-a-glance: Incident classification box, Safety confirm, Immediate containment 5 bullets, Incident Commander & phone, Communications: internal & customer template lines, Recovery priority top 3.
Appendix: practical notes and templates
Use the following as editable templates in your local copy of this playbook: phone-tree spreadsheet, incident log spreadsheet, customer notice Word template, supplier contact list, alternate-sourcing decision log, and post-incident CAPA tracker. If you acquire or copy this playbook into an enterprise domain, tailor role names, contact lists, approval limits, and customer escalation requirements to local needs.
This playbook is intended to be operational. Convert the listed templates and checklists into accessible forms and printed quick cards for every shift. Practice regularly — a tested plan beats an assumed plan during a crisis.
Discussion
Comments and conversation will live here.