Business Continuity & Crisis Runbook Template
A practical playbook and runbook template to prepare for and respond to major supply, plant, IT, or natural-disaster disruptions. Includes role definitions, decision thresholds, response checklists, communication templates, alternate-supplier planning, testing cadence, and after-action steps so teams respond quickly, consistently, and recover faster.
Purpose and Scope
This playbook helps teams prepare for, respond to, and recover from major disruptions (plant outages, supplier failures, IT incidents, natural disasters). It provides clear ownership, decision thresholds, response runbooks, communications templates, and a testing cadence so you can reduce downtime, protect customers, and restore operations faster.
How to use this playbook
- Assign ownership for the playbook at the site and enterprise levels.
- Maintain a short Critical Process Inventory for each location and update quarterly.
- Tailor each runbook to the local plant, supplier network, and regulatory needs—keep the structure consistent so people can act under pressure.
- Run exercises and capture after-action improvements; treat the playbook as a living document.
Core Components
-
Critical Process Inventory & Single-Point-of-Failure (SPOF) List
Document and prioritize the processes, machines, systems, and suppliers whose failure would stop production or create major customer impact. For each item record:
- Business impact (financial, safety, regulatory)
- Maximum acceptable outage (RTO) and data loss tolerance (RPO) where relevant
- Primary owner and backup owner
- Single point of failure notes (supplier, part, skill, software)
- Mitigations already in place (spare parts, alternate suppliers, SOPs)
Use this inventory to focus preparedness and to create prioritized runbooks.
-
Escalation Paths & Decision Tree
Define simple, measurable triggers that move an incident up the chain. Keep decision rules crisp—avoid vague wording under stress. Example thresholds:
- Level 1 (Local): Operator or shift supervisor handles events causing <2 hours production loss and no customer impact.
- Level 2 (Site): Plant manager and core response team when outage >2 hours, product quality at risk, or safety concern.
- Level 3 (Enterprise): Executive response when outage >8 hours, major customer impact, regulator notification, or multi-site effect.
Decision tree (example):
- Is production stopped? If no, monitor and record. If yes, continue.
- Can a local workaround restore operation within RTO? If yes, implement and monitor. If no, escalate to Site.
- Does the issue affect multiple customers or sites? If yes, escalate to Enterprise and notify stakeholders.
-
Communications Plan & Contact Matrix
Prepare concise templates and a maintained contact matrix. Include:
- Internal notification cascade (shift, operations, maintenance, safety, HR)
- External notifications (customers, suppliers, regulators, media) and when to notify each
- Primary and secondary communication channels (phone, SMS, email, mass-notify tool)
- Spokesperson assignment and approved message templates
Contact matrix example (maintain in the runbook):
- Plant Incident Lead: name, role, mobile, backup
- Maintenance Lead: name, role, mobile
- Supply Chain Lead: name, role, mobile, alternate supplier contact
- Site Safety: name, role, mobile
- Customer Success / Sales: name, role, mobile
-
Alternate Suppliers, Locations & Workarounds
For each SPOF, document the validated alternate options and the steps to engage them. Record commercial terms, lead times, and any qualification steps. Include physical alternatives such as production-transfer checklists for moving production to another site or approved contract manufacturers.
-
Scenario-based Runbooks
Create short, actionable runbooks for common major scenarios. Each runbook should include:
- Trigger (how the scenario is detected)
- Immediate safety steps
- Initial containment actions (first 60–120 minutes)
- Escalation path and decision thresholds
- Recovery steps and owner for each step
- Customer and regulator notification actions
- Estimated timeframes and critical dependencies
Suggested scenario templates to prepare: utility outage, raw-material supplier failure, critical equipment breakdown, major quality event, IT/OT outage, cyber incident, natural disaster, pandemic.
-
Response Checklists (Templates)
Provide short checklists that front-line responders can follow under pressure. Example quick checklist for a plant outage:
- Ensure worker safety — confirm safe status of personnel.
- Stabilize process if safe to do so; isolate faulty equipment.
- Notify Plant Incident Lead and log incident start time.
- Run initial diagnostics and confirm whether workaround is possible.
- Activate communications: internal team and customers as required by thresholds.
- Engage maintenance or vendor support per SPOF list.
- Record decisions and next steps in incident log.
-
Recovery & Restoration Plan
Document stepwise restoration tasks, quality checks required before returning product to normal flow, disposition of product made during the event, and validation steps required for safety or regulatory compliance.
-
Testing, Exercises & Metrics
Maintain a test calendar and run regular tabletop or live exercises. Track simple KPIs:
- Mean time to detect (MTTD)
- Mean time to acknowledge (MTTA)
- Mean time to restore (MTTR) vs target RTO
- Number of successful alternate supplier activations
- Exercise completion rate and action-item closure rate
Run at least one cross-functional tabletop per year and a focused drill for high-risk SPOFs every 6–12 months.
-
After-Action Review & Continuous Improvement
After every incident or exercise, run a structured after-action review to capture what worked, what didn’t, root causes, and committed corrective actions with owners and dates. Update the playbook and runbooks promptly and communicate changes to affected teams.
-
Versioning & Ownership
Maintain a short change log at the front of the playbook with version, date, author, and summary of changes. Assign a playbook owner responsible for quarterly reviews and post-incident updates.
Quick-start: Build a Runbook in One Day (90–120 minute core, then expand)
- Gather the plant manager, maintenance lead, supply chain lead, and shift supervisor for 30 minutes.
- Identify 3 highest-priority SPOFs and record owners, RTO, current mitigations.
- Write a short response checklist for one SPOF (the one most likely to occur).
- Create a minimal contact matrix and one customer notification template.
- Schedule a tabletop within 30 days to validate the runbook and capture improvements.
Templates & Examples
Include these as annexes in your maintained playbook (copy and adapt):
- Incident Log template (time, event, decision, owner)
- Customer notification template (short factual message and expected next update)
- Supplier engagement checklist (order, confirm, expedited shipping, quality checks)
- Plant-transfer checklist (qualifications, tooling, approvals, shipping)
Tailoring and Ownership Guidance
Keep the runbooks concise and action-oriented. Local sites should copy the playbook and tailor runbooks to their lines, equipment, suppliers, and local regulations. Enterprises should keep a canonical version that defines thresholds, reporting requirements, and corporate escalation rules.
Where to start
Begin with the Critical Process Inventory and one simple runbook for the most likely SPOF. Run an exercise to test it, then iterate. The best playbooks are short, tested, and current.
Related Resources
- Supplier continuity checklist
- IT disaster recovery runbook
- Mass notification templates
- After-action review guide
Annex: Minimal Runbook Template (copyable)
- Title and scenario (who, what, where)
- Trigger and detection method
- Immediate safety actions
- Containment actions (first 0–2 hours)
- Escalation steps and owners
- Recovery steps and quality checks
- Customer/regulator communications and templates
- Log of decisions, time stamps, and actions
- Post-incident actions and assigned owners
Keep this playbook near operations, with digital copies accessible to the response team and key stakeholders. Treat it as a living instruction set: test, learn, update, and share improvements.
Discussion
Comments and conversation will live here.