Data Incident Triage & Communication Form
Structured interactive triage form to capture incident facts, assign owners, record mitigation and rollback plans, notify stakeholders, and drive post‑incident follow-up. Saves submissions for audit, timelines, and analytics.
{"Title":"Data Incident Triage & Communication Form","IntroductionHtml":" ","SubmitLabel":"Record Triage","SuccessMessage":"Incident triage saved. Share the reference with your team and follow the escalation timeline.","DataType":"DataIncidentTriageV1","SchemaVersion":1,"Fields":[{"Key":"incident_summary","Label":"Brief incident summary","FieldType":"textarea","HelpText":"Summarize what happened in 1–2 sentences (what, observed when, who reported).","Required":true},{"Key":"detection_timestamp","Label":"Detection timestamp (UTC)","FieldType":"text","HelpText":"Use ISO 8601 where possible (e.g., 2026-08-27T15:04:00Z).","Required":true},{"Key":"reported_by","Label":"Reported by (name & role)","FieldType":"text","HelpText":"Person or system that detected or reported the incident.","Required":false},{"Key":"teams_affected","Label":"Teams affected","FieldType":"checkbox","Options":[{"Value":"data_engineering","Label":"Data Engineering"},{"Value":"analytics","Label":"Analytics / BI"},{"Value":"product","Label":"Product"},{"Value":"ops","Label":"Operations / SRE"},{"Value":"security","Label":"Security"},{"Value":"compliance","Label":"Compliance"},{"Value":"customer_support","Label":"Customer Support"},{"Value":"sales","Label":"Sales"},{"Value":"other","Label":"Other"}],"HelpText":"Select all teams impacted.","Required":true},{"Key":"data_products_impacted","Label":"Data products / pipelines / reports impacted","FieldType":"textarea","HelpText":"List dataset names, pipeline IDs, dashboards, reports, or artifacts affected.","Required":true},{"Key":"initial_severity","Label":"Initial severity (S1–S4)","FieldType":"select","Options":[{"Value":"S1","Label":"S1 — Critical (major outage or data loss)"},{"Value":"S2","Label":"S2 — High (significant degradation/incorrect data)"},{"Value":"S3","Label":"S3 — Medium (localized, workaround exists)"},{"Value":"S4","Label":"S4 — Low (minor, non-urgent)"}],"HelpText":"Choose the best-fit severity based on business impact.","Required":true},{"Key":"likely_root_cause","Label":"Likely root cause (initial)","FieldType":"textarea","HelpText":"Share early hypotheses (config change, schema change, bad upstream data, deployment, etc.).","Required":false},{"Key":"immediate_mitigation_steps","Label":"Immediate mitigation steps taken or planned","FieldType":"textarea","HelpText":"Describe actions taken so far (stop jobs, revert change, apply patch, isolate dataset) and next immediate steps.","Required":true},{"Key":"rollback_required","Label":"Is a rollback required?","FieldType":"yesno","HelpText":"If yes, provide the rollback plan in the following field.","Required":true},{"Key":"rollback_plan","Label":"Rollback plan (if applicable)","FieldType":"textarea","HelpText":"Include specific commands, versions, snapshot timestamps, and an owner for the rollback. If not applicable, enter 'N/A'.","Required":false},{"Key":"communications_recipients","Label":"Stakeholders to notify (roles/groups)","FieldType":"textarea","HelpText":"e.g., Execs, Product Owners, Customers (if impacted), Legal, Compliance, Support. Include channel (email, Slack, ticket).","Required":true},{"Key":"communications_template","Label":"Suggested communications (copy & adapt)","FieldType":"textarea","HelpText":"Template: '[Short summary of incident], impacted systems: [list], known impact: [summary], mitigation in progress: [actions], ETA for next update: [time], contact: [name/email].' Include links to incident page or ticket.","Required":false},{"Key":"escalation_contacts","Label":"Escalation contacts (names, roles, contact info)","FieldType":"textarea","HelpText":"List primary and secondary contacts with preferred contact method and expected SLA to respond.","Required":true},{"Key":"sla_window_hours","Label":"Suggested SLA response window (hours)","FieldType":"number","HelpText":"Typical suggestions: S1=1 hour, S2=4 hours, S3=24 hours, S4=72 hours. Enter a numeric value.","Required":true},{"Key":"evidence_links","Label":"Evidence links and saved artifacts","FieldType":"textarea","HelpText":"Links to logs, dashboards, dataset snapshots, queries, tickets, or storage locations where evidence is preserved.","Required":false},{"Key":"post_incident_checklist","Label":"Post‑incident checklist (select items to include in postmortem)","FieldType":"checkbox","Options":[{"Value":"preserve_evidence","Label":"Preserve raw evidence (snapshots, logs) and store securely"},{"Value":"document_timeline","Label":"Document timeline of events and actions taken"},{"Value":"root_cause_analysis","Label":"Perform root cause analysis and validate with data"},{"Value":"corrective_actions","Label":"Define corrective actions and owners"},{"Value":"test_and_verify","Label":"Test fixes in staging and verify data correctness"},{"Value":"external_notifications","Label":"Notify customers or external parties if required"},{"Value":"schedule_postmortem","Label":"Schedule postmortem meeting and assign facilitator"},{"Value":"update_runbooks","Label":"Update runbooks, checks, and monitoring as needed"}],"HelpText":"Choose items to ensure the incident is fully addressed.","Required":true},{"Key":"post_incident_owner","Label":"Owner for post‑incident follow‑up","FieldType":"text","HelpText":"Person responsible for ensuring postmortem actions are completed.","Required":true},{"Key":"postmortem_schedule_date","Label":"Planned postmortem date/time (UTC)","FieldType":"text","HelpText":"Provide a date/time for the review. Use ISO 8601 where possible.","Required":false},{"Key":"additional_notes","Label":"Additional notes","FieldType":"textarea","HelpText":"Any other relevant context, blockers, or constraints.","Required":false}]}
Purpose
Use this form to capture essential information when a data incident is detected. It standardizes triage, owner assignment, mitigation, communications, and post‑incident follow‑up so teams can act quickly and preserve evidence.
Severity guide
- S1 (Critical): Major outage or data loss affecting many customers or core business flows — immediate response required.
- S2 (High): Significant degradation or incorrect data impacting key processes or decision‑making.
- S3 (Medium): Localized issues with limited business impact; workaround available.
- S4 (Low): Minor anomaly or non‑urgent data quality issue requiring investigation.
Roles
- Incident Lead: Owns the operational response and decisions during the incident.
- Engineering Lead: Manages technical mitigations, rollbacks, and root cause investigation.
- Communications Lead: Prepares stakeholder updates and external messages.
- Forensics / Compliance: Engaged when legal, regulatory, or sensitive data issues are suspected.
This form supports operational response only. For legal, regulatory, or forensic investigations, escalate to Security or Compliance immediately.
Discussion
Comments and conversation will live here.