Agent Acceptance Test Suite & SLO Checklist

An interactive, saveable checklist that captures deterministic and stochastic acceptance tests, safety injects, performance buckets, SLO targets, and incident/runbook details needed before deploying agent-based automation to production.

Interactive Tool

Agent Acceptance Test Suite & SLO Checklist

Use this checklist to confirm the functional correctness, safety constraints, latency expectations, observability, and operational runbooks for an agent before it reaches production. Save the results to preserve acceptance evidence, support post-deployment reviews, and trigger follow-up actions.

The unique name or identifier for the agent under test (include version/tag).
Person or team responsible for the agent and for incident response.
How to reach the owner or on-call rotation when an incident occurs.
Select every acceptance test or artifact you have completed.
Yes if the CI/CD pipeline shows green for the agent's branch/tag.
Number of runs or examples used for stochastic sampling tests.
Summarize observed failure modes, flakiness, and confidence level.
Which safety or adversarial injects were run?
List any failing cases and how they were mitigated or why risk is acceptable.
Desired p50 end-to-end latency for the agent's critical path.
Desired p95 end-to-end latency.
Desired p99 end-to-end latency.
Choose the acceptance bucket used by monitoring and SLO calculations.
Typical and peak requests per minute or concurrent tasks the agent must support.
Select the monitoring and observability items validated.
Percent of requests that may fail without an SLO breach (e.g., 1.0 for 1%).
Target availability over the measurement window (e.g., 99.9).
Period over which the SLO is measured.
Steps to stabilize service on breach (auto-rollback, scale up, reduce traffic, disable feature, etc.).
Canonical runbook or playbook for on-call responders.
Who to notify and how (paging, slack, email), including escalation timing.
Planned canary rollout approach for production.
Concrete, measurable criteria that require immediate rollback.
Checks to run after deployment (smoke, sampling, golden signals verification).
List the monitoring alert names or IDs that indicate serious degradation.
Final readiness judgement after executing the checklist.
Summarize rationale, outstanding risks, and list reviewers who signed off.
Links to CI job runs, test reports, sampling logs, dashboards, and recordings.
Anything else the team should track after deployment.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.