Experimentation Playbook — Experiment Plan Template

A practical, structured interactive experiment plan template you can fill, save, and reuse. Captures hypothesis, metrics, sample assumptions, randomization, instrumentation checks, analysis plan, monitoring and stopping rules, pre-registration signatures, and a filled example for a conversion test.

Interactive Tool

Experiment Plan Template

Use this structured experiment plan to design rigorous, low-risk tests that produce trustworthy, actionable results. Complete the fields below to pre-register your test, ensure measurement readiness, and make the analysis and rollout straightforward. Save the plan to your project or team library so results are auditable and learnings are reusable.

A short, descriptive name for the experiment (e.g., 'Homepage CTA color test - Q3').
Person or role accountable for running the experiment.
Planned launch date (YYYY-MM-DD). Update when scheduling is finalized.
Explain the problem, why it matters, and the opportunity you expect this test to address. Keep it focused on customer behavior or business outcome.
State the hypothesis clearly: 'If we [treatment], then [measurable outcome] because [rationale]'.
The single metric you will use to decide success (include exact definition and numerator/denominator).
List metrics you will monitor to ensure you are not causing harm or unintended regressions (e.g., revenue per session, page load time, error rate).
The level at which randomization occurs.
Describe how units are identified and any clustering or dependency to consider.
Who is in scope (geography, platform, audience segments) and exclusion rules.
Enter the historical baseline value used for power/sample calculations (as a percent for rates, or raw units for other metrics).
Smallest relative change you care about detecting (e.g., 5 for 5%).
Typical default is 0.05. Use a lower alpha if multiple tests or comparisons increase false positive risk.
Typical default is 0.8 (80%). Higher power requires larger samples.
Record how sample size was calculated, assumptions used, and the final required sample per group. If you used an external calculator, paste the link and inputs. (Note: plan saves assumptions; platform may be extended to calculate sample sizes automatically.)
Describe exactly what the treatment is, including copy, layout, timing, funnel changes, exposures, and content. Include asset names or version IDs so engineers can implement an exact match.
Describe the control condition so comparisons are exact and reproducible.
How units are assigned to variants.
Include hashing, seeding, stratification variables, and any rollout granularity constraints.
Ensure events and metrics are tracked and tied to the experimental unit.
Specify the statistical tests, aggregation windows, handling of outliers, transformation, multiple comparison corrections, and subgroup analyses. Clearly state the decision rule tied to the primary metric.
How often the experiment will be monitored for anomalies.
Define thresholds for automated alerts (e.g., >10% drop in revenue per session) and who is notified.
Describe conditions for stopping early for harm, futility, or clear success, and how interim looks will be managed to avoid false positives.
If the test is successful, how will you roll out? If harmful, how will you roll back? Include timeline, risk mitigation, and communication steps.
Person who pre-registered the plan (typed name is acceptable for pre-registration).
Date of pre-registration (YYYY-MM-DD).
Anything the implementation or analysis teams must know (engineering tickets, QA, analytics work, external vendors).
An illustrative example you can copy and adapt.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.