Experiment Protocol Template

Interactive protocol template to pre-specify hypothesis, metrics, sampling, instrumentation, guardrails, analysis, and rollout. Saves structured protocols for reproducible, trustworthy experiments.

Interactive Tool

Experiment Protocol

Use this protocol to pre-specify the design and analysis of a controlled test so results are trustworthy and actionable. Complete each field before running the experiment. Where helpful, link to instrumentation, dashboards, or analytic notebooks. Saving this form archives a record for review, preregistration, and decision-making.

Short descriptive name for this experiment (team, feature, or hypothesis).
Person responsible for running the experiment and making the final decision. Include role/title.
Where questions about the protocol should be directed.
State the hypothesis in causal form: if we do X (treatment), then Y (primary metric) will change by Z because... Explain why this matters. Keep it testable and specific.
Name the main metric you'll use to judge success and give an operational definition (numerator, denominator, aggregation window, user/unit level).
Define the minimum effect size, direction, or range that would change decisions (e.g., > 3% lift in conversion, or reduction of 10s in latency). Be concrete.
List additional metrics to monitor (e.g., retention, error rate, revenue per user). Mark which are guardrails (safety checks) and why.
Describe who/what units are eligible (users, sessions, stores) and any exclusion or inclusion rules. Specify time windows or geography if relevant.
Describe how units are identified and sampled (user-id, session-id, store-id) and the experimental unit for analysis. Explain any clustering or correlated structure.
Report sample size calculation or justification. Include baseline rate, minimum detectable effect, power, alpha, and any assumptions. If a sequential or adaptive plan is used, describe stopping rules. If you cannot calculate a sample size, describe pragmatic constraints and planned sensitivity.
How will treatment be assigned? Choose the method and explain why it fits the design.
If using blocking or stratification, list strata and how blocks are formed and assigned. Explain any balancing variables.
Describe the treatment(s): what users see or receive, timing, variations, dosage, and any technical rollout details. Include control group definition.
Select the data and instrumentation items you have verified. Provide links or notes below if necessary.
Paste links to event definitions, analytic notebooks, dashboards, and relevant tickets. Helpful for auditors and analysts.
Identify potential harms or business risks and describe automated and manual rollback plans, monitoring thresholds, and contacts for rollbacks.
Note any user consent, privacy, regulatory, or ethics concerns and approvals required. Mention data retention and access controls.
How will you monitor the experiment in real time? Specify cadence (hourly/daily), dashboards, and alert thresholds for guardrail metrics.
Describe any planned interim analyses, how often they occur, and statistical or business rules that would stop or change the experiment. If none, say so.
Pre-specify the primary analysis method (e.g., difference-in-means, regression with covariates, survival analysis), covariates, handling of missing data, and multiple comparisons correction. Include model formulas or notebook references.
Specify the alpha (e.g., 0.05) and target power used in sample-size calculation or reasoning.
Who will run the analysis, and when will results be delivered? Include expected dates for unlocking and final report.
Who can access raw and aggregated data? Where will results be stored and how long will they be retained?
State the rule you will use to make decisions (e.g., deploy if primary metric >= threshold with p < 0.05 and no guardrail violations). Describe possible post-experiment actions.
If the experiment succeeds, how will the change be rolled out? Include phasing, feature flags, and monitoring during rollout.
Mark yes if this protocol will be archived and locked before exposing units to treatment.
Optional: link to archived preregistration or version identifier.
Call out assumptions, anticipated threats to validity, dependencies, or known limitations that reviewers should consider.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.