Causal Inference Practical Checklist
A practical, decision-focused checklist to validate causal claims: refine the causal question, choose an appropriate design, state identification assumptions, pre-specify measurement and analysis, run balance and robustness checks, evaluate threats to validity, and prepare transparent communication and decision guidance.
Causal Inference Practical Checklist
Use this checklist when you must assess whether an observed effect plausibly reflects a causal relationship (for decisions, policies, or operational changes) and when a randomized experiment is not feasible. For each item, collect evidence, record decisions, and flag outstanding risks before acting.
-
Clear causal question
Write a concise, actionable causal question that links an intervention to a specific outcome and population. Example: "Does implementing overtime caps in Plant A reduce defect rate among Line 2 workers within 6 months?" Specify the exposure, outcome, population, timing, and estimand (e.g., average treatment effect for treated units).
-
Is an experiment feasible?
Confirm whether a randomized controlled trial (RCT) is possible. If not, document why and move to the quasi-experimental design selection step.
-
Design selection that fits constraints
Choose the best design available and justify why it fits your question and data:
- Difference-in-differences (DiD) — when you have treated and comparison units observed over time.
- Interrupted time series (ITS) — when you have many pre- and post-intervention time points for the treated unit.
- Matching / propensity score — when you can match treated and untreated units on observed covariates.
- Regression discontinuity (RD) — when treatment assignment hinges on a known cutoff.
- Instrumental variables (IV) — when you have a credible instrument that affects treatment but not outcome directly.
Record why alternate designs were rejected and any design-specific caveats.
-
Make your causal model explicit (DAG or theory)
Draw a simple directed acyclic graph (DAG) or narrative model showing assumed causal paths and potential confounders, mediators, and colliders. Use this to identify variables you must measure and paths you must block.
-
Identification assumptions listed and falsifiable implications
For your chosen design, explicitly list the core assumptions (e.g., parallel trends for DiD, no manipulation around the cutoff for RD, instrument exogeneity for IV). For each assumption, record at least one observable implication you can test (placebo test, pre-trend check, density test, etc.).
-
Pre-specify measurement & analysis plan
Create a short analysis plan that specifies primary outcome(s), time windows, covariates, functional forms, clustering and standard error strategy, subgroup analyses, and decision thresholds. Where practical, pre-register this plan or lock it in an internal memo before looking at post-treatment outcomes.
-
Data quality & sample size / power
Verify data completeness, coding consistency, and timing alignment. Run a power or minimum detectable effect calculation (or report expected precision) to judge whether the study can detect practically important effects.
-
Baseline balance and pre-treatment trends
For designs comparing groups or time series, check covariate balance and test for similar pre-treatment trends. Document any imbalances and whether they can be adjusted for or indicate design weakness.
-
Robustness and falsification checks
Run multiple checks to probe threats to validity:
- Pre-trend/placebo tests (lead effects, fake intervention dates).
- Placebo outcomes unlikely to be affected.
- Alternate specifications (functional form, control sets).
- Window/sampling sensitivity (vary time windows or matching calipers).
- RD: McCrary density and manipulation checks around the cutoff.
- IV: test instrument strength and report over-identification / monotonicity concerns if applicable.
-
Sensitivity analysis for unobserved confounding
Quantify how strong an unobserved confounder would need to be to overturn your findings (e.g., Rosenbaum bounds, E-values, or simulated bias scenarios). Report whether conclusions are robust to plausible bias.
-
Heterogeneity and mechanism checks
Test whether effects differ across meaningful subgroups and run basic mediation checks to explore mechanisms (while acknowledging limits of observational mediation analysis).
-
Missing data and attrition
Document patterns of missingness or attrition, test whether missingness correlates with treatment or outcomes, and apply appropriate corrections (multiple imputation, bounding, worst-case scenarios).
-
External validity and scope notes
Describe the population, context, and time period. State limits to generalization (which units, settings, or outcomes the result likely does not apply to).
-
Transparent reporting and reproducibility
Prepare a short reproducible package: analysis code, data dictionary, key datasets (or synthetic/aggregated versions if data are confidential), and a decision-oriented summary of findings, assumptions, and limitations.
-
Decision guidance & risk framing
Translate the causal conclusion into decision terms: expected benefit, uncertainty range, potential harms, cost, and whether trial implementation, phased rollout, or monitoring is recommended. Use risk-based thresholds (e.g., pilot before broad roll-out when uncertainty remains).
-
Sign-off and living action items
Record who reviewed the study, key unresolved risks, monitoring metrics if action proceeds, and a date for re-evaluation. Prefer staged adoption with monitoring when evidence is suggestive but not definitive.
Discussion
Comments and conversation will live here.