Data pipelines & reproducible compute

Patterns and practical steps to build reproducible ETL and containerized compute workflows that preserve provenance and support reliable re-runs.