Pipelines
Turning notebooks and scripts into pipelines that run on a schedule or after each other, following their runs, and promoting them from one environment to the next.
A pipeline is one or more steps that run together — a notebook, a script, a container — started by hand, on a schedule, or when other pipelines finish. It runs as its environment's service identity, not as you, which is what lets it write where people cannot: rows in a production catalog change only through pipelines.
- From a notebook
- Creating a pipeline
- Steps
- Runs
- Versions
- Promoting between environments
- Limits a domain can set
From a notebook
The fastest way in. In a notebook attached to a kernel, run your cells: the Lineage pane in the right-hand sidebar draws the tables each cell read and wrote, as Marquez recorded them. When it looks right, choose Commit as job to make the notebook a new job or the task of an existing one, with its input and output tables recorded for lineage. Group jobs into a pipeline to run them in order.
Creating a pipeline
Open Jobs and create a pipeline. The form has four steps:
| Step | What you choose |
|---|---|
| 1. Pipeline | Its name — how runs, schedules and other pipelines refer to it — and how its steps run: in series (one at a time), in parallel (all at once), or as a custom graph where each step names the steps it waits for. |
| 2. Catalog | A domain and an environment. Together they name the catalog its steps read and write, and the environment's service identity is who they run as. The form shows that identity and its roles. |
| 3. Trigger | Manual, a Schedule (a cron expression), or After other pipelines: each successful run of a pipeline you tick starts this one — once per run, not once all have finished — and a failed run starts nothing. Any pipeline can also be run by hand from its page. |
| 4. Compute | The compute profile its steps run on, each listed with what one worker costs per hour. |
Steps
A pipeline can mix step types. Apache Spark™ runs notebook steps and Spark Submit steps; the other types run whatever their image holds.
| Step type | Runs |
|---|---|
| Notebook | A notebook, top to bottom, on the same image the notebook editor uses — so a notebook that works in the editor works in a pipeline. Needs the pipeline to have an environment. |
| Script | A script. |
| Container | A container image you name. |
| Spark Submit | A Spark application, with its own configuration and main class. |
| Resource | A Kubernetes manifest. |
Each step can also set its order, the steps it depends on (for a custom graph), and whether its driver and executors run on On-demand or cheaper Spot capacity, which can be reclaimed mid-run. Listing a step's Input tables and Output tables puts it in the catalog's lineage.
Runs
Every run is listed under Runs with its status — Pending, Running, Succeeded, Failed or Error. A run's page shows each step, and its logs: the last 500 lines of each container's output.
- Re-run starts the pipeline again from the beginning, as a new run of its current definition — usually what you want after fixing whatever failed.
- Resume continues a failed run from the step that failed, keeping what earlier steps produced.
A promotion's sandbox run is not re-run from here; it is started again from its promotion.
Versions
Each change to a pipeline's definition freezes a new version, and every run records the version it ran. A version never changes afterwards, so what an old run executed can always be seen. Promotion moves a specific version, not whatever the pipeline looks like at the time.
Promoting between environments
A pipeline built in dev reaches prod by promotion: a request that pins one version and passes the checks the path between the two environments sets. Administrators define the paths on the Promotion paths page; each one sets:
- how many sign-offs it needs, from which roles, and whether the author may sign their own;
- whether a sandbox run must pass first, how long it may take, and whether a failure blocks the request or only records the failure;
- whether an approved request is promoted automatically or by someone pressing Promote.
An environment can also require a sandbox run on every path into it, under Settings → Environments.
To promote, search the command palette for Environment Promotions and choose New promotion request. Pick the pipeline and the path; the version is pinned as it is at that moment. Then:
- Sandbox run, if the path needs one. A reviewer starts it — opening a request starts nothing.
- Sign-off, once the sandbox run has reported: the required number of people from the allowed roles, not counting the author where the path says so. A reviewer can also reject, with a reason.
- Promote. The page records who deployed which version, and when.
A pipeline that starts after other pipelines in the source environment arrives in the target as manually started, because those pipelines are not there. The request says so before you promote.
The Review queue lists every request with its path, sandbox result, sign-offs and author, filtered by what is waiting on whom.
Limits a domain can set
An organisation's policy is a floor every domain sits on; a domain can tighten it for its own pipelines, never loosen it. The policy can set the shortest schedule interval, which runtimes pipelines may name, how many pipelines one may trigger, the most CPU, memory and running time a run may use, and whether deploying to production needs a second person. See Settings and administration.