Documentation
How the platform actually fits together, written against the objects you will see in the product: workspaces, runtimes, jobs and the catalog.
Quickstart
Sign in and start a kernel on a runtime. A kernel is compute, kept separate from any one notebook, so closing a notebook does not stop the kernel underneath it.
- Start a kernel and pick a runtime version for it
- Attach one or more notebooks to that kernel
- Run cells on the kernel, powered by Apache Spark™
- Turn a notebook into a scheduled pipeline once it does what you want
Getting your tenant live
The control plane is SaaS; the compute and storage a workspace or warehouse touches run in your own account. Standing that up is not a self-service installer today — our team provisions the network boundaries and compute ceilings with you before handing the domain over, so get in touch to start that.
Runtime versions
A runtime bundles a Spark version with the Iceberg, Hudi and OpenLineage builds known to work with it, so picking a runtime is one decision instead of four. See Kernels for what each one holds.
- 0.1 — Spark 3.5.0, Scala 2.12 — supported, default
- 0.2 — Spark 3.5.0, Scala 2.13 — preview
- 0.3 — Spark 3.5.0, Scala 2.12 — preview
- 0.4 — Spark 3.5.0, Scala 2.12 — preview, applies the same column masks and row filters as the SQL warehouses
- 1.0 — Spark 4.0.1, Scala 2.13 — preview, first Spark 4 runtime
Pipelines and triggers
A pipeline runs notebooks, scripts or containers on a compute profile, started by a trigger. Its steps run in series, in parallel, or as a custom graph with per-step dependencies. See Pipelines.
- Manual — run it on demand
- Schedule — a cron expression
- After other pipelines — runs after each successful run of the pipelines it follows
User guides
Step-by-step help with each part of the product.
- The notebook editor — tabs, the toolbar, choosing compute and resizing the sidebars.
- Kernels — starting compute for Python and Scala, runtimes, and which catalog a kernel uses.
- SQL warehouses — creating, sizing and scaling the compute that answers SQL.
- SQL and dashboards — querying in SQL notebooks, charts, and dashboards.
- Pipelines — steps, triggers, runs, versions, and promoting between environments.
- Sandboxes — working against real tables without changing them, in notebooks and promotions.
- The catalog — finding tables, lineage, permissions, quality rules and maintenance.
- Governance — profiling, classifying personal data, review and watches.
- Sharing data — data products between domains, and the marketplace between organisations.
- Databases — Postgres schemas and Cassandra keyspaces, queried without a password.
- AI assist — asking about a notebook, and setting up your organisation's plan.
- Settings and administration — roles, domains, policy, environments, billing and notifications.
See it against your own data
We will walk through your workloads and show what runs where.