Kernels
Starting the compute your notebooks run Python and Scala on, choosing its runtime, size and catalog, and what happens when it stops.
A kernel is one pod running one compute session, powered by Apache Spark™. Notebooks attached to the same kernel share variables and run one cell at a time. A kernel is separate from any notebook: closing a notebook leaves its kernel running, and a kernel stops by itself once it has been idle for a while.
- Starting a kernel
- Runtimes
- Which catalog it uses
- Stopping, restarting, terminating
- The kernel page
- SQL from a Python notebook
Starting a kernel
Open Kernels under Compute and choose New kernel.
| Field | What it does |
|---|---|
| Name | How you and the notebooks you attach refer to it. |
| Runtime | Pins the Spark version and every library. See Runtimes. |
| Size | Small — 2 GB RAM, 1 CPU. Medium — 4 GB RAM, 2 CPUs. Large — 8 GB RAM, 4 CPUs. Your organisation's plan may cap the largest size. |
| Compute profile | What the kernel's running time is charged at. The form shows the hourly cost for the size you picked: a kernel is charged for every minute it is up, idle or not, until it stops — so auto-stop is also what keeps an idle kernel cheap. Everyone who can start a kernel can use the organisation's default profile; others appear when an admin grants them. The kernel page shows what it has cost this month. |
| Domain and Environment | Together they name the catalog the kernel reads and writes. Choose both, or neither. See Which catalog it uses. |
| Auto-stop after (minutes idle) | Between 10 and 480; 60 by default. The kernel shows Idle — stopping soon before it stops. |
A kernel takes a couple of minutes to start; the page updates as it does. Your organisation limits how many kernels can run at once — the Kernels page shows how many are running and the limit. At the limit, New kernel is disabled until you stop one.
Starting kernels is open to owners, admins and members. Viewers can read notebooks but not run them.
Runtimes
A runtime bundles a Spark version with the Iceberg, Hudi and OpenLineage builds known to work with it, so choosing one is one decision instead of four. Every runtime ships Python 3.11 and Java 17. A published runtime never changes; a new version is a new runtime.
| Runtime | Spark | Scala | Status |
|---|---|---|---|
| 0.1 | 3.5.0 | 2.12 | Supported — the default |
| 0.2 | 3.5.0 | 2.13 | Preview |
| 0.3 | 3.5.0 | 2.12 | Preview |
| 0.4 | 3.5.0 | 2.12 | Preview — applies the same column masks and row filters as the SQL warehouses |
| 1.0 | 4.0.1 | 2.13 | Preview — the first Spark 4 runtime |
A domain's policy can limit which runtimes its pipelines may name. The kernel page lists the exact libraries and versions in its runtime.
Which catalog it uses
A domain and an environment together name one catalog — sales in dev is sales_dev. A table written without a catalog prefix lands there, and an unqualified table name is read from there.
If you choose neither, the kernel uses the catalog its runtime image defaults to. That is a real catalog holding real tables: if another catalog has a table of the same name, the kernel reads that one, not the one you may be browsing. The notebook toolbar marks this with a dashed outline around the catalog name.
The catalog is written into the kernel before it starts, so it cannot be changed on a running kernel. Start a new kernel to use a different one.
Stopping, restarting, terminating
| Action | What happens |
|---|---|
| Stop | Releases the compute. Attached notebooks stay attached. |
| Restart | A full stop and fresh start. Every variable in attached notebooks is gone; the kernel starts empty. |
| Terminate | Stops the kernel and detaches every notebook. You confirm by typing the kernel's name. |
None of these touches the files on your volume. Stopped kernels are listed under Terminated with the reason they stopped — idle, stopped by someone, or failed — and count against nothing. Start a new kernel to run their notebooks again.
The kernel page
- Overview — the notebooks attached to it, with Detach, and the catalog it uses.
- Resources — the runtime's versions, the exact image, its libraries, and the resolved Spark configuration (read-only).
- Spark — the live application behind the kernel.
- Events — when it started, stopped, and why.
- Logs — the last 500 lines of its output.
SQL from a Python notebook
Two cell commands run SQL from a kernel notebook. Put the command on the first line of the cell and the query below it.
| Command | Runs |
|---|---|
%%sql | Spark SQL, on the kernel itself, against its catalog. |
%%trino | SQL through your organisation's running SQL warehouse — the way to reach Postgres, Cassandra and anything else that is not a table in the catalog. Needs a warehouse to be running. |
Both take the same options: -o name assigns the result to a DataFrame called name, -n N previews N rows, and -q runs without showing the result. %%sql also takes -e to show the query plan instead of running it.
%%trino signs in with a token issued when the kernel started, which lasts several hours. If a long session starts failing with an authentication error, restart the kernel.