Where data takes root.
Build pipelines in notebooks, orchestrate workloads on the web, browse a unified catalog across AWS, GCP and Azure. Sown Data is the lakehouse and warehouse that grows with you, not against you.
Everything your data team needs, nothing they don't.
Sown unifies the four pieces every modern data team rebuilds from scratch: a notebook IDE, a pipeline orchestrator, a queryable warehouse, and a catalog. One control plane. Any cloud.
Polyglot notebooks
SQL, Python and Scala in the same notebook. Compile straight to a production pipeline with one click.
- DuckDB, Apache Spark™ and Polars engines
- Live collaboration, comments, branches
- Git-native, every cell is reviewable
Web orchestration
Schedule, retry, branch and backfill from the browser. SLAs, alerts, lineage-aware triggers, out of the box.
- Cron, event and webhook triggers
- Visual DAG with live run status
- PagerDuty, Slack, OpsGenie alerts
Unified catalog
Browse every table, view and stream across every cloud. Search, schemas, lineage and quality, all queryable from one place.
- Cross-cloud column-level lineage
- Faceted search with semantic tags
- Certified assets & PII tagging
Open warehouse
Iceberg-native. Bring your own storage. Query through Snowflake, BigQuery, Synapse, Trino™, or all of them.
- Apache Iceberg + Delta + Parquet
- BYO S3 / GCS / ADLS, your buckets
- No vendor lock-in. Ever.
Data quality & observability
Freshness, drift, schema and custom checks, declared next to the data and alerted before downstream notices.
- SLA tracking + auto-incident creation
- Drift detection on every refresh
- Quality scores per asset
Enterprise-ready
SSO, SCIM, audit, region pinning and customer-managed keys. The boring parts that make security teams smile.
- Compliance requirements reviewed case-by-case
- Okta / Azure AD / Google SSO
- VPC-peered private deployments
Write it once. Ship it as a pipeline.
SQL and Python cells in one notebook. sd.publish() turns any cell's output into a versioned, scheduled, governed asset. No rewrite, no copy-paste.
- Reproducible. Every run is git-pinned with a content hash.
- Testable. Inline assertions become DQ checks in production.
- Observable. Runtime, rows scanned and cost, visible per cell.
SELECT order_id, customer_id,
date_trunc('hour', placed_at) AS hour,
total_cents / 100.0 AS total_usd
FROM raw.shopify_orders_raw
WHERE placed_at >= now() - INTERVAL 2 HOUR;
| order_id | customer_id | hour | total_usd |
|---|---|---|---|
| o_4f3a8e91 | c_8821 | 2026-05-07 12:00 | 147.50 |
| o_4f3a8e92 | c_3104 | 2026-05-07 12:00 | 89.99 |
| o_4f3a8e93 | c_9182 | 2026-05-07 12:00 | 1,204.00 |
enriched = sd.ref("src_orders") \
.join(sd.ref("cohort_dimensions"), on="customer_id")
sd.publish(enriched, name="orders_hourly") # → schedules pipeline
Find any table, in any cloud, in seconds.
Search 1,200+ assets across Snowflake, BigQuery, Synapse, Iceberg, Kafka. See schema, owners, lineage, freshness and quality at a glance, and query right from the browser.
- Cross-cloud lineage. Trace a column from S3 raw to a Tableau dashboard in BigQuery.
- Semantic search. Type "monthly revenue" and find
fct_revenue_dailyranked first. - Certified assets. Trust the data your team trusts.
Your platform knows your data. Ask it anything.
An enterprise AI assistant is embedded across your workflows — not bolted on. Generate code, tune queries, troubleshoot failed jobs, and discover certified data assets using natural language.
From natural language to production PySpark.
Generate PySpark transformations, SQL queries, and Argo workflow specs from a prompt. Paste a slow query and receive refactored SQL, partition pruning optimizations, and explain-plan diagnostics. Ask about your data — get certified table references back.
- Code & SQL generation. Describe the transformation; get production-ready PySpark or SQL.
- Query tuning. Automated refactoring with partition pruning and broadcast hints.
- Job diagnostics. Analyzes driver logs, executor errors, and system metrics to pinpoint root causes with step-by-step fixes.
- Natural language discovery. "Which dataset contains active customer churn rates for Southern Africa?" returns a certified table reference instantly.
orders_hourly pipeline fail at stage 4?Fix: Increase Core NVMe from 500 GB → 1 TB in your compute profile, or enable
spark.shuffle.compress=true to reduce shuffle footprint by ~40%.
orders join query./*+ BROADCAST(dim_customer) */ hint (2.1 MB table), pushed WHERE region = 'ZA' before the join, and replaced SELECT * with explicit column projection. Estimated scan reduction: 74%.
Plant the seed today.
We will walk through your workloads and show what runs where.