Core concepts
Six ideas explain how Kaveon fits together. Once these click, the rest of the documentation is detail.
1 · One authority, two data planes
Kaveon separates durable product state from customer lake data while keeping both under one governed platform. KaveonDB is the transactional authority for datasets, charts, dashboards, saved queries, DLM definitions, audit records and other product metadata. The analytical plane reads Parquet, Delta and supported Iceberg snapshots in place; it does not require copying the source into a proprietary warehouse.
| KaveonDB transactional plane | Distributed analytical plane | |
|---|---|---|
| Holds | Versioned product records, ownership, revisions, audit and transaction state | Customer tables and immutable source snapshots in object storage |
| Execution | Validate → revision check → commit → audit → recover | Plan → split → execute on workers → Arrow exchange → merge result |
| Storage | Durable KaveonDB product store, local or ADLS-backed | ADLS Gen2/local Parquet and Delta; registered catalog definitions |
| Consistency | Immutable revisions, compare-and-set heads and owner-scoped reads | Version-pinned snapshots, exact statistics and fail-closed pruning |
This is why a dashboard update and a billion-row aggregate do not compete for the same execution path: transactional writes stay bounded and revisioned, while analytical work fans out across workers and returns columnar results.
2 · The content chain
Each layer is reusable by the next, and each one exists so you stop repeating yourself:
data source ──▶ dataset ──▶ chart ──▶ dashboard
connection meaning one view many views
+ filters- Data source — a connection to a database (docs).
- Dataset — a semantic layer over tables: which columns are dimensions, which are metrics, which is time (docs).
- Chart — a visualization bound to a dataset (docs).
- Dashboard — a canvas of charts sharing filters (docs).
3 · A dataset encodes intent once
This is the concept that pays for itself. A dataset says what your columns mean, so nothing downstream has to restate it. Define it once:
dataset orders
table public.orders
dimensions region, plan
metrics Revenue = SUM(total)
Orders = COUNT(*)
time orderedNow “Revenue by region” is fully specified — as a chart, as a dashboard tile, or as a question typed in English. Kaveon assembles the aggregation, grouping, joins, and time filtering for you, in the dialect of the source it is talking to:
SELECT region, SUM(total) AS "Revenue"
FROM public.orders
GROUP BY region
ORDER BY "Revenue" DESC4 · Transactional and distributed execution
Kaveon has two cooperating execution paths. The transactional path protects product state; the distributed path executes analytical SQL. The deterministic DLM can answer a third way — from a compiled context artifact — when the requested shape is already materialized.
| Transactional path | Distributed analytical path | Context path | |
|---|---|---|---|
| Runs | KaveonDB coordinator and transaction API | Rust coordinator plus distributed workers | DLM context service and compiled artifacts |
| Work | Product records, permissions, revisions and audit | Scans, joins, aggregates, windows, sorting and exchange | Precomputed totals, dimensions, pairs and sketches |
| Storage | KaveonDB product records | Cataloged Parquet, Delta and supported Iceberg | Versioned context in governed product storage |
| Entry point | Studio/API product operations | Studio SQL Lab, kaveon CLI, or Engine HTTP API | Studio Chat, DLM API, or CLI .ask |
5 · Catalogs — how the Engine sees storage
Where the platform path has data sources, the Engine has catalogs. A catalog points at storage and the tables inside it are addressed as catalog.schema.table. Tables in a local catalog are discovered from the .parquet files present:
# ~/.kaveon/catalogs/warehouse.toml — the filename becomes the catalog name
type = "local"
base_path = "/data/warehouse"SHOW CATALOGS;
USE warehouse.default;
SELECT count(*) FROM orders;The direction here is the Live Lake Path: read data where it already lives, with no mandatory import. Local filesystem and ADLS Gen2 are qualified paths; supported Iceberg snapshots are available through the Engine catalog, while S3 remains a staged connector target. Details live inStorage & Catalogs.
6 · Ask, don’t query
Questions are resolved by the DLM — a compiled per-dataset context artifact — not by a hosted language model. Compiling a dataset precomputes each metric’s total, its breakdown per dimension, and low-cardinality dimension pairs, so common questions are answered without touching the database at all. Kaveon always tells you which path answered:
| Badge | Meaning | Cost |
|---|---|---|
| From context | Served from precomputed context | No database scan |
| From sketch · ≈ | Approximate distinct count from a HyperLogLog sketch | No database scan |
| Live query · Xs | Shape was not precomputed; SQL was assembled and run | One source query |
Because resolution is rule-based, the same question produces the same SQL every time — reproducible, inspectable, and free of token cost. See DLM · NL→SQL for how routing decides, and Freshness for when context is preferred over a live read.
Access: two independent axes
Roles gate what you can do: Viewer → Analyst → Editor → Admin. Visibility gates who can see a given object: private, internal, published. They compose — an Editor still cannot read someone else’s private dashboard.
The API defines all four roles, but OAuth sign-in resolves a user to just two of them: Admin if their email is listed in AUTH_ADMIN_EMAILS, otherwise Viewer. Full model in Auth & RBAC.