Kaveon documentation
Kaveon is a unified data intelligence platform: KaveonDB transactional records, a distributed columnar query engine, a deterministic Data Language Model, and a complete analytics studio — built as one system over data that stays in storage you own.
The idea
Most analytics stacks are assembled. You run a BI tool, point it at someone else’s query engine, bolt an LLM onto the front for natural language, and move data between all three. Every seam is a place where cost, latency, and correctness leak.
Kaveon is built as one system with three pillars. The engine is ours, so query execution is not rented. The language layer is deterministic, so a question resolves through inspectable rules rather than a model’s guess. The studio is the surface over both, not the product boundary.
Three pillars
| Pillar | What it is | Maturity |
|---|---|---|
| Kaveon Engine | A vectorized columnar query engine in Rust, over Arrow. Its own SQL parser, planner, optimizer, and distributed runtime — no DuckDB, Trino, or Spark embedded inside it. | Alpha |
| Kaveon DLM | The Data Language Model: compiles per-dataset context, then answers questions by resolving them against that context. No hosted LLM on this path. | Shipping |
| Kaveon Studio | SQL Lab, semantic datasets, 37 chart types, dashboards with cross-filtering, and administration. | Shipping |
What makes it different
Deterministic language, not a model call. Ask “active users by plan last quarter” and the DLM routes it to a dataset, resolves the entities, and either answers from precomputed context or assembles SQL. The same question yields the same query every time, at no token cost and with no model latency. Generative approaches cover broader phrasing; this one is reproducible and auditable. See DLM · NL→SQL.
Answers without a scan. The DLM precomputes metric totals, per-dimension breakdowns, and low-cardinality dimension pairs when it compiles a dataset. Common questions are served from that context with no database round trip; only uncovered shapes fall through to a live query. The UI always tells you which path answered you. See Freshness.
Your data stays where it is. The direction is the Live Lake Path — Engine reads Parquet and Delta in your storage without a mandatory import step. Today that means the local filesystem and ADLS Gen2; S3 and broader Iceberg qualification remain target work, tracked honestly in Storage & Catalogs.
Where the project actually is
These docs separate what runs today from what is designed but unbuilt, and they say which is which on every page. The two things worth knowing before you read further:
- Studio, the API, and the DLM are the shipping product. They run in production and query your registered SQL sources directly.
- Engine is alpha and its distributed path is qualified through the CLI, HTTP API, and KaveonDB product path. Broader SQL, cloud-format, and direct end-user authentication coverage is still staged. Keep the Engine behind the platform trust boundary.
Choose your path
| If you want to… | Start here |
|---|---|
| Run Kaveon locally and try it | Quickstart |
| Understand the model before committing | Core concepts |
| See how the pieces fit | Architecture |
| Evaluate or develop the query engine | Kaveon Engine |
| Understand deterministic NL→SQL | Data Language Model |
| Integrate programmatically | API reference |
| Deploy and operate it | Deployment · Operations |
| Read the underlying research | Papers & patents |
How these docs are organized
- Getting Started — what Kaveon is, a runnable quickstart, and the core concepts reused everywhere.
- Platform — architecture, the API surface, and the connector matrix.
- Studio — SQL Lab, charts, dashboards, semantic datasets, and data sources.
- Intelligence — the Data Language Model, NL→SQL, and freshness-based routing.
- Engine — the Rust engine manual: architecture, SQL, distributed runtime, storage, and memory.
- CLI — installation, authenticated connections, interactive commands, catalog administration, and automation.
- Deploy & Operate — deployment, auth and RBAC, operations, troubleshooting, upgrades, and releases.
- Research — technical papers, the patent disclosure, and comparisons with other engines.
Every page carries a status badge and a verification date. Desktop pages have an outline on the right; the sidebar filter searches titles, descriptions, keywords, and body text.