Quickstart
A five-minute local walkthrough for Studio, the API, and the Engine. Use Install and deploy for VM, AKS, or Vercel hosting; use Deployment topology and Operations for production operation.
Prerequisites
| Requirement | Why |
|---|---|
| Docker with Compose v2 | Runs the whole stack. docker compose version should print v2.x. |
| ~4 GB free RAM | Five containers: Studio, the API, one Engine coordinator and two Engine workers. There is no database to install. |
| Ports 3000, 8080, 8081, 5433 | All bound to 127.0.0.1 only. |
| Git | To clone the repository. |
You do not need an OAuth application, a cloud account, an API key, or an existing warehouse to complete this page.
1 · Start the stack
git clone https://github.com/PruthviProdduturi/Kaveon.git
cd Kaveon
docker compose up -d --buildThe first build compiles the Rust Engine and takes several minutes; subsequent starts are fast. Watch the containers become healthy:
docker compose psNAME STATUS
kaveon-engine-coordinator Up (healthy)
kaveon-engine-worker-1 Up
kaveon-engine-worker-2 Up
kaveon-api Up (healthy)
kaveon-studio Up (healthy).env to use personal or work sign-in; without provider credentials, local mode uses a development Admin. That mode is for localhost only — see Auth & RBAC before exposing Kaveon to a network.2 · Verify
Check each tier independently before opening the UI:
curl -s localhost:8080/api/health # platform API
curl -s localhost:8081/health # Engine coordinatorThen open http://localhost:3000. If Microsoft is configured, sign in with your Microsoft account; otherwise Studio uses the local development Admin.
3 · Run your first read query
Open Catalog to choose a registered catalog and schema, then open SQL Lab. The local showcase uses read-only Parquet and Delta tables; Kaveon catalog DDL registers existing lake data and does not create arbitrary PostgreSQL-style user tables.
SELECT *
FROM OpenSource.public.nyc_taxi_borough
LIMIT 25;This table is part of the OpenSource showcase manifest. For another lake, replace it with a table shown by SHOW TABLES. Press Ctrl/Cmd + Enter to run the query.
Results are cached by SHA of the query text, so re-running is instant. Full editor reference: SQL Lab.
4 · Define a semantic dataset
A dataset names your dimensions and metrics once so charts and questions can be built without rewriting SQL. Go to Datasets → New Dataset, choose the table you just queried, and mark:
- Dimensions — choose categorical columns such as borough, country, or status
- Metrics — choose a numeric column and aggregation such as
SUMorCOUNT - Time column — choose a date or timestamp column when the table has one
This is the input the DLM compiles against. Details: Semantic Datasets.
5 · Compile the DLM context
On the dataset page, click Generate. Kaveon scans the table once and precomputes each metric’s grand total, its breakdown per dimension, and low-cardinality dimension pairs. This is what lets common questions be answered with no database round trip.
6 · Ask a question
On the home page, type:
revenue by regionKaveon routes the question to the dataset, matches the metric, groups by the dimension, and renders the result. Look at the badge above the answer:
| Badge | Meaning |
|---|---|
| From context · no DB scan | Served from precomputed context. No query ran. |
| From sketch · ≈ estimate | Approximate distinct count from a HyperLogLog sketch. |
| Live query · Xs | The shape was not precomputed, so SQL was assembled and run. |
No hosted model is involved on any of those paths. How the routing decides: DLM · NL→SQL and Freshness.
7 · Query a Parquet file with the Engine
Studio sends analytical statements through the Engine. For direct Engine access, point the CLI at a directory of Parquet or Delta files and restart:
KAVEON_DATA_PATH=/path/to/parquet docker compose up -dInstall the CLI and connect it to the running coordinator:
curl -fsSL https://raw.githubusercontent.com/PruthviProdduturi/Kaveon/dev/scripts/install.sh | bash
kaveon --server http://localhost:8081Tables are auto-discovered from .parquet files in that directory, under the configured local catalog:
SHOW CATALOGS;
SHOW TABLES;
SELECT region, count(*) FROM orders GROUP BY region ORDER BY 2 DESC LIMIT 5;Or run a single statement without entering the shell:
kaveon --server http://localhost:8081 -e "SELECT count(*) FROM orders"8 · Shut down
docker compose down # stop, keep everything
docker compose down -v # stop and delete the volumes — see below-v deletes the catalog-data volume, and that volume holds the Engine’s catalog: every catalog, schema and table definition you registered, the planner’s statistics, and every cube. Table data in your mounted directory is untouched, but the map to it is gone and the cubes have to be rebuilt, which is not quick. Use plain docker compose down unless you mean to start over.Where to go next
| To… | Read |
|---|---|
| Understand datasets, charts, and questions properly | Core concepts |
| Connect a real warehouse instead of the local one | Data Sources · Connector matrix |
| Build charts and dashboards | Chart Builder · Dashboards |
| Go deeper on the Engine | Kaveon Engine |
| Deploy beyond localhost | Install & deploy · Deployment topology |