Freshness
Kaveon answers a great deal without reading the data, which only works if it can tell when a precomputed answer has stopped being true. Deciding that is itself a metadata read, never a scan.
Two layers precompute, and they decide freshness differently. The Engine’s statistics and cube cells are tied to an exact version of a table and are discarded the moment it moves. The DLM’s context artifact carries a score, but what invalidates it is a measured change, not age.
| Layer | What it precomputes | Freshness rule |
|---|---|---|
| The Engine | Table statistics, cube cells, HyperLogLog sketches | Exact: the recorded source version either is the table’s current one or it is not |
| The DLM | A dataset’s context artifact — totals, per-dimension breakdowns, low-cardinality combinations | Scored, but only a detected change makes it stale |
The Engine: source version identity
Every statistics record and every cube carries the SourceVersion it was computed over — a digest of the table’s location together with its version identity, which is the Delta log version, the Iceberg snapshot id, or the file listing for a directory of Parquet.
pub struct SourceVersion {
/// Location plus the Delta version, Iceberg snapshot, object version
/// or listing digest, hashed together.
pub identity_sha256: String,
pub kind: SourceVersionKind, // DeltaVersion | IcebergSnapshot | Listing
}
impl TableCube {
pub fn is_current_for(&self, identity_sha256: &str) -> bool {
self.source_version.identity_sha256 == identity_sha256
}
}There is no score and no tolerance here. A cube built over version 41 is not used to answer a question about version 42; the planner falls back to a distributed scan and the cube waits for the next ANALYZE. A changed shape declaration rebuilds for the same reason. This is why a tile that used to return in under a second can suddenly take much longer after a table is rewritten: the answer is still correct, it is just no longer precomputed. See Kaveon Engine for what makes a question cube-shaped in the first place.
The DLM: a score, and what actually invalidates it
A dataset’s context artifact gets a score from two factors, and the score is reported — but the decision does not rest on it alone:
score = time_factor(age) * change_factor(rows_changed)
fresh = (not data_modified) or score >= 0.50.0 alongside "fresh": true — that is an artifact built weeks ago over data that has not moved, and it is the correct answer.The change signal depends on where the dataset lives, and this is the part worth understanding:
- An Engine-backed dataset compares the source version its artifact recorded against the table’s current one. The digests match or they do not, so the change is a fact rather than an estimate. A moved version is scored as at least the half fraction, which puts any artifact with age past the threshold — in practice, a version that moved means rebuild.
- A registered external database has no such identity, so drift is inferred from the database’s own modification counter. This path needs a source that exposes those counters and is not available in a deployment without one.
data_modified stays false, so it reports use_context forever and is never rebuilt automatically — even if the table underneath is replaced. An artifact compiled before its dataset was bound to an Engine table is in exactly this position, and the only fix is to recompile it so the binding is recorded. Check the signal field to tell which kind you have.The numbers
| Constant | Value | What it does |
|---|---|---|
BASE_HALF_LIFE_SECONDS | 6 hours | The time factor halves every 6 hours. It lowers confidence in the score; it does not by itself make context stale. |
CHANGE_HALF_FRACTION | 0.05 | Five percent of rows changed halves the change factor. |
| Staleness threshold | 0.5 | With a detected change, the score must reach this to keep using context. |
Usage weighting exists in the scorer — frequently relied-upon elements were meant to decay faster, shortening the effective half-life — but dataset freshness does not pass a usage count, so it is not in effect on this path. It is described here because the code is there, not because it changes an answer.
What rebuilds a stale artifact
| Trigger | Status |
|---|---|
| On ask. A question that routes through the DLM checks freshness after serving, and starts a background rebuild if the artifact is stale. The asker gets their answer immediately; the next asker gets the rebuilt context. | Active |
Pipeline notification. POST /dlm/notify-data-change?dataset_id=… clears the in-memory context and starts a rebuild, so a load that has just finished does not wait to be noticed. | Active |
Manual sweep. POST /dlm/sweep checks every compiled artifact the caller can read and rebuilds the stale ones. Administrator only. | Active |
| The automatic background sweep. A daemon thread was intended to run the same check every 30 minutes. | Not running |
Reading a freshness report
GET /datasets/24/freshness
{
"fresh": true,
"score": 0.0,
"computed_at": "2026-09-12T08:35:02.373996",
"data_modified": false,
"recommendation": "use_context",
"signal": "engine_source_version",
"source_version": { "identity_sha256": "…", "kind": "delta_version", "version": 41 },
"current_source_version": { "identity_sha256": "…", "kind": "delta_version", "version": 41 },
"observed_at_ms": 1760042435000
}| Field | How to read it |
|---|---|
recommendation | use_context, rebuild, or no_context when there is no ready artifact. This is the decision; the score is evidence for it. |
data_modified | Whether a change was actually detected. False with a low score means old but unchanged, which is fresh. |
signal | Which change signal was used. engine_source_version is the exact one. Anything else means the artifact has no Engine binding — see the warning above. |
source_version vs current_source_version | What the artifact was built over, and what the table is now. Equal digests are the reason a precomputed answer is allowed to stand. |
Studio reports the same distinction per answer, so a reader never has to guess whether a number came from precomputed context or from a query that has just run.