A code-to-column change-impact knowledge graph for AI coding agents.
Six ideas hold this engine together. Each is a mechanism in the code, not a posture: the file that implements it is named beside it.
Every edge in the graph carries one of five grades. They are a lattice, and the whole point is that nothing moves up it.
| Grade | What it means | Example |
|---|---|---|
EXACT |
proved unique by syntax, symbols and constants | a handler calling a mapper method directly |
SOUND_SET |
a conservative candidate set the real target is guaranteed to be inside | the possible implementations behind an interface dispatch |
HEURISTIC |
plausible from a project convention; a recovered or partial binding | a column name derived from a field name with no declared naming strategy |
RUNTIME_ONLY |
not statically decidable; needs runtime evidence | a bean or URL chosen from external configuration |
UNRESOLVED |
the analysis failed, or the shape is unsupported | a parse failure, a missing dependency |
A candidate set of one is never promoted to EXACT, and no refactor may
change that. Narrowing is not proving. The lattice is computed in exactly one place —
src/core/policy.mjs — and the workers emit evidence (what kind of call,
what binding state), never a grade. A property test checks that no evidence
combination classifies above the lattice, and a second checks totality: an
evidence shape the table does not know falls to a conservative HEURISTIC with
a POLICY_GAP diagnostic, never to silence.
A walk is graded by its weakest link: a chain that passes through one
SOUND_SET call is SOUND_SET, however exact the rest of it was.
Query modes pick a floor: strict uses confirmed edges only, conservative
adds candidate calls, heuristic also admits guessed rules. A question that
returns nothing under conservative and something under heuristic has told
you something real about the code.
src/mcp/contract.mjs is the only place a valid response is built. The fields
are stamped with a private Symbol, so a hand-assembled object that merely has
the right shape is rejected by assertContract() before it is serialised
(invariant I-8). A server response that breaks the contract is a 500
contract-violation, not a 200.
basis — the answer’s anchor: project, pack digest, when it was built, and
a freshness verdict, one of current · behind · provisional-overlay ·
unknown. unknown is a real answer and never reads as current. A basis with
no pack digest is refused outright: an answer with nothing to anchor it to is
not an answer.
trust — a computed level, never a literal. Three names exist
(UNCERTIFIED, GOLDEN_FAIL, GOLDEN_PASS) and src/core/trust.mjs is the
only file allowed to write them down (a test greps the tree and fails the build
if the strings reappear elsewhere). The computation is pessimistic in order: no
calibration state at all → UNCERTIFIED; a gate that is not GREEN/BOOTSTRAP →
UNCERTIFIED; fewer than 30 approved golden cases → UNCERTIFIED. Note the
implication the module states plainly — 30 flawless cases give a Wilson lower
bound of 0.8865, so no 95% target can be shown at N=30. Thirty cases are
where a corpus may start being scored, not where it becomes sufficient.
trust.knownGaps also carries the axis declarations (column-axis-degraded,
code-axis-not-shipped, …).
limits — what the engine could not see, in sentences, each scoped. A depth
cap that bit, a ${} substitution it refused to guess at, an axis that was
never built.
truncated — per list: how many were shown, how many there are, in what
order, and where to continue. A truncated list that called itself complete is
the failure this field exists to prevent.
Empty means something. not-shipped (the axis was never built) ·
degraded (built, but without something it needed) · none (looked, found
nothing). Those are three answers, and a 0 results never reads as “safe”.
| the certified base | the working-tree overlay | |
|---|---|---|
| what it describes | a specific commit’s blobs | the bytes on your disk right now |
| determinism | byte-for-byte reproducible | explicitly not; outside every digest |
| cost | seconds to minutes | ~0.2 s on a 900-statement project |
| written down | yes — the pack | nothing: not the pack, not the fact cache |
| how it reads | current / behind |
provisional-overlay |
The base is the pack: same commit + same catalog snapshot + same engine and
profile → the same digest, on any machine. Build time, hostname and absolute
paths are carried in meta and are outside the digest.
The overlay re-parses your dirty files on every call and splices them over
the cached shards of everything else (src/core/overlay.mjs). A node the base
never had is marked provisional — a marker beside the grade, not a grade. The
overlay may over-approximate; it may not omit, and that is a test, not a
promise: test/overlay_integration.test.mjs compares it against a full
re-analysis of the same bytes. Commit your edit and the overlay is discarded,
answering behind and naming cascade analyze as the cure — never a quiet fall
back to the pre-edit answer.
The second speed also covers reruns: a run after a small edit reuses the
content-addressed shards of everything that did not change, and the result must
equal a cold run of the same state byte for byte (invariant I-9,
test/incremental.test.mjs, which mutates a random subset of files each round
with a seeded PRNG). Four things it refuses to do quietly: unknown is never
“nothing changed”; a damaged shard is never loaded (it is reported
SHARD_UNUSABLE and that unit alone is recomputed); a dirty tree is stated in
meta.base.dirty; and the pack records the mode, the base commit and the
reparsed/reused counts.
Every lane input is optional, and every lane can be switched off by name
(--no-ddl, --no-mappers, --no-java, --no-web, --no-openapi). The run
still produces a valid pack,
and the pack records meta.axes: per axis (catalog statements column
code jpa web screen) one of shipped / degraded / not-shipped
and why.
That declaration then rides on every answer: the axis name lands in
trust.knownGaps, the reason in limits, and an empty list says not-shipped
rather than none.
Nothing is dropped to make the answer look clean. Analysing a project with
--no-ddl declares column: degraded, and the numbers behind it are real —
what cannot be attributed without a catalog is recorded as unresolved, not
discarded.
The screen axis is the far end of that round trip: the router’s own
declarations, composed into the paths a user is really on, each joined by a
RENDERS edge to the frontend functions of the component it mounts. RENDERS is
EXACT onto the file the route DECLARES and SOUND_SET onto a file that
component imports, because which of an imported component’s functions really runs
is a run-time question. It is shipped only when the profile turned the axis on,
at least one screen reaches a function, and nothing about the reading was a
guess; a router that fetches its menu from the server at run time makes it
degraded, because the screens in the source are then not the screens the
product has. The gate itself, screenAxis.enabled, has three states: true and
false are your word and are obeyed whatever the run reads, and null (the
default, and what cascade init leaves when it found no router package in the
tree) means “decide it from what this run READS” — which is what lets a backend
analysed with --web-src ../front/src build screens with nothing configured. Between that component and the route it hits sits one more hop,
symbol --CALLS--> symbol, graded the same way: EXACT for an import or a name
inside one file, SOUND_SET through an export * barrel or when the function
was never called here at all but handed to some other call AS A VALUE, and
HEURISTIC through an assumed alias. See
the web lane setup page.
A browser recording (--har) is a different KIND of fact and is graded as one:
RUNTIME_ONLY sits below every mode’s floor, so a recorded screen → route call
is shown (observed: true) and never walked, and it never raises the
grade of the static edge beside it. A recording proves a request happened once;
it proves nothing about what the code can do.
degraded is not only about a missing input. The web axis is shipped only
when nothing about the frontend had to be guessed: a URL prefix this engine
worked out by counting matches, or a path alias it assumed because the project
declares none, makes the axis degraded and the reason names what to declare to
fix it.
It can also mean the axis stops earlier than it looks. A pack whose routes
came from an OpenAPI document and not from source has endpoints and nothing under
them, so the code axis is degraded with the reason “endpoints come from an
OpenAPI document, not from source: the routes exist, but nothing below them is
walked, so a frontend call reaches an endpoint and stops there” — and a question
that would need the chain below a route answers not-shipped, never an empty
list that would read as “we looked”.
cascade estimate answers the same question before you analyze: what this
tree will ship, degrade, or not ship.
Each project keeps a sealed baseline in .cascade/calibration/
(src/core/calibration.mjs). Every analyze is compared against it, and the
gate first asks why this run differs:
| Mode | When |
|---|---|
NO_SEAL |
no baseline yet — this run becomes it (BOOTSTRAP) |
NO_CHANGE |
same engine, same commit pins |
ENGINE_MOVED |
the engine fingerprint moved — an upgrade |
REPIN |
the analyzed commits moved |
BOTH_MOVED |
both |
ENGINE_MOVED is the strict one: “the analyzer changed” is exactly when a loss
must not slip through, so a drop is treated as a regression and blocked. The
gate is comparative, not absolute, and it reports improvements too — a ratio
whose numerator rose is a wider lane, not a smaller answer, and the finding
says so in those words.
A RED run is never silently discarded: its pack is written to
<packDir>-rejected/, the certified pack is left exactly where it was, and the
command exits 3. The one override is --accept-baseline, which re-seals the
baseline from that run — a human decision, by design.
Alongside it, cascade golden keeps the project’s own labelled corpus: the tool
proposes cases, a human approves them (--ids …, or an explicit --all),
a hash decides which are held out, and check scores the approved ones through
the shipped MCP tools. The tool never approves itself, which is why the trust
level above can mean anything.
UNRESOLVED —
never rounded up.${} substitution, an unresolvable
include) is recorded as a diagnostic, not guessed and not dropped.create table item (…) in the DDL and FROM ITEM I in a mapper name the
same table in HSQLDB, in Oracle, and in a MySQL server with
lower_case_table_names=1. Matching them by exact string would put two tables in
the graph and split every fact between them. Matching them by folding always
would merge two genuinely different tables in a database that does distinguish
them. So the rule is neither: it is declared per dialect, from what each
database’s own manual says, and it is one small table
(src/core/identifier_case.mjs, mirrored for the worker in
adapters/sql/identifier_case.py, cross-checked by a test that runs both).
fold-lower; Oracle, HSQLDB and H2 →
fold-upper; a dialect this engine does not know → exact, folding nothing."Item", `Item`) is exact in every one of them,
so it is matched only against the spelling the catalog declares.sqlIdentifierCase — fold-lower | fold-upper | exact |
null (the dialect’s own rule) — overrides the table, and every run states
which rule it used and where the rule came from.The fold produces a matching key, never a new name. A table and a column
keep the spelling the catalog gave them, so pms_product stays pms_product in
the ERD and in every answer — which also means the fold direction cannot
change the facts, only what matches. Two catalog names that fold to one key are
reported as a folded_identifier_collision and the first declaration keeps the
key; both tables stay in the pack as declared. And a table the catalog does not
have under the folded comparison is still unresolved — folding makes the
comparison correct, not lenient.
A tool argument is folded the same way. The pack records the rule it was
built under (pack.meta.identifierCase), so table_usage ORDERS — the spelling
you read off your own SQL — answers about the catalog’s orders, and the answer
carries one limits line saying so. The same goes for column_impact,
endpoint_impact, flow, neighborhood and erd’s focus table. A name that
matches nothing keeps its unknown-table / unknown-column code and names the
closest ones it does have; a name that folds onto two declared names is
ambiguous and is not resolved to either. Under exact nothing folds, and a pack
built before the engine recorded the rule declares none — such a pack matches
arguments exactly, as it always did.