Name the code a run ran, and let an interrupted sync finish
Three faults with one root: the stored body of a code-defined node is an import shim, and nothing that mattered was ever read from the code itself. - The run stamp could not identify what ran. The shim imports whatever is on disk when the worker starts, and an uncommitted tree stamps <commit>-dirty for every run it ever produces. Run.code_digest hashes the repository's .py files, memoized on their stat state, and it is read again when the run is actually claimed -- so a sweep queued for hours records the code each of its runs executed, not the code that was there when it was submitted. - The stage cache adopted code that was too new. The fingerprint hashed the shim, which is invariant under any edit to the imported function or anything it calls into, so a re-run was served from cache and answered without the outputs the edit added. It now carries the repo digest and the node's declared ports. Every fingerprint changes once, which invalidates the existing cache; a canvas flow has no repository and keys as before. - An interrupted sync looked like a hand-edited canvas. The engine answers a new-node template for a node with no stored body, and the template carries no marker, so the drift check read "somebody edited this" and demanded --force -- for the one state that re-running the sync is the fix for. NodeSource.missing states the fact, and sync skips those and reuses the bodies it read instead of asking for each one twice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+16
-6
@@ -177,18 +177,28 @@ downloadable at `GET /api/v1/artifacts/{digest}`.
|
||||
## Stage caching
|
||||
|
||||
A run mostly does not redo what an earlier one already did. Before a node
|
||||
executes it is fingerprinted — a sha256 over its source, its settings and the
|
||||
values it is about to read — and if some earlier run of that same fingerprint
|
||||
finished, what that one returned is restored into this run's state and the node
|
||||
is skipped. It is recorded with the status `cached` and a duration of zero, and
|
||||
its artifacts are listed on the new run as well, so they stay downloadable from
|
||||
either.
|
||||
executes it is fingerprinted — a sha256 over its source, its settings, the ports
|
||||
it declares and the values it is about to read — and if some earlier run of that
|
||||
same fingerprint finished, what that one returned is restored into this run's
|
||||
state and the node is skipped. It is recorded with the status `cached` and a
|
||||
duration of zero, and its artifacts are listed on the new run as well, so they
|
||||
stay downloadable from either.
|
||||
|
||||
The settings go into the key raw, so a secret contributes its `{"$secret":
|
||||
name}` reference and never its value. An artifact input counts as its content
|
||||
digest: the same bytes under a different filename are the same input. Each node
|
||||
on a run carries the `cache_key` it was looked up by.
|
||||
|
||||
For a [code-defined flow](../getting-started/data-science.md), "its source" is
|
||||
the generated shim, which imports the real function and does not change when
|
||||
that function does. So the key carries one thing more: a digest of every
|
||||
`.py` file under the repository the flow was declared in, read when the run
|
||||
starts. Editing a helper three calls down from the node invalidates it, which
|
||||
is the point — the alternative is a re-run answering with the previous code's
|
||||
numbers. It is deliberately blunt: any edit anywhere in the repository re-runs
|
||||
every node of its flows. An engine that cannot see the repository — a worker on
|
||||
another machine — records no digest and keys as it did before.
|
||||
|
||||
The run history *is* the cache; there is no second store. A node's returned
|
||||
outputs are kept on its run record as canonical JSON, up to 256000 characters —
|
||||
a node returning more than that is simply not cacheable that run. An entry
|
||||
|
||||
Reference in New Issue
Block a user