Name the code a run ran, and let an interrupted sync finish

Three faults with one root: the stored body of a code-defined node is an
import shim, and nothing that mattered was ever read from the code itself.

- The run stamp could not identify what ran. The shim imports whatever is on
  disk when the worker starts, and an uncommitted tree stamps <commit>-dirty
  for every run it ever produces. Run.code_digest hashes the repository's .py
  files, memoized on their stat state, and it is read again when the run is
  actually claimed -- so a sweep queued for hours records the code each of its
  runs executed, not the code that was there when it was submitted.
- The stage cache adopted code that was too new. The fingerprint hashed the
  shim, which is invariant under any edit to the imported function or anything
  it calls into, so a re-run was served from cache and answered without the
  outputs the edit added. It now carries the repo digest and the node's
  declared ports. Every fingerprint changes once, which invalidates the
  existing cache; a canvas flow has no repository and keys as before.
- An interrupted sync looked like a hand-edited canvas. The engine answers a
  new-node template for a node with no stored body, and the template carries
  no marker, so the drift check read "somebody edited this" and demanded
  --force -- for the one state that re-running the sync is the fix for.
  NodeSource.missing states the fact, and sync skips those and reuses the
  bodies it read instead of asking for each one twice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-26 21:27:10 +02:00
co-authored by Claude Opus 5
parent 1f7c6646f1
commit 4a38c6ed31
15 changed files with 422 additions and 28 deletions
+26 -1
View File
@@ -150,6 +150,10 @@ class RunContext:
"""
run_id: str
#: What the code-defined flow's repository hashed to when this run started.
#: Part of every node's fingerprint, so editing a function the node calls
#: into invalidates the cache the way editing the node itself does.
code_digest: str = ""
@dataclass(frozen=True, slots=True)
@@ -919,9 +923,30 @@ class FlowController:
# The raw params, not the resolved ones: a secret's value
# must not end up in a key, and its name is what changes
# when the node is reconfigured anyway.
#
# `source` is the stored body, which for a code-defined
# flow is a shim that imports the real function — it does
# not move when that function does, and it says nothing
# about what the function calls into. `code` is what
# covers both: the digest of the repository the shim
# imports from. The ports are here because declaring a new
# one changes what the node answers with, which a hit
# would otherwise restore without.
node.fingerprint = hashlib.sha256(
json.dumps(
{"source": code, "params": node_def.params},
{
"source": code,
"params": node_def.params,
"requires": [
spec.model_dump(mode="json")
for spec in node_def.requires
],
"provides": [
spec.model_dump(mode="json")
for spec in node_def.provides
],
"code": run.code_digest,
},
sort_keys=True,
separators=(",", ":"),
).encode()