Stage caching for batch runs, and an engine that lives in the command
Docs / docs (push) Successful in 19s
Playwright Tests / test-playwright (1, 2) (push) Failing after 1m5s
Playwright Tests / test-playwright (2, 2) (push) Failing after 20s
pre-commit / pre-commit (push) Failing after 2m33s
Test Backend / test-backend (push) Successful in 2m7s
Compose Smoke Test / test-compose (push) Failing after 20s
Playwright Tests / merge-reports (push) Failing after 1m3s
Publish / publish (push) Failing after 12s
Docs / docs (push) Successful in 19s
Playwright Tests / test-playwright (1, 2) (push) Failing after 1m5s
Playwright Tests / test-playwright (2, 2) (push) Failing after 20s
pre-commit / pre-commit (push) Failing after 2m33s
Test Backend / test-backend (push) Successful in 2m7s
Compose Smoke Test / test-compose (push) Failing after 20s
Playwright Tests / merge-reports (push) Failing after 1m3s
Publish / publish (push) Failing after 12s
A code node in a batch run is now fingerprinted by its source, its raw settings and the values it reads — an artifact input counting as its digest, which is what the content addressing was always for. A run that finds the key restores what the earlier one returned and skips the node, recorded as `cached`. The run history is the cache: `run_node.outputs` beside the `cache_key` the schema already had, no second store. On for code nodes, never for the built-in and connector types that have side effects; off per node with `@node(cache=False)` and per run with `--no-cache`. Emissions are not replayed on a hit, so a cached training node returns its result without redrawing its curve. Recorded in NOTEPAD.md with the two other deliberate limits. `fluksio run --local` boots the real app in the command's own process and drives it through its ASGI interface behind the ordinary client, so a run no longer needs a `serve` terminal beside it — same data directory, same history, and the cache carries between the two. It always waits, because the engine it starts lives exactly as long as the command. Also: `fluksio sweep --param lr=0.1,0.01` for the product of the lists, `run --follow` for a run's numbers as they arrive, Ctrl-C cancelling a waited run rather than abandoning it, coloured statuses on a terminal, and `name` made optional on the metrics endpoint so a follower can ask for every series. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,77 @@
|
||||
"""The stage cache, from the side that needs a database.
|
||||
|
||||
The pipeline half — what a hit restores and what a key is made of — is in
|
||||
`tests/flow/test_runs.py`, which runs without one.
|
||||
"""
|
||||
|
||||
import json
|
||||
|
||||
from sqlmodel import Session
|
||||
|
||||
from fluksio.core.db import engine as db_engine
|
||||
from fluksio.flow.artifacts import ArtifactStore
|
||||
from fluksio.flow.pipeline import NodeOutcome
|
||||
from fluksio.flow.runs import OUTPUT_CAP, RunCache, _cacheable
|
||||
from fluksio.models import RunNode
|
||||
|
||||
|
||||
def test_a_run_cache_finds_what_an_earlier_run_recorded(tmp_path):
|
||||
"""The run history is the cache; there is no second store to keep."""
|
||||
store = ArtifactStore(tmp_path / "artifacts")
|
||||
reference = store.put([b"payload"], name="data.bin")
|
||||
plain, with_artifact, collected = "k-plain", "k-artifact", "k-collected"
|
||||
with Session(db_engine) as session:
|
||||
session.add(
|
||||
RunNode(
|
||||
run_id="cache-1",
|
||||
node="study.a",
|
||||
status="ok",
|
||||
cache_key=plain,
|
||||
outputs=json.dumps({"study.loss": 1.5}),
|
||||
)
|
||||
)
|
||||
session.add(
|
||||
RunNode(
|
||||
run_id="cache-2",
|
||||
node="study.b",
|
||||
status="ok",
|
||||
cache_key=with_artifact,
|
||||
outputs=json.dumps({"study.data": reference}),
|
||||
)
|
||||
)
|
||||
session.add(
|
||||
RunNode(
|
||||
run_id="cache-3",
|
||||
node="study.c",
|
||||
status="ok",
|
||||
cache_key=collected,
|
||||
outputs=json.dumps(
|
||||
{"study.data": {**reference, "digest": "sha256:" + "1" * 64}}
|
||||
),
|
||||
)
|
||||
)
|
||||
session.commit()
|
||||
|
||||
cache = RunCache(store)
|
||||
assert cache.lookup(plain) == (True, {"study.loss": 1.5})
|
||||
assert cache.lookup(with_artifact) == (True, {"study.data": reference})
|
||||
# Its bytes have gone from the store, so the reference names nothing a
|
||||
# restored run could open. That is a miss, not a broken run.
|
||||
assert cache.lookup(collected) == (False, None)
|
||||
assert cache.lookup("never-seen") == (False, None)
|
||||
assert cache.lookup("") == (False, None)
|
||||
|
||||
|
||||
def test_what_may_be_stored_as_a_cache_entry():
|
||||
"""A row carries a key and its outputs together, or neither."""
|
||||
ok = NodeOutcome(
|
||||
node="study.a", ok=True, cache_key="k", output_values={"study.loss": 1.0}
|
||||
)
|
||||
assert _cacheable(ok) == '{"study.loss":1.0}'
|
||||
# A node that published nothing is still an answer worth reusing.
|
||||
assert _cacheable(ok.model_copy(update={"output_values": None})) == "null"
|
||||
# Not cacheable: it failed, it has no key, or it returned too much.
|
||||
assert _cacheable(ok.model_copy(update={"ok": False})) is None
|
||||
assert _cacheable(ok.model_copy(update={"cache_key": ""})) is None
|
||||
big = {"study.data": "x" * (OUTPUT_CAP + 1)}
|
||||
assert _cacheable(ok.model_copy(update={"output_values": big})) is None
|
||||
Reference in New Issue
Block a user