Close the nine open SDK tasks: one engine per directory, a tabbed dashboard, re-pairing, run recovery
Docs / docs (push) Successful in 27s
Playwright Tests / test-playwright (1, 2) (push) Failing after 17s
Playwright Tests / test-playwright (2, 2) (push) Failing after 12s
pre-commit / pre-commit (push) Failing after 1m59s
Test Backend / test-backend (push) Failing after 2m30s
Compose Smoke Test / test-compose (push) Failing after 13s
Playwright Tests / merge-reports (push) Failing after 2m19s

serve: refuse a second engine for one data directory whatever port it was
asked for, using the pidfile and a token this directory signed. The check
runs before the database is touched and before the credential is written,
which is what left every later CLI call pointing at a dead port.

The terminal dashboard is three tabs (Overview, Runs, Logs) with the toolbar
following the focused pane, the engine's output goes to serve.log rather than
down a pipe, and closing the screen stops both reader threads so the prompt
comes back. It adopts a running engine on every start, so stop/start and
restart work on one it did not start, and a stop waits for the process to be
gone before the next start. Enrolment reports itself in the modal.

enroll: a new claim code replaces the pairing instead of being refused. The
code is redeemed before anything is written, mappings to a portal being left
are cleared, and a running engine redials when the stored enrolment changes.

runs: an engine re-queues the runs left `queued` by the one before it, and
`fluksio retry <id>` / `retry --group <sweep>` submits an interrupted run
again with the same inputs and group, recorded through Run.parent_id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9BoNGq6V9MdRWAte7JBuC
This commit is contained in:
2026-08-31 17:37:13 +02:00
co-authored by Claude Opus 5
parent bdad6d7fc2
commit 8a94bf10d7
14 changed files with 873 additions and 140 deletions
+10 -1
View File
@@ -296,11 +296,20 @@ one the automations use: a burst of five hundred sweep runs must not stand
between a house and its heating. An engine that is down when a run is
submitted picks it up when it starts.
The row is what makes that true rather than the queue. A run is written before
the work item is added, so an engine reads its own history at startup and
wakes itself for anything still `queued` — which is what a run submitted
seconds before a restart is, and what an in-memory queue would otherwise have
lost with the process.
From the moment a run is claimed, its database row is the record and the queue
is finished with it. Redelivering two hours of training because an
acknowledgement was late is not recovery; instead a running run refreshes a
lease, and one whose lease goes stale is marked `abandoned`, which is what a
run whose engine was killed mid-training becomes.
run whose engine was killed mid-training becomes. Starting it over is a
decision rather than something that happens: [`fluksio
retry`](../code/cli.md#fluksio-retry) submits it again, and `--group` does
that for the runs of a sweep that did not finish.
## Looking at what ran