Close the nine open SDK tasks: one engine per directory, a tabbed dashboard, re-pairing, run recovery
Docs / docs (push) Successful in 27s
Playwright Tests / test-playwright (1, 2) (push) Failing after 17s
Playwright Tests / test-playwright (2, 2) (push) Failing after 12s
pre-commit / pre-commit (push) Failing after 1m59s
Test Backend / test-backend (push) Failing after 2m30s
Compose Smoke Test / test-compose (push) Failing after 13s
Playwright Tests / merge-reports (push) Failing after 2m19s

serve: refuse a second engine for one data directory whatever port it was
asked for, using the pidfile and a token this directory signed. The check
runs before the database is touched and before the credential is written,
which is what left every later CLI call pointing at a dead port.

The terminal dashboard is three tabs (Overview, Runs, Logs) with the toolbar
following the focused pane, the engine's output goes to serve.log rather than
down a pipe, and closing the screen stops both reader threads so the prompt
comes back. It adopts a running engine on every start, so stop/start and
restart work on one it did not start, and a stop waits for the process to be
gone before the next start. Enrolment reports itself in the modal.

enroll: a new claim code replaces the pairing instead of being refused. The
code is redeemed before anything is written, mappings to a portal being left
are cleared, and a running engine redials when the stored enrolment changes.

runs: an engine re-queues the runs left `queued` by the one before it, and
`fluksio retry <id>` / `retry --group <sweep>` submits an interrupted run
again with the same inputs and group, recorded through Run.parent_id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U9BoNGq6V9MdRWAte7JBuC
This commit is contained in:
2026-08-31 17:37:13 +02:00
co-authored by Claude Opus 5
parent bdad6d7fc2
commit 8a94bf10d7
14 changed files with 873 additions and 140 deletions
+18 -2
View File
@@ -60,6 +60,11 @@ MAX_IN_FLIGHT = 256
ENROL_POLL_S = 3.0
def _identity(config: cloud_config.CloudConfig | None) -> tuple[str, str] | None:
"""Which enrolment a link is running on, so a replaced one is noticed."""
return (config.instance_id, config.token) if config is not None else None
def start(app: FastAPI) -> None:
"""Dial the portal, replacing any link already up."""
existing = getattr(app.state, "cloud_task", None)
@@ -67,6 +72,7 @@ def start(app: FastAPI) -> None:
existing.cancel()
connector = CloudConnector(app)
app.state.cloud_connector = connector
app.state.cloud_identity = _identity(cloud_config.load())
app.state.cloud_task = asyncio.create_task(
connector.serve_forever(), name="cloud-connector"
)
@@ -85,7 +91,7 @@ async def watch_enrolment(app: FastAPI) -> None:
task = getattr(app.state, "cloud_task", None)
if task is not None and task.done():
# It returns of its own accord when the config goes away, which is
# what `fluksio disconnect` and the portal's own Disconnect do.
# what Disconnect, here or on the portal, does.
app.state.cloud_task = None
app.state.cloud_connector = None
task = None
@@ -94,7 +100,17 @@ async def watch_enrolment(app: FastAPI) -> None:
# started — which, started from here, is a restart every few seconds.
# A config that is fine but unreachable keeps its task, and the
# retrying belongs to the connector rather than to this.
if task is None and cloud_config.load() is not None:
config = cloud_config.load()
if task is not None and _identity(config) != getattr(
app.state, "cloud_identity", None
):
# Enrolled again, at this portal or another one. `fluksio enroll`
# is its own process and cannot cancel this task, so the link would
# otherwise stay up on the credential that was replaced.
logger.info("Enrolment replaced while running; redialling")
task.cancel()
task = None
if task is None and config is not None:
logger.info("Enrolled while running; dialling the portal")
start(app)