Let a quarantine expire, and stop two tasks spending one budget
Docs / docs (push) Successful in 23s
Playwright Tests / test-playwright (1, 2) (push) Successful in 3m1s
Playwright Tests / test-playwright (2, 2) (push) Successful in 1m44s
pre-commit / pre-commit (push) Failing after 2m50s
Test Backend / test-backend (push) Successful in 2m39s
Compose Smoke Test / test-compose (push) Successful in 31s
Playwright Tests / merge-reports (push) Successful in 1m9s

A house's inverter broker dropped at 04:27 and the power flow was quarantined
20 seconds later. Quarantine was terminal — the supervised task returned and
only a publish or an engine restart could bring it back — so five hours of
power and battery readings are missing, and what ended it was an unrelated
`git pull` restarting uvicorn.

Two changes, both in that path:

- the failure budget is per task, not per flow. `power` runs an MQTT subscriber
  and a Victron keepalive publisher against the same broker; they died together
  and spent one shared budget in 41s, giving up before the 60s backoff step was
  ever reached.
- quarantine is now a rest. The task sits out 5min, then 15, then an hour, and
  each time gets its budget back and tries again, so a broker that comes back
  is picked up without anyone watching. `quarantined` reads from whichever
  tasks are currently resting.

The alert for it says when it will try again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5H4uLCCpsbipL1R7WKCee
This commit is contained in:
2026-08-28 10:35:13 +02:00
co-authored by Claude Opus 5
parent e40bbb9245
commit 8f1e685526
3 changed files with 140 additions and 36 deletions
+3 -1
View File
@@ -114,10 +114,12 @@ def describe(event: dict[str, Any]) -> Alert | None:
node=node,
)
if kind == "flow_quarantined":
retry = event.get("retry_in_s")
again = f" Trying again in {round(float(retry) / 60)} min." if retry else ""
return Alert(
title=f"Flow '{flow}' was quarantined",
body=(
f"It kept crashing, so the engine stopped restarting it. {error}"
f"It kept crashing, so the engine stood it down. {error}{again}"
).strip(),
flow=flow,
)
+71 -30
View File
@@ -3,9 +3,15 @@
A node's subscription, schedule or poll loop is an `asyncio.Task`, and a task
that raises is simply gone — the node stays listed as running while nothing
listens any more. The supervisor restarts those tasks with a growing delay and
gives up on a flow that keeps failing, because a flow crash-looping every
stands a flow down when one keeps failing, because a flow crash-looping every
second is worse than a flow that is visibly stopped.
Standing down is a rest, not a verdict. It used to be the end of the task: the
flow stayed dead until somebody published it or the engine restarted. On a
house that runs unattended that turns a broker rebooting at four in the morning
into a hole in the record until someone notices — so the rest expires, the
budget comes back, and the task tries again on a much longer clock.
Deliberately a plain task registry rather than a `TaskGroup`: a group cancels
its siblings when one member fails, which is the opposite of what supervision
means here.
@@ -25,10 +31,20 @@ from fluksio.flow.events import EventBus
logger = logging.getLogger(__name__)
BACKOFF = (1.0, 5.0, 30.0, 60.0)
# A flow that burns through this many restarts in the window is not going to
# recover by being restarted again.
# A task that burns through this many restarts in the window is not going to
# recover by being restarted again *now*.
#
# Per task rather than per flow: a flow whose subscriber and its keepalive
# publisher both go down with the same broker spends one shared budget twice as
# fast as it was sized for, and gives up before the backoff above has reached
# its longest step. That is what took the house's inverter readings out for
# nearly five hours — six crashes across two tasks inside 41 seconds.
FAILURE_BUDGET = 5
FAILURE_WINDOW = 300.0
# How long a task sits out before its budget is handed back. Minutes rather
# than seconds because a device that is really gone should not be hammered, and
# capped in the hour because nothing here is worth losing a day of readings to.
QUARANTINE_BACKOFF = (300.0, 900.0, 3600.0)
# How long a cancelled task gets to notice. A loop that is still waiting after
# this is not going to stop on its own — a client closing a socket the broker
# no longer answers on is the case seen in the wild — and the rebuild asking
@@ -48,8 +64,15 @@ class Supervisor:
#: task's name, because names are a flow and a node joined by a dot and
#: matching on that prefix would let 'hea' cancel 'heating'.
self._flows: dict[str, str] = {}
#: Crash times per *task*, keyed by task name.
self._failures: dict[str, deque[float]] = {}
self.quarantined: set[str] = set()
#: The tasks sitting out a quarantine, and the flow each belongs to.
self._resting: dict[str, str] = {}
@property
def quarantined(self) -> set[str]:
"""The flows with at least one task currently standing down."""
return set(self._resting.values())
def spawn(self, name: str, flow: str, factory: TaskFactory) -> None:
"""Run `factory()` and keep running it until told to stop."""
@@ -62,6 +85,7 @@ class Supervisor:
async def _supervise(self, name: str, flow: str, factory: TaskFactory) -> None:
attempt = 0
rests = 0
while True:
try:
await factory()
@@ -69,7 +93,10 @@ class Supervisor:
raise
except Exception as exc:
if not self._record_failure(name, flow, exc):
return
await self._rest(name, flow, exc, rests)
rests += 1
attempt = 0
continue
else:
# A clean return means the loop decided it was done.
return
@@ -78,8 +105,37 @@ class Supervisor:
attempt += 1
await asyncio.sleep(delay)
async def _rest(self, name: str, flow: str, exc: Exception, rests: int) -> None:
"""Stand the task down, then hand its budget back and let it retry."""
delay = QUARANTINE_BACKOFF[min(rests, len(QUARANTINE_BACKOFF) - 1)]
logger.error(
"Flow '%s' crashed %d times in %.0fs — quarantined, retrying in %.0fs",
flow,
FAILURE_BUDGET,
FAILURE_WINDOW,
delay,
)
self._resting[name] = flow
self._publish(
{
"type": "flow_quarantined",
"flow": flow,
"task": name,
"error": f"{type(exc).__name__}: {exc}",
"retry_in_s": delay,
"ts": time.time(),
}
)
try:
await asyncio.sleep(delay)
finally:
# Both, and in this order: a task woken by a cancellation is being
# torn down, and must not leave the flow looking quarantined.
self._resting.pop(name, None)
self._failures.pop(name, None)
def _record_failure(self, name: str, flow: str, exc: Exception) -> bool:
"""Note the crash; False when the flow has spent its budget."""
"""Note the crash; False when this task has spent its budget."""
logger.warning("Supervised task '%s' crashed: %s", name, exc, exc_info=True)
self._publish(
{
@@ -92,36 +148,20 @@ class Supervisor:
)
now = time.monotonic()
window = self._failures.setdefault(flow, deque())
window = self._failures.setdefault(name, deque())
window.append(now)
while window and window[0] < now - FAILURE_WINDOW:
window.popleft()
if len(window) < FAILURE_BUDGET:
return True
logger.error(
"Flow '%s' crashed %d times in %.0fs — quarantined",
flow,
len(window),
FAILURE_WINDOW,
)
self.quarantined.add(flow)
self._publish(
{
"type": "flow_quarantined",
"flow": flow,
"error": f"{type(exc).__name__}: {exc}",
"ts": time.time(),
}
)
return False
return len(window) < FAILURE_BUDGET
async def cancel_all(self) -> None:
"""Stop supervising. Idempotent, and safe to call mid-restart."""
tasks = list(self._tasks.values())
self._tasks.clear()
self._flows.clear()
self._resting.clear()
self._failures.clear()
await self._cancel(tasks)
async def cancel_flow(self, flow: str) -> None:
@@ -136,11 +176,12 @@ class Supervisor:
tasks = [self._tasks.pop(name) for name in names if name in self._tasks]
for name in names:
del self._flows[name]
# The clean slate: whatever these tasks spent before, the build
# that follows starts their budgets again — and a rebuild is not
# made to wait out a rest somebody has just fixed the cause of.
self._failures.pop(name, None)
self._resting.pop(name, None)
await self._cancel(tasks)
# The clean slate: whatever this flow spent before, the build that
# follows starts its budget again.
self.quarantined.discard(flow)
self._failures.pop(flow, None)
async def _cancel(self, tasks: list[asyncio.Task[None]]) -> None:
"""Ask these tasks to stop, and wait no longer than the grace period."""