Refuse a port the function cannot take, and read a failing node as degraded
Docs / docs (push) Successful in 33s
Playwright Tests / test-playwright (1, 2) (push) Failing after 1m26s
Playwright Tests / test-playwright (2, 2) (push) Failing after 15s
pre-commit / pre-commit (push) Failing after 1m43s
Test Backend / test-backend (push) Failing after 2m46s
Compose Smoke Test / test-compose (push) Failing after 12s
Playwright Tests / merge-reports (push) Canceled after 0s
Docs / docs (push) Successful in 33s
Playwright Tests / test-playwright (1, 2) (push) Failing after 1m26s
Playwright Tests / test-playwright (2, 2) (push) Failing after 15s
pre-commit / pre-commit (push) Failing after 1m43s
Test Backend / test-backend (push) Failing after 2m46s
Compose Smoke Test / test-compose (push) Failing after 12s
Playwright Tests / merge-reports (push) Canceled after 0s
A python node's ports and settings arrive as keyword arguments, so a declared name its `process` does not take was a TypeError on every call — and a node that loads fine and fails every time it runs is the quiet kind of broken: the hosted demo did it 720 times an hour for two days and the health badge read ok throughout. `_build_node` now reads a written body with `ast` and refuses the mismatch at load, so the node is an error on the canvas and an issue on publish. Skipped for `**kwargs`, a decorated or absent `process`, and the template a new node opens with. The SDK's generated shim always takes `**settings`, so synced flows are untouched. `/observability/summary` names a node that has failed in the last fifteen minutes and reads degraded while it does, which is what would have made the badge amber. `nodes.failing` carries the count. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkgNtaR6JspnHBFFP6crZj
This commit is contained in:
@@ -7,6 +7,7 @@ which the generated SDK turns into a thrown error — and a health page that
|
||||
cannot render while the engine is degraded is the wrong way round.
|
||||
"""
|
||||
|
||||
import time
|
||||
from datetime import UTC, datetime, timedelta
|
||||
from typing import Annotated, Any, Literal
|
||||
|
||||
@@ -191,6 +192,20 @@ async def read_summary(request: Request, controller: FlowControllerDep) -> Any:
|
||||
f"{len(unhealthy)} node(s) down: "
|
||||
f"{', '.join(sorted(e.id for e in unhealthy))}"
|
||||
)
|
||||
# A node that loads and then fails on every call is not in `errored`, and
|
||||
# until this it read as healthy. ponytail: one failure reads degraded for
|
||||
# 15 minutes; a counter with decay if that proves noisy.
|
||||
now = time.time()
|
||||
failing = [
|
||||
e
|
||||
for e in entries
|
||||
if e.last_error_ts is not None and now - e.last_error_ts < 900
|
||||
]
|
||||
if failing:
|
||||
problems.append(
|
||||
f"{len(failing)} node(s) failing: "
|
||||
f"{', '.join(sorted(e.id for e in failing))}"
|
||||
)
|
||||
|
||||
# What the canvas flags on a flow — a dependency loop, an input nothing
|
||||
# feeds — stops that flow running just as surely as a node that will not
|
||||
@@ -234,6 +249,7 @@ async def read_summary(request: Request, controller: FlowControllerDep) -> Any:
|
||||
"total": len(entries),
|
||||
"error": len(errored),
|
||||
"unhealthy": len(unhealthy),
|
||||
"failing": len(failing),
|
||||
},
|
||||
queue=queue,
|
||||
loop_lag=(
|
||||
|
||||
Reference in New Issue
Block a user