Backend: real alert test results, state cleanup on delete/rename, node trigger errors, queue and collector fixes
`AlertManager.send` swallowed every delivery failure, so the alerts screen's Test button answered 200 whatever happened — the one thing it exists for. It takes `raise_on_error` now, which only the test route passes; the per-channel loop keeps the swallow, because one dead channel must not stop the others hearing about the same fault. A refused delivery answers 502 with whatever the sender said. Renaming a flow left its values under the old name for good: the delete path already swept them, the rename path never did. It calls the same `forget_flow`, which covers the messages and the `__ts__`/`__version__`/`__history__` bookkeeping keyed by message name. Cleanup, not migration — they repopulate under the new name on the next run. Triggering a node by hand ran `Node.__call__` with nothing catching it, so a node that raised produced a 500 and a stack trace in the server log, and nothing at all on the canvas. `Pipeline.publish_error` is the reporting half of `_execute_node` lifted out; both paths go through it, so a manual failure now reads the same on the canvas and in the metrics as a queued one. The route answers 400 with the node's error. `MemoryWorkQueue.stats()` counts claimed-but-unacknowledged work rather than reporting zero, so the health tile means something without Redis. The metrics collector's held tracebacks are capped at `DETAIL_CAP` and swept on the same `RUN_STALE_S` cutoff the open runs use, instead of one untruncated traceback per node kept for the life of the process — a traceback still survives the flush between the log and the failure it belongs to. `GET /observability/events` takes `since`/`until`, the window `/runs` already took, so a failures list can cover the span the charts beside it are drawn from. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XC2jX6Hdj7pxGGKzBTrbqB
This commit is contained in:
@@ -134,3 +134,35 @@ def test_runs_narrow_to_one_minute(
|
||||
# The upper bound is exclusive, so the run starting the next minute is not
|
||||
# in this one.
|
||||
assert [run["id"] for run in runs] == ["minute-in"]
|
||||
|
||||
|
||||
def test_events_narrow_to_one_minute(
|
||||
client: TestClient, superuser_token_headers: dict[str, str], db: Session
|
||||
) -> None:
|
||||
"""The failures list reaches as far back as the charts beside it."""
|
||||
minute = datetime.now(timezone.utc).replace(second=0, microsecond=0) - timedelta(
|
||||
hours=4
|
||||
)
|
||||
db.add(EngineEvent(ts=minute, type="node_error", flow=FLOW, detail="minute-in"))
|
||||
db.add(
|
||||
EngineEvent(
|
||||
ts=minute + timedelta(minutes=1),
|
||||
type="node_error",
|
||||
flow=FLOW,
|
||||
detail="minute-after",
|
||||
)
|
||||
)
|
||||
db.commit()
|
||||
|
||||
events = client.get(
|
||||
f"{PREFIX}/events",
|
||||
headers=superuser_token_headers,
|
||||
params={
|
||||
"flow": FLOW,
|
||||
"since": minute.isoformat(),
|
||||
"until": (minute + timedelta(minutes=1)).isoformat(),
|
||||
},
|
||||
).json()
|
||||
|
||||
# The upper bound is exclusive, so the failure a minute later is not in it.
|
||||
assert [event["detail"] for event in events] == ["minute-in"]
|
||||
|
||||
Reference in New Issue
Block a user