Media dtypes: image, audio and video as narrowed artifact references
Docs / docs (push) Successful in 30s
Playwright Tests / test-playwright (1, 2) (push) Successful in 3m7s
Playwright Tests / test-playwright (2, 2) (push) Successful in 1m54s
pre-commit / pre-commit (push) Failing after 4m24s
Test Backend / test-backend (push) Successful in 3m8s
Compose Smoke Test / test-compose (push) Successful in 40s
Playwright Tests / merge-reports (push) Successful in 1m33s
Docs / docs (push) Successful in 30s
Playwright Tests / test-playwright (1, 2) (push) Successful in 3m7s
Playwright Tests / test-playwright (2, 2) (push) Successful in 1m54s
pre-commit / pre-commit (push) Failing after 4m24s
Test Backend / test-backend (push) Successful in 3m8s
Compose Smoke Test / test-compose (push) Successful in 40s
Playwright Tests / merge-reports (push) Successful in 1m33s
A port may now declare `image`, `audio` or `video`. Each is the artifact
reference the engine already had, narrowed by the `media_type` on it, so a
speech recogniser declares what it eats rather than taking any bytes at all and
finding out. Bytes still never travel as a message and nothing on the wire
stops being JSON: a camera publishes one reference per frame, a microphone one
per chunk, and a reference may carry a `meta` dict nothing here interprets.
Streaming media is therefore an ordinary streaming port — with one change to
what that means. An emission used to journal an item with no payload, so
downstream read whatever was current when the item was claimed; a consumer
slower than its producer saw only the newest chunk and the ones between were
lost. That is right for a training curve and wrong for a second of speech, so
an emission now journals a `kind="emission"` item carrying its values, and the
executor hands them to the nodes reading that message instead of writing them
to state again. The value in state stays the latest, which is what everything
else reads, and the wave is filtered by what actually changed rather than
walking everything reachable. No queue serialization change — the existing
`outputs` field carries it.
Continuous media makes the store's missing GC a real problem, so this closes
it: `sweep_artifacts` runs hourly, keeps every digest a `run_artifact` row
records or a live message holds, spares anything written in the last hour, and
stands aside entirely while a run is in flight, since a node may store a
checkpoint long before it returns the reference to it. That also collects the
orphans a deleted flow has always left behind. `ARTIFACT_GC_INTERVAL_S=0` turns
it off.
Around the edges: `GET /artifacts/{digest}` serves the media type the caller
passes and answers ranged requests, so a browser plays a clip rather than
downloading it; `PUT` spools to disk instead of holding the whole body in
memory, as does `save_artifact` given a path; a Media widget draws whatever its
message points at, and a wall panel may fetch the bytes its own tiles are
showing and nothing else; and a connector gets `save_artifact`, for a device
whose readings are bytes.
What this cannot do is live video: a frame every second or two is a glance, and
the honest answer above that is the camera's own stream, which the widget takes
as a URL and the browser plays from source.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -19,6 +19,7 @@ import hashlib
|
||||
import logging
|
||||
import os
|
||||
import tempfile
|
||||
import time
|
||||
from collections.abc import Iterable, Iterator
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
@@ -96,11 +97,13 @@ class ArtifactStore:
|
||||
"name": name,
|
||||
}
|
||||
|
||||
def put_file(self, path: Path, media_type: str = "") -> dict[str, Any]:
|
||||
def put_file(
|
||||
self, path: Path, media_type: str = "", name: str = ""
|
||||
) -> dict[str, Any]:
|
||||
with path.open("rb") as handle:
|
||||
return self.put(
|
||||
iter(lambda: handle.read(CHUNK), b""),
|
||||
name=path.name,
|
||||
name=name or path.name,
|
||||
media_type=media_type,
|
||||
)
|
||||
|
||||
@@ -119,20 +122,29 @@ class ArtifactStore:
|
||||
while chunk := handle.read(CHUNK):
|
||||
yield chunk
|
||||
|
||||
def collect(self, keep: set[str]) -> int:
|
||||
"""Delete what no run refers to any more. Returns how many went.
|
||||
def collect(self, keep: set[str], grace_s: float = 0.0) -> int:
|
||||
"""Delete what nothing refers to any more. Returns how many went.
|
||||
|
||||
The caller passes every digest still recorded; anything else in the
|
||||
store was produced by a run that has since been pruned, or never got a
|
||||
row at all because the run failed between writing and recording.
|
||||
|
||||
``grace_s`` spares anything written that recently. Storing bytes and
|
||||
recording the reference to them are two steps, and a sweep landing
|
||||
between them would take an artifact its run is about to name — so
|
||||
recent files are left for the next pass, by which time they are either
|
||||
referenced or genuinely orphaned.
|
||||
"""
|
||||
removed = 0
|
||||
cutoff = time.time() - grace_s
|
||||
for entry in self.root.glob("*/*"):
|
||||
if not entry.is_file():
|
||||
continue
|
||||
if DIGEST_PREFIX + entry.name in keep:
|
||||
continue
|
||||
try:
|
||||
if grace_s > 0 and entry.stat().st_mtime > cutoff:
|
||||
continue
|
||||
entry.unlink()
|
||||
removed += 1
|
||||
except OSError:
|
||||
|
||||
Reference in New Issue
Block a user