Files
app/backend/fluksio/api/routes/artifacts.py
T
stroblmeandClaude Opus 5 0ffcabfdb9
Docs / docs (push) Successful in 30s
Playwright Tests / test-playwright (1, 2) (push) Successful in 3m7s
Playwright Tests / test-playwright (2, 2) (push) Successful in 1m54s
pre-commit / pre-commit (push) Failing after 4m24s
Test Backend / test-backend (push) Successful in 3m8s
Compose Smoke Test / test-compose (push) Successful in 40s
Playwright Tests / merge-reports (push) Successful in 1m33s
Media dtypes: image, audio and video as narrowed artifact references
A port may now declare `image`, `audio` or `video`. Each is the artifact
reference the engine already had, narrowed by the `media_type` on it, so a
speech recogniser declares what it eats rather than taking any bytes at all and
finding out. Bytes still never travel as a message and nothing on the wire
stops being JSON: a camera publishes one reference per frame, a microphone one
per chunk, and a reference may carry a `meta` dict nothing here interprets.

Streaming media is therefore an ordinary streaming port — with one change to
what that means. An emission used to journal an item with no payload, so
downstream read whatever was current when the item was claimed; a consumer
slower than its producer saw only the newest chunk and the ones between were
lost. That is right for a training curve and wrong for a second of speech, so
an emission now journals a `kind="emission"` item carrying its values, and the
executor hands them to the nodes reading that message instead of writing them
to state again. The value in state stays the latest, which is what everything
else reads, and the wave is filtered by what actually changed rather than
walking everything reachable. No queue serialization change — the existing
`outputs` field carries it.

Continuous media makes the store's missing GC a real problem, so this closes
it: `sweep_artifacts` runs hourly, keeps every digest a `run_artifact` row
records or a live message holds, spares anything written in the last hour, and
stands aside entirely while a run is in flight, since a node may store a
checkpoint long before it returns the reference to it. That also collects the
orphans a deleted flow has always left behind. `ARTIFACT_GC_INTERVAL_S=0` turns
it off.

Around the edges: `GET /artifacts/{digest}` serves the media type the caller
passes and answers ranged requests, so a browser plays a clip rather than
downloading it; `PUT` spools to disk instead of holding the whole body in
memory, as does `save_artifact` given a path; a Media widget draws whatever its
message points at, and a wall panel may fetch the bytes its own tiles are
showing and nothing else; and a connector gets `save_artifact`, for a device
whose readings are bytes.

What this cannot do is live video: a frame every second or two is a glance, and
the honest answer above that is the camera's own stream, which the widget takes
as a URL and the browser plays from source.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 23:44:55 +02:00

125 lines
4.2 KiB
Python

"""Artifacts over HTTP: the one way bytes get in and out of the store.
A node on this host could reach the directory itself, but a node on a remote
worker cannot — and having one path rather than two is what keeps a flow's
code the same wherever it runs.
"""
import re
import tempfile
from pathlib import Path
from typing import Any
from fastapi import APIRouter, Depends, HTTPException, Query, Request
from fastapi.responses import FileResponse
from jwt.exceptions import InvalidTokenError
from pydantic import BaseModel
from sqlmodel import Session
from starlette.concurrency import run_in_threadpool
from fluksio.api.deps import user_from_token
from fluksio.core import security
from fluksio.core.db import engine
from fluksio.flow.artifacts import ArtifactStore
#: What may be echoed back as a response content type. The caller holds the
#: reference and passes its media type, so this guards a header rather than
#: trusting one — anything else is served as bytes.
_MEDIA_TYPE = re.compile(r"^[\w.+-]+/[\w.+-]+$")
def artifact_caller(request: Request) -> str:
"""Who may move artifacts: a signed-in person, or an attached worker.
A worker's node stores its checkpoints through this endpoint, so its own
credential has to open it — and only it. The token is no use anywhere else
in the API, which is why this check is here rather than in the shared
dependency every other route uses.
"""
header = request.headers.get("Authorization", "")
token = header[7:] if header.lower().startswith("bearer ") else ""
if not token:
raise HTTPException(status_code=401, detail="Not authenticated")
try:
claims = security.decode_worker_token(token)
except InvalidTokenError:
pass
else:
return f"worker:{claims.get('sub')}"
with Session(engine) as session:
# With the request, so a credential that is scoped by route — a wall
# panel's — is judged against this one rather than waved through.
user = user_from_token(session, token, request)
if user is None:
raise HTTPException(status_code=401, detail="Not authenticated")
return user.email
router = APIRouter(
prefix="/artifacts", tags=["artifacts"], dependencies=[Depends(artifact_caller)]
)
class ArtifactRef(BaseModel):
digest: str
size: int
media_type: str
name: str = ""
def _store(request: Request) -> ArtifactStore:
store: ArtifactStore | None = getattr(request.app.state, "artifact_store", None)
if store is None:
raise HTTPException(status_code=503, detail="The artifact store is not ready")
return store
@router.put("", response_model=ArtifactRef)
async def put_artifact(
request: Request,
name: str = Query(default=""),
media_type: str = Query(default=""),
) -> Any:
"""Store the request body and answer with the reference to it.
Spooled to disk as it arrives rather than buffered: a video segment is as
legitimate a body here as a checkpoint, and neither should have to fit in
memory twice.
"""
store = _store(request)
handle = tempfile.NamedTemporaryFile(dir=store.root, delete=False)
try:
with handle:
async for chunk in request.stream():
handle.write(chunk)
return await run_in_threadpool(
store.put_file, Path(handle.name), media_type, name
)
finally:
Path(handle.name).unlink(missing_ok=True)
@router.get("/{digest}")
def get_artifact(
digest: str, request: Request, media_type: str = Query(default="")
) -> Any:
"""Serve one artifact back.
The caller passes the media type off the reference it holds, which is what
lets a browser play a clip rather than download it; the store itself keeps
only bytes. Ranged requests are answered because an audio or video element
scrubbing through a file asks for them.
"""
store = _store(request)
path = store.path(digest)
if path is None:
raise HTTPException(status_code=404, detail="No such artifact")
return FileResponse(
path,
media_type=(
media_type if _MEDIA_TYPE.match(media_type) else "application/octet-stream"
),
# The digest is the content, so it is also the perfect validator.
headers={"ETag": f'"{digest}"'},
)