Two container images provide the actual GPU compute the Cloudiy gateway drives:
| Image | Port | API | Purpose |
|---|---|---|---|
ghcr.io/cloudiy/worker-sdxl:latest |
7860 | /sdapi/v1/txt2img, /sdapi/v1/img2img |
Stable Diffusion image generation (AUTOMATIC1111 webui, API-only) |
ghcr.io/cloudiy/worker-ltx:latest |
7861 | POST /generate → writes <id>.mp4 to /out |
LTX-Video text-to-video |
onerahmet/openai-whisper-asr-webservice |
9000→9977 | POST /asr (multipart audio_file) |
Speech-to-text for whisper-ep. Public image, CPU — real today (like the Ollama text worker); the playground uploads a file as audio_b64. Model size is env-selectable — CLOUDIY_WHISPER_MODEL=tiny|base|small|medium|large-v3 (default base); a larger model transcribes more accurately but wants more RAM and a bigger first-run download. Changing it recreates the worker on the next call. |
ghcr.io/cloudiy/worker-tts:latest |
8000→9978 | POST /tts {"text"} → {"wav_b64"} |
Piper text-to-speech for chatterbox (CPU). Buildable from workers/tts — publish it and the endpoint serves for real. |
ghcr.io/cloudiy/worker-audio:latest |
— | (text-to-audio) | Not built yet — audio generation (stable-audio) still awaits a worker; the endpoint reports this honestly. |
Image and audio endpoints are GPU/worker-gated and report honestly ("needs" in
the JSON) until the image exists on the serving node; text (llama-ep, via a
CPU Ollama worker) runs today.
Both are GPU-only: they require a Linux host with an NVIDIA GPU and the
NVIDIA Container Toolkit, and must be run with --gpus all. They will not run on
macOS or on a CPU-only host.
workers/
├── sdxl/
│ ├── Dockerfile # CUDA base + AUTOMATIC1111 webui (--api --listen --port 7860 --nowebui)
│ └── launch.sh # optional checkpoint download, then boots the API server
└── ltx/
├── Dockerfile # CUDA base + diffusers LTXPipeline + FastAPI wrapper
└── server.py # FastAPI server on :7861, writes /out/<id>.mp4
Publishing images to GHCR and running them on GPU hardware are human steps — they require GPU hardware, a registry you control, and moving large images around.
Two ways to build & push:
-
Locally with the helper script (requires
docker login ghcr.iofirst):export GHCR_PAT=<github PAT with write:packages> echo "$GHCR_PAT" | docker login ghcr.io -u <github-username> --password-stdin ./scripts/build-workers.sh
-
Via CI — push a tag matching
worker-v*and.github/workflows/workers.ymlbuilds and pushes both images using the repo'sGITHUB_TOKEN:git tag worker-v0.1.0 git push origin worker-v0.1.0
# Image worker
docker run --gpus all -p 7860:7860 ghcr.io/cloudiy/worker-sdxl:latest
# then: curl -X POST localhost:7860/sdapi/v1/txt2img -d '{"prompt":"a cat","steps":20}'
# Video worker (mount an /out volume so you can retrieve the mp4)
docker run --gpus all -p 7861:7861 -v "$PWD/out:/out" ghcr.io/cloudiy/worker-ltx:latest
# then: curl -X POST localhost:7861/generate -d '{"prompt":"a rocket launch"}'- sdxl: no checkpoint is baked in. Mount one at
/app/models/Stable-diffusionor setDOWNLOAD_MODEL_URLto a.safetensorsURL (fetched at boot). Recommended: SDXL base 1.0. - ltx: weights (
Lightricks/LTX-Video) are pulled from HuggingFace on first request. Mount a volume at/root/.cache/huggingfaceto persist the cache.
The gateway runs each worker with a reduced blast radius so a compromised model or prompt can't easily pivot on the host:
--cap-drop ALL(no Linux capabilities),--security-opt no-new-privileges(no setuid escalation),--pids-limit 512(anti fork-bomb).--memorycap on every worker (host RAM can't be exhausted). The text worker also runs with a read-only root filesystem (--read-only), its only writable surfaces being the model volume and a--tmpfs /tmp, so a compromised model can't tamper with the image. The GPU workers aren't read-only yet (the webui / HF cache write to several paths) — tune per image when validating on a GPU node.- Input caps: prompts are bounded (16 KB) and text generation is capped
(
num_predict), plus per-request timeouts, so one call can't pin the GPU. - Per-consumer rate limit on the provider path (30 endpoint runs / 60 s per wallet identity) against flooding / cost-grief.
- Supply chain — pin images by digest. Each worker image is overridable by
env so you can pin a reviewed digest without recompiling:
CLOUDIY_OLLAMA_IMAGE,CLOUDIY_IMAGE_WORKER,CLOUDIY_VIDEO_WORKER(e.g.ollama/ollama@sha256:…). SetCLOUDIY_REQUIRE_PINNED_IMAGES=1to fail closed — refuse any worker image that isn't@sha256:-pinned, so a repointed:latesttag can't slip a malicious image in. - Seccomp. Docker's default seccomp profile is always in force (we never
pass
unconfined). SupplyCLOUDIY_SECCOMP_PROFILE=/path/to/profile.jsonto apply a stricter, validated profile of your own (test it against the GPU/CUDA workers first — an over-tight profile breaks them). - Signature verification (cosign). Set
CLOUDIY_COSIGN_VERIFY=1to verify each worker image's signature before it runs (fail closed). Configure a key (CLOUDIY_COSIGN_KEY) or keyless (CLOUDIY_COSIGN_IDENTITY+CLOUDIY_COSIGN_ISSUER). Requires thecosignbinary and signed images; pairs with digest pinning (pin what runs, verify who signed it). - Non-root user (opt-in).
CLOUDIY_WORKER_USER=1000:1000runs workers as a non-root UID. Off by default because several images assume root (they write to/root, the HF cache, …); enable only with an image and volumes prepared for that UID. - Egress-less serving with
CLOUDIY_WORKER_NO_EGRESS=1: workers run on a dedicated--internalDocker network (cloudiy-sealed) with no outbound (no exfil/callback), while the gateway can still reach the container's published port (unlike--network none, which also cuts the gateway off). The text worker caches models in a persistent volume (cloudiy-ollama-models); a pull needs egress, so warm each model once without the flag (or bake weights), then enable sealed serving — the weights persist. A sealed run for an un-warmed model fails with a clear message instead of hanging.
The software path is now wired; the remaining steps need your infra:
- Publish the worker images (human): build & push
worker-sdxl,worker-ltx(and, when built, the audio/tts/whisper images) to a registry the serving node can pull. See "Building & publishing" above. - Run a serving node (human): on a Linux + NVIDIA host,
cloudiy share --require-payment --rpc-url <solana-rpc>. It serves model endpoints over iroh and admits each run through the x402/escrow gate — an endpoint run is now paid and signed exactly like a kernel job (core::run_endpoint_guarded→authorize→serve_endpoint→signed_response). - Route runs to that node: the gateway's
POST /api/endpointaccepts an optionalto(the provider's EndpointId) andpayment; withtoset it dispatchesproto::Request::RunEndpointover iroh to that provider (payment enforced there) instead of running on the gateway host. Withoutto, the model runs on the local gateway (free, for solo/dev use). - Point the browser at a gateway: the UI reads
?gw=<url>(persisted) or defaults to the local gateway. Do not expose the gateway itself to the public internet — it holds the machine identity and is deliberately loopback-only (guard_local_origin, anti-CSRF/rebinding). Public reach is: browser → localcloudiy os(loopback) → iroh → remote provider (step 3), not a CORS-opened gateway.