Your render farm reports 100% success. Some of those frames are wrong.
Check the gate is the call on a film set to inspect the camera gate before the crew strikes a setup. A hair in the gate means the take looked perfect through the viewfinder and the negative is ruined. You only find out later, when it is expensive.
Render farms have the same failure. A renderer that cannot find a texture does
not stop — it substitutes a fallback and writes the frame. A scene saved with
the wrong colour transform renders happily. A sample count clobbered by a bad
submit produces a noisy frame at full speed. Every one of them exits 0. The
queue is green, the dashboards are green, the wrangler's shift report says the
night went fine, and three days later a lighting supervisor asks why the crates
are grey.
GATECHECK is a night-shift render wrangler that goes looking for those frames.
| Demo | three minutes, showing it run |
| Console | https://gatecheck-console-594428061375.europe-west1.run.app |
| Dashboard | what it looks like with real data |
| Five-minute review | JUDGES.md · PROOF.md · MOCKS.md |
- A deterministic verifier scores every rendered frame against its own neighbours — histogram divergence, unique-colour collapse, edge density, noise estimate, block variance in the tail of the file — and writes a verdict as a Prometheus metric and a Loki line. No model is involved. The statistics decide where the problem is.
- An agent on Gemini reads that through Grafana, correlates it with the farm's own queue metrics and the renderer's logs, and explains what happened: which shot, which host, which scene change, and the log line that proves it.
- It prices the fix. Roughly three quarters of the cost of a wasted frame is renderer licence time, not compute. Retrying a frame whose texture is genuinely absent from disk spends that money to produce the same wrong picture. The agent says so, in dollars, and recommends one action.
- It asks permission, and it can be refused. The action is held at a graph interrupt for a human. Then it is checked again, server-side, against limits the agent cannot argue with.
- It writes the shift up — a dashboard annotation, an incident with the studio runbook for that defect class, and a snapshot so the evidence outlives the retention window. Including when the answer was "do nothing".
The agent is not trusted, and the distrust is not implemented in its prompt.
The rule that stops it lives in Grafana, not in the agent. Grafana Agent Observability supports guards: rules stored in the tenant that run inline on the request path and can deny a tool call. They are inert until an application calls the hooks endpoint and obeys the answer — so GATECHECK calls it before every action that touches the farm.
One guard ships with the project. It denies farm_action_retry when the frame's
defect class is one that rerunning cannot fix. Same tool, same arguments, only
the evidence differs:
| tool | defect class | decision |
|---|---|---|
farm_action_retry |
missing_texture |
deny |
farm_action_retry |
wrong_view_transform |
deny |
farm_action_retry |
oom_kill |
allow |
farm_action_retry |
stuck_frame |
allow |
farm_action_kill |
missing_texture |
allow |
query_prometheus |
missing_texture |
allow |
An operator can edit that rule without reading a line of the agent's source, and
the agent obeys it whether or not it agrees. The vendor SDK defaults to
fail_open=True so that a transport error never blocks an LLM call; for
something that can spend money on a render farm that default is wrong, and we
override it. If the guard endpoint is unreachable, nothing happens.
And the agent has to show it has been right lately. Its own generations and evaluation scores are recorded in the same Grafana stack. Before proposing an action it reads its own report card back and refuses to act when the recent score is below the floor, or when there are too few graded samples to say anything at all. A wrangler who cannot prove he has been right does not get to touch the farm.
The model is never handed a tool that can act. farm_action_retry and its
siblings live on the same MCP server as everything else, but no language model
in this system is given one. They are called by a deterministic node, after a
human approves a specific validated action, and the server re-checks it anyway.
Blender on OpenCue deterministic Grafana Cloud
┌────────────────────┐ ┌──────────────┐ ┌────────────────┐
│ cuebot · rqd-01/02 │ frames │ verifier │ metrics │ Mimir · Loki │
│ rest-gateway ├───────►│ (no model) ├─────────────►│ Tempo │
└─────────┬──────────┘ └──────────────┘ logs │ Agent Obs. │
│ actions └───────┬────────┘
│ │ MCP
│ ┌──────────────────────────────┐ │
└──────────────┤ mcp-gatecheck │◄───────────┘
│ Grafana's own tools, plus │
│ ours, plus a policy engine │
└──────────────┬───────────────┘
│ MCP
┌──────────────▼───────────────┐
│ wrangler agent (ADK graph) │
│ Gemini 3.8 Flash │
└──────────────────────────────┘
mcp-gatecheck is one binary that imports github.com/grafana/mcp-grafana as a
library and registers Grafana's own tool implementations unmodified
alongside the render-farm tools we added, over a single MCP endpoint. The agent
has no Prometheus client, no Loki client, no dashboard API wrapper and no auth
plumbing of its own.
The graph:
START → triage → price → plan → self_check ─┬─ passed ─→ approval_gate → act → verify ─┐
│ ├→ write_back
└─ refused ─→ escalate ────────────────────┘
triage and price are model nodes with tools. plan, self_check,
approval_gate, act and escalate are ordinary Python. Between a proposal and
the render farm there are three independent refusals: a schema that rejects a
malformed proposal, a guard in Grafana that can deny the call, and a server-side
policy that re-checks the action after a human has approved it.
| path | what it is |
|---|---|
farm/ |
a real OpenCue render farm in Docker, rendering real Blender frames |
gaffer/ |
deterministic, seeded fault injection — the failures are real, not simulated |
verifier/ |
the frame forensics. Pure NumPy and Pillow. No model, no network in the analysis path |
mcp-gatecheck/ |
the MCP server: Grafana's tools, our tools, the policy engine, and an MCP App |
agent/ |
the wrangler graph on ADK and Gemini |
grafana/ |
guards and dashboards, provisioned as code |
alloy/ |
telemetry collection into Grafana Cloud |
eval/ |
the graded scorecard, and the answer key the agent cannot reach |
console/ |
the operator's front door |
cp .env.example .env # fill in your Grafana and Google Cloud values
./scripts/farm.sh up # OpenCue, Prometheus, Loki, two render nodes
./scripts/farm.sh submit # render a shot, with seeded faults
./scripts/apply-guards.sh # install the guards into your Grafana tenant
adk web agent # or: adk run agentFull setup, including the Grafana Cloud free-tier prerequisites, is in
docs/SETUP.md. A five-minute reviewer path is in
JUDGES.md.
MOCKS.md draws the line explicitly: what is real, what is standing in for something, and what was not built. The short version is that the farm, the frames, the faults, the detection, the telemetry and the guards are real, and the show is invented — there is no studio here, and deliberately no real film titles or logos anywhere.
- The verifier finds frames that are statistically unlike their neighbours. A shot that is uniformly wrong from frame 1 — every frame equally miscoloured, with no clean neighbour to compare against — is invisible to it. It needs either a reference render or a human.
- It will flag a legitimate hard cut or a deliberate strobe as suspect. The
scorecard in
eval/reports that false-positive rate rather than hiding it. - The cost model uses published render-farm unit economics, not your studio's contract. The dollar figures are directionally right and specifically wrong.
p(success)for a retry comes from a fixed table keyed on defect class. It is a stated prior, not a learned estimate, and it is in the repository so you can disagree with it.- Agent Observability guards and score read-back are Grafana Cloud features. Against a self-hosted Grafana the agent runs, and the self-check correctly refuses to authorise anything, because it cannot prove it has been reliable.
Apache-2.0. See LICENSE.