Skip to content
RaYYeR220Public

About

A night-shift render wrangler: finds the frames a render farm reports as successful and proves are wrong, prices the fix, and can only act inside limits it cannot argue past.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

GATECHECK

Your render farm reports 100% success. Some of those frames are wrong.

Check the gate is the call on a film set to inspect the camera gate before the crew strikes a setup. A hair in the gate means the take looked perfect through the viewfinder and the negative is ruined. You only find out later, when it is expensive.

Render farms have the same failure. A renderer that cannot find a texture does not stop — it substitutes a fallback and writes the frame. A scene saved with the wrong colour transform renders happily. A sample count clobbered by a bad submit produces a noisy frame at full speed. Every one of them exits 0. The queue is green, the dashboards are green, the wrangler's shift report says the night went fine, and three days later a lighting supervisor asks why the crates are grey.

GATECHECK is a night-shift render wrangler that goes looking for those frames.

Demo three minutes, showing it run
Console https://gatecheck-console-594428061375.europe-west1.run.app
Dashboard what it looks like with real data
Five-minute review JUDGES.md · PROOF.md · MOCKS.md

What it actually does

  1. A deterministic verifier scores every rendered frame against its own neighbours — histogram divergence, unique-colour collapse, edge density, noise estimate, block variance in the tail of the file — and writes a verdict as a Prometheus metric and a Loki line. No model is involved. The statistics decide where the problem is.
  2. An agent on Gemini reads that through Grafana, correlates it with the farm's own queue metrics and the renderer's logs, and explains what happened: which shot, which host, which scene change, and the log line that proves it.
  3. It prices the fix. Roughly three quarters of the cost of a wasted frame is renderer licence time, not compute. Retrying a frame whose texture is genuinely absent from disk spends that money to produce the same wrong picture. The agent says so, in dollars, and recommends one action.
  4. It asks permission, and it can be refused. The action is held at a graph interrupt for a human. Then it is checked again, server-side, against limits the agent cannot argue with.
  5. It writes the shift up — a dashboard annotation, an incident with the studio runbook for that defect class, and a snapshot so the evidence outlives the retention window. Including when the answer was "do nothing".

The part we think is new

The agent is not trusted, and the distrust is not implemented in its prompt.

The rule that stops it lives in Grafana, not in the agent. Grafana Agent Observability supports guards: rules stored in the tenant that run inline on the request path and can deny a tool call. They are inert until an application calls the hooks endpoint and obeys the answer — so GATECHECK calls it before every action that touches the farm.

One guard ships with the project. It denies farm_action_retry when the frame's defect class is one that rerunning cannot fix. Same tool, same arguments, only the evidence differs:

tool defect class decision
farm_action_retry missing_texture deny
farm_action_retry wrong_view_transform deny
farm_action_retry oom_kill allow
farm_action_retry stuck_frame allow
farm_action_kill missing_texture allow
query_prometheus missing_texture allow

An operator can edit that rule without reading a line of the agent's source, and the agent obeys it whether or not it agrees. The vendor SDK defaults to fail_open=True so that a transport error never blocks an LLM call; for something that can spend money on a render farm that default is wrong, and we override it. If the guard endpoint is unreachable, nothing happens.

And the agent has to show it has been right lately. Its own generations and evaluation scores are recorded in the same Grafana stack. Before proposing an action it reads its own report card back and refuses to act when the recent score is below the floor, or when there are too few graded samples to say anything at all. A wrangler who cannot prove he has been right does not get to touch the farm.

The model is never handed a tool that can act. farm_action_retry and its siblings live on the same MCP server as everything else, but no language model in this system is given one. They are called by a deterministic node, after a human approves a specific validated action, and the server re-checks it anyway.


Architecture

      Blender on OpenCue            deterministic                  Grafana Cloud
    ┌────────────────────┐        ┌──────────────┐              ┌────────────────┐
    │ cuebot · rqd-01/02 │ frames │  verifier    │  metrics     │  Mimir · Loki  │
    │ rest-gateway       ├───────►│  (no model)  ├─────────────►│  Tempo         │
    └─────────┬──────────┘        └──────────────┘  logs        │  Agent Obs.    │
              │ actions                                          └───────┬────────┘
              │                                                          │ MCP
              │              ┌──────────────────────────────┐            │
              └──────────────┤  mcp-gatecheck               │◄───────────┘
                             │  Grafana's own tools, plus   │
                             │  ours, plus a policy engine  │
                             └──────────────┬───────────────┘
                                            │ MCP
                             ┌──────────────▼───────────────┐
                             │  wrangler agent (ADK graph)  │
                             │  Gemini 3.8 Flash            │
                             └──────────────────────────────┘

mcp-gatecheck is one binary that imports github.com/grafana/mcp-grafana as a library and registers Grafana's own tool implementations unmodified alongside the render-farm tools we added, over a single MCP endpoint. The agent has no Prometheus client, no Loki client, no dashboard API wrapper and no auth plumbing of its own.

The graph:

START → triage → price → plan → self_check ─┬─ passed ─→ approval_gate → act → verify ─┐
                                            │                                          ├→ write_back
                                            └─ refused ─→ escalate ────────────────────┘

triage and price are model nodes with tools. plan, self_check, approval_gate, act and escalate are ordinary Python. Between a proposal and the render farm there are three independent refusals: a schema that rejects a malformed proposal, a guard in Grafana that can deny the call, and a server-side policy that re-checks the action after a human has approved it.


Repository layout

path what it is
farm/ a real OpenCue render farm in Docker, rendering real Blender frames
gaffer/ deterministic, seeded fault injection — the failures are real, not simulated
verifier/ the frame forensics. Pure NumPy and Pillow. No model, no network in the analysis path
mcp-gatecheck/ the MCP server: Grafana's tools, our tools, the policy engine, and an MCP App
agent/ the wrangler graph on ADK and Gemini
grafana/ guards and dashboards, provisioned as code
alloy/ telemetry collection into Grafana Cloud
eval/ the graded scorecard, and the answer key the agent cannot reach
console/ the operator's front door

Running it

cp .env.example .env      # fill in your Grafana and Google Cloud values
./scripts/farm.sh up      # OpenCue, Prometheus, Loki, two render nodes
./scripts/farm.sh submit  # render a shot, with seeded faults
./scripts/apply-guards.sh # install the guards into your Grafana tenant
adk web agent             # or: adk run agent

Full setup, including the Grafana Cloud free-tier prerequisites, is in docs/SETUP.md. A five-minute reviewer path is in JUDGES.md.


Is any of this fake?

MOCKS.md draws the line explicitly: what is real, what is standing in for something, and what was not built. The short version is that the farm, the frames, the faults, the detection, the telemetry and the guards are real, and the show is invented — there is no studio here, and deliberately no real film titles or logos anywhere.

Honest limits

  • The verifier finds frames that are statistically unlike their neighbours. A shot that is uniformly wrong from frame 1 — every frame equally miscoloured, with no clean neighbour to compare against — is invisible to it. It needs either a reference render or a human.
  • It will flag a legitimate hard cut or a deliberate strobe as suspect. The scorecard in eval/ reports that false-positive rate rather than hiding it.
  • The cost model uses published render-farm unit economics, not your studio's contract. The dollar figures are directionally right and specifically wrong.
  • p(success) for a retry comes from a fixed table keyed on defect class. It is a stated prior, not a learned estimate, and it is in the repository so you can disagree with it.
  • Agent Observability guards and score read-back are Grafana Cloud features. Against a self-hosted Grafana the agent runs, and the self-check correctly refuses to authorise anything, because it cannot prove it has been reliable.

Licence

Apache-2.0. See LICENSE.

About

A night-shift render wrangler: finds the frames a render farm reports as successful and proves are wrong, prices the fix, and can only act inside limits it cannot argue past.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages