|
| 1 | +# Agent End-to-End Proof of Life — Implementation Plan |
| 2 | + |
| 3 | +**Created**: 2026-04-26 |
| 4 | +**Status**: Plan |
| 5 | +**Goal**: Watch a real session go pending → running → succeeded with an |
| 6 | +agent making a real LLM call and producing observable output. |
| 7 | +**LLM**: Kimi K2.5 Turbo via Fireworks (OpenAI-compatible endpoint). |
| 8 | + |
| 9 | +## Why this is the right next step |
| 10 | + |
| 11 | +Bare-minimum platform shipped today (PASS 29 / FAIL 1 / GAP 0 / SKIP 5). |
| 12 | +Every primitive contract is wired. What we have NOT yet observed end-to-end |
| 13 | +is the closing-the-loop step: orchestrator picks pending → spawns sandbox → |
| 14 | +agent makes a real LLM call → emits events → finishes. Once that lands |
| 15 | +once on prod, every Ramp-style use case stops being theoretical. |
| 16 | + |
| 17 | +## Non-goals for this run |
| 18 | + |
| 19 | +- No GitHub integration yet (no PR creation, no clone). Just an LLM call |
| 20 | + that produces text output captured as an artifact or session message. |
| 21 | +- No multi-tenant tests — single smoke account, single workspace. |
| 22 | +- No frontend involvement — entirely terminal-driven, per the |
| 23 | + terminal-first principle. |
| 24 | +- No Daytona — Modal only (free credits, already wired). |
| 25 | + |
| 26 | +## Pre-flight checks (run first) |
| 27 | + |
| 28 | +```bash |
| 29 | +# 1. Orchestrator is on |
| 30 | +railway variables | grep -E "MONITORING_ORCHESTRATOR_ENABLED|SANDBOX_PROVIDER" |
| 31 | +# Expect: MONITORING_ORCHESTRATOR_ENABLED=true, SANDBOX_PROVIDER=modal |
| 32 | + |
| 33 | +# 2. Smoke account credentials are fresh |
| 34 | +ls -la backend/.env.smoke-test |
| 35 | +# If older than ~1 hour: re-run setup_prod_smoke_account.py to refresh JWT |
| 36 | + |
| 37 | +# 3. Fireworks key exists in Railway (set in this session, see step 1) |
| 38 | +railway variables | grep FIREWORKS_API_KEY |
| 39 | +``` |
| 40 | + |
| 41 | +## Step 1 — Stash the Fireworks key on Railway |
| 42 | + |
| 43 | +The key the user provided lives in Railway, not in the repo. Set once: |
| 44 | + |
| 45 | +```bash |
| 46 | +railway variables --set "FIREWORKS_API_KEY=<the key>" |
| 47 | +``` |
| 48 | + |
| 49 | +This is the env-var fallback for the bootstrap path. The PRIMARY credential |
| 50 | +delivery is via the broker — see step 3. |
| 51 | + |
| 52 | +## Step 2 — Teach the renderer about Fireworks |
| 53 | + |
| 54 | +`backend/omoi_os/services/opencode_config_renderer.py` knows about |
| 55 | +anthropic / openai / openrouter / google / groq / xai today. Add a |
| 56 | +`fireworks` entry: |
| 57 | + |
| 58 | +```python |
| 59 | +"fireworks": { |
| 60 | + "npm": "@ai-sdk/openai-compatible", # OpenAI-compatible adapter |
| 61 | + "name": "Fireworks AI", |
| 62 | + "options": { |
| 63 | + "baseURL": "https://api.fireworks.ai/inference/v1", |
| 64 | + "apiKey": "{env:FIREWORKS_API_KEY}", |
| 65 | + }, |
| 66 | +}, |
| 67 | +``` |
| 68 | + |
| 69 | +And the model entry: |
| 70 | + |
| 71 | +```python |
| 72 | +"fireworks": "fireworks/accounts/fireworks/routers/kimi-k2p5-turbo", |
| 73 | +``` |
| 74 | + |
| 75 | +Add `fireworks` to `_PREFERENCE_ORDER` *before* `anthropic` (so it's the |
| 76 | +default when present — that's the whole point of this run). |
| 77 | + |
| 78 | +**Risk**: OpenCode's provider format may not handle slashes in model IDs |
| 79 | +cleanly. If it chokes, fall back to declaring an explicit `models` block |
| 80 | +in the provider entry. Verifiable from inside the sandbox by running |
| 81 | +`opencode --version` then checking what model it picks. |
| 82 | + |
| 83 | +## Step 3 — Create the credential binding via SDK |
| 84 | + |
| 85 | +This is the broker path — the secret never touches the repo. |
| 86 | + |
| 87 | +```python |
| 88 | +import asyncio, os |
| 89 | +from omoios import AsyncOmoiOSClient |
| 90 | + |
| 91 | +async def main(): |
| 92 | + async with AsyncOmoiOSClient( |
| 93 | + base_url=os.environ["OMOIOS_API_BASE_URL"], |
| 94 | + api_key=os.environ["OMOIOS_PLATFORM_API_KEY"], |
| 95 | + ) as c: |
| 96 | + binding = await c.credentials.create( |
| 97 | + workspace_id=os.environ["OMOIOS_TEST_WORKSPACE_A"], |
| 98 | + kind="bearer_secret", |
| 99 | + name="fireworks-kimi", |
| 100 | + value=os.environ["FIREWORKS_API_KEY"], |
| 101 | + ) |
| 102 | + print(f"binding_id={binding.id}") |
| 103 | + |
| 104 | +asyncio.run(main()) |
| 105 | +``` |
| 106 | + |
| 107 | +## Step 4 — Create an EnvironmentVersion bound to that credential |
| 108 | + |
| 109 | +There's no SDK method for environment-version yet (we have |
| 110 | +`environments` resource). The CLI for this should ship in next |
| 111 | +iteration — but for proof-of-life we can call the route directly: |
| 112 | + |
| 113 | +```python |
| 114 | +async with c as client: |
| 115 | + env = await c.environments.create( |
| 116 | + organization_id=ORG_ID, |
| 117 | + name="kimi-fireworks-poc", |
| 118 | + ) |
| 119 | + # Direct POST until SDK gets envversion support: |
| 120 | + r = await client._request("POST", |
| 121 | + f"/api/v1/environments/{env.id}/versions", |
| 122 | + json={ |
| 123 | + "image": "nikolaik/python-nodejs:python3.12-nodejs22", |
| 124 | + "credentials": { |
| 125 | + "fireworks": { |
| 126 | + "kind": "bearer_secret", |
| 127 | + "binding_id": binding.id, |
| 128 | + } |
| 129 | + }, |
| 130 | + } |
| 131 | + ) |
| 132 | + env_version_id = r.json()["id"] |
| 133 | +``` |
| 134 | + |
| 135 | +The credential alias `fireworks` is what the bootstrap reads to render |
| 136 | +auth.json. The renderer (step 2) maps that alias to the OpenCode |
| 137 | +provider config, which uses `{env:FIREWORKS_API_KEY}` — but inside the |
| 138 | +sandbox the bootstrap actually rewrites this from the broker fetch (see |
| 139 | +`sandbox/bootstrap.sh:render_auth_entry`). |
| 140 | + |
| 141 | +**Spec subtlety**: the bootstrap renders auth.json from broker data, NOT |
| 142 | +from env vars. OpenCode reads `auth.json` for keys regardless of what |
| 143 | +opencode.json says about `{env:FIREWORKS_API_KEY}`. So as long as |
| 144 | +`fireworks` is in `auth.json`, the LLM call works. Verify by exec'ing |
| 145 | +`cat ~/.local/share/opencode/auth.json` inside the sandbox — should show |
| 146 | +`{"fireworks": {"type": "api", "key": "Fw_..."}}`. |
| 147 | + |
| 148 | +## Step 5 — Spawn the session with a real prompt |
| 149 | + |
| 150 | +```python |
| 151 | +session = await c.sessions.create( |
| 152 | + workspace_id=os.environ["OMOIOS_TEST_WORKSPACE_A"], |
| 153 | + environment_id=env.id, |
| 154 | + prompt="Explain in 3 bullets how OpenCode finds its provider keys.", |
| 155 | + metadata={"source": "proof-of-life-2026-04-26"}, |
| 156 | +) |
| 157 | +print(f"session_id={session.id}") |
| 158 | +``` |
| 159 | + |
| 160 | +## Step 6 — Stream events and watch the trajectory |
| 161 | + |
| 162 | +```python |
| 163 | +async for evt in c.sessions.events(session.id): |
| 164 | + print(f" seq={evt.seq:>3} {evt.type:<24} actor={evt.actor}") |
| 165 | + if evt.type in ("session.succeeded", "session.failed", "session.cancelled"): |
| 166 | + print(f"\nTERMINAL: {evt.type}") |
| 167 | + break |
| 168 | +``` |
| 169 | + |
| 170 | +What we expect to see, in order: |
| 171 | +1. `session.created` — actor=user |
| 172 | +2. `session.started` — actor=agent (orchestrator picked up + sandbox spawned) |
| 173 | +3. `session.message` events from `actor=agent` carrying the LLM output |
| 174 | + chunks (token-by-token if streaming is wired) or full text |
| 175 | +4. `session.succeeded` — actor=system |
| 176 | + |
| 177 | +If we see `session.failed` or `session.cancelled`, capture the |
| 178 | +`error_message` and the last few events for triage. |
| 179 | + |
| 180 | +## Step 7 — Pull artifacts (if any) |
| 181 | + |
| 182 | +```python |
| 183 | +artifacts = await c.sessions.artifacts(session.id) |
| 184 | +for a in artifacts: |
| 185 | + print(f" {a.name} ({a.size_bytes} bytes, {a.content_type})") |
| 186 | +``` |
| 187 | + |
| 188 | +For a pure-text response, OpenCode may or may not write an artifact — |
| 189 | +the `session.message` events should already contain the answer. Artifact |
| 190 | +verification is a bonus check, not a gate. |
| 191 | + |
| 192 | +## Step 8 — Hardening before the next session |
| 193 | + |
| 194 | +A single proof-of-life run validates the path. Before relying on it, |
| 195 | +fix anything that the run exposed: |
| 196 | + |
| 197 | +- If OpenCode rejects `fireworks/accounts/...` model IDs → patch the |
| 198 | + renderer's model formatter |
| 199 | +- If bootstrap doesn't run → either invoke it via `sb.exec` post-spawn |
| 200 | + (Modal) or move auth.json render into the spawner directly (mirrors |
| 201 | + what we already did for opencode.json) |
| 202 | +- If the orchestrator picks up the session but never spawns a sandbox → |
| 203 | + trace through `orchestrator_worker.py` claim → spawn path; likely |
| 204 | + related to the spawner factory not seeing `SANDBOX_PROVIDER=modal` |
| 205 | +- If the agent runs but takes >10min → likely Modal cold start; bake |
| 206 | + OpenCode + OmO into the image rather than relying on the bootstrap |
| 207 | + to install them |
| 208 | + |
| 209 | +## Hand-off package for next session |
| 210 | + |
| 211 | +This file is the input. Next session should: |
| 212 | + |
| 213 | +1. Confirm orchestrator + smoke account are still live (ScheduleWakeup |
| 214 | + usually keeps them warm; if not, re-run `setup_prod_smoke_account.py`) |
| 215 | +2. Set `FIREWORKS_API_KEY` on Railway |
| 216 | +3. Land the renderer patch from step 2 + push |
| 217 | +4. Run a single Python script that does steps 3-7 in sequence |
| 218 | +5. Capture the event stream + final status to |
| 219 | + `.sisyphus/evidence/agent-proof-of-life-<date>.json` |
| 220 | +6. Whatever fails, iterate from step 8 |
| 221 | + |
| 222 | +## Spec reference |
| 223 | + |
| 224 | +- spec §3 sessions API (create / events / cancel / reply) |
| 225 | +- spec §4 broker (credential alias delivery to sandbox) |
| 226 | +- spec §5 environment versions (immutable credential binding) |
| 227 | +- spec §14 OmO config schema |
| 228 | +- spec §18 §3 primitive patterns (this is Pattern B — sync wait) |
| 229 | + |
| 230 | +## Discipline notes (from the user) |
| 231 | + |
| 232 | +- **Terminal first, UIs second.** Provider mgmt + GitHub auth + onboarding |
| 233 | + must work as CLI flows before anyone builds a frontend page for them. |
| 234 | + This proof-of-life script is itself the first such CLI. |
| 235 | +- **Bare minimum is shipped.** Resist building extras (Slack bot, Chrome |
| 236 | + extension, voice, etc.) until the agent loop has been observed end-to- |
| 237 | + end at least once. Stuff built on top of an unobserved loop is built on |
| 238 | + faith. |
0 commit comments