Skip to content

Commit 134b488

Browse files
kivo360claude
andcommitted
docs(plan): agent end-to-end proof-of-life with Fireworks Kimi K2.5
Captures the next-session task: observe a real session go pending → running → succeeded with an agent making a real LLM call. Uses Fireworks (OpenAI-compatible) instead of Anthropic because we have free credits there + it's a useful regression test for the opencode_config_renderer's non-canonical-provider path. The Fireworks API key is set on Railway as FIREWORKS_API_KEY (not in the repo). Plan walks through: 1. Pre-flight (orchestrator + smoke account checks) 2. Renderer patch — teach opencode_config_renderer about Fireworks 3. Credential binding via SDK (broker path) 4. EnvironmentVersion bound to that credential 5. Session create with workspace + env_version 6. Stream events, watch the trajectory 7. Pull artifacts (bonus) 8. Hardening notes for follow-up runs Includes spec references (§3, §4, §5, §14, §18) and discipline notes captured from the user (terminal-first, no extras until proof-of-life lands). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
1 parent 2d60a1b commit 134b488

1 file changed

Lines changed: 238 additions & 0 deletions

File tree

‎tasks/agent-proof-of-life-plan.md‎

Lines changed: 238 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,238 @@
1+
# Agent End-to-End Proof of Life — Implementation Plan
2+
3+
**Created**: 2026-04-26
4+
**Status**: Plan
5+
**Goal**: Watch a real session go pending → running → succeeded with an
6+
agent making a real LLM call and producing observable output.
7+
**LLM**: Kimi K2.5 Turbo via Fireworks (OpenAI-compatible endpoint).
8+
9+
## Why this is the right next step
10+
11+
Bare-minimum platform shipped today (PASS 29 / FAIL 1 / GAP 0 / SKIP 5).
12+
Every primitive contract is wired. What we have NOT yet observed end-to-end
13+
is the closing-the-loop step: orchestrator picks pending → spawns sandbox →
14+
agent makes a real LLM call → emits events → finishes. Once that lands
15+
once on prod, every Ramp-style use case stops being theoretical.
16+
17+
## Non-goals for this run
18+
19+
- No GitHub integration yet (no PR creation, no clone). Just an LLM call
20+
that produces text output captured as an artifact or session message.
21+
- No multi-tenant tests — single smoke account, single workspace.
22+
- No frontend involvement — entirely terminal-driven, per the
23+
terminal-first principle.
24+
- No Daytona — Modal only (free credits, already wired).
25+
26+
## Pre-flight checks (run first)
27+
28+
```bash
29+
# 1. Orchestrator is on
30+
railway variables | grep -E "MONITORING_ORCHESTRATOR_ENABLED|SANDBOX_PROVIDER"
31+
# Expect: MONITORING_ORCHESTRATOR_ENABLED=true, SANDBOX_PROVIDER=modal
32+
33+
# 2. Smoke account credentials are fresh
34+
ls -la backend/.env.smoke-test
35+
# If older than ~1 hour: re-run setup_prod_smoke_account.py to refresh JWT
36+
37+
# 3. Fireworks key exists in Railway (set in this session, see step 1)
38+
railway variables | grep FIREWORKS_API_KEY
39+
```
40+
41+
## Step 1 — Stash the Fireworks key on Railway
42+
43+
The key the user provided lives in Railway, not in the repo. Set once:
44+
45+
```bash
46+
railway variables --set "FIREWORKS_API_KEY=<the key>"
47+
```
48+
49+
This is the env-var fallback for the bootstrap path. The PRIMARY credential
50+
delivery is via the broker — see step 3.
51+
52+
## Step 2 — Teach the renderer about Fireworks
53+
54+
`backend/omoi_os/services/opencode_config_renderer.py` knows about
55+
anthropic / openai / openrouter / google / groq / xai today. Add a
56+
`fireworks` entry:
57+
58+
```python
59+
"fireworks": {
60+
"npm": "@ai-sdk/openai-compatible", # OpenAI-compatible adapter
61+
"name": "Fireworks AI",
62+
"options": {
63+
"baseURL": "https://api.fireworks.ai/inference/v1",
64+
"apiKey": "{env:FIREWORKS_API_KEY}",
65+
},
66+
},
67+
```
68+
69+
And the model entry:
70+
71+
```python
72+
"fireworks": "fireworks/accounts/fireworks/routers/kimi-k2p5-turbo",
73+
```
74+
75+
Add `fireworks` to `_PREFERENCE_ORDER` *before* `anthropic` (so it's the
76+
default when present — that's the whole point of this run).
77+
78+
**Risk**: OpenCode's provider format may not handle slashes in model IDs
79+
cleanly. If it chokes, fall back to declaring an explicit `models` block
80+
in the provider entry. Verifiable from inside the sandbox by running
81+
`opencode --version` then checking what model it picks.
82+
83+
## Step 3 — Create the credential binding via SDK
84+
85+
This is the broker path — the secret never touches the repo.
86+
87+
```python
88+
import asyncio, os
89+
from omoios import AsyncOmoiOSClient
90+
91+
async def main():
92+
async with AsyncOmoiOSClient(
93+
base_url=os.environ["OMOIOS_API_BASE_URL"],
94+
api_key=os.environ["OMOIOS_PLATFORM_API_KEY"],
95+
) as c:
96+
binding = await c.credentials.create(
97+
workspace_id=os.environ["OMOIOS_TEST_WORKSPACE_A"],
98+
kind="bearer_secret",
99+
name="fireworks-kimi",
100+
value=os.environ["FIREWORKS_API_KEY"],
101+
)
102+
print(f"binding_id={binding.id}")
103+
104+
asyncio.run(main())
105+
```
106+
107+
## Step 4 — Create an EnvironmentVersion bound to that credential
108+
109+
There's no SDK method for environment-version yet (we have
110+
`environments` resource). The CLI for this should ship in next
111+
iteration — but for proof-of-life we can call the route directly:
112+
113+
```python
114+
async with c as client:
115+
env = await c.environments.create(
116+
organization_id=ORG_ID,
117+
name="kimi-fireworks-poc",
118+
)
119+
# Direct POST until SDK gets envversion support:
120+
r = await client._request("POST",
121+
f"/api/v1/environments/{env.id}/versions",
122+
json={
123+
"image": "nikolaik/python-nodejs:python3.12-nodejs22",
124+
"credentials": {
125+
"fireworks": {
126+
"kind": "bearer_secret",
127+
"binding_id": binding.id,
128+
}
129+
},
130+
}
131+
)
132+
env_version_id = r.json()["id"]
133+
```
134+
135+
The credential alias `fireworks` is what the bootstrap reads to render
136+
auth.json. The renderer (step 2) maps that alias to the OpenCode
137+
provider config, which uses `{env:FIREWORKS_API_KEY}` — but inside the
138+
sandbox the bootstrap actually rewrites this from the broker fetch (see
139+
`sandbox/bootstrap.sh:render_auth_entry`).
140+
141+
**Spec subtlety**: the bootstrap renders auth.json from broker data, NOT
142+
from env vars. OpenCode reads `auth.json` for keys regardless of what
143+
opencode.json says about `{env:FIREWORKS_API_KEY}`. So as long as
144+
`fireworks` is in `auth.json`, the LLM call works. Verify by exec'ing
145+
`cat ~/.local/share/opencode/auth.json` inside the sandbox — should show
146+
`{"fireworks": {"type": "api", "key": "Fw_..."}}`.
147+
148+
## Step 5 — Spawn the session with a real prompt
149+
150+
```python
151+
session = await c.sessions.create(
152+
workspace_id=os.environ["OMOIOS_TEST_WORKSPACE_A"],
153+
environment_id=env.id,
154+
prompt="Explain in 3 bullets how OpenCode finds its provider keys.",
155+
metadata={"source": "proof-of-life-2026-04-26"},
156+
)
157+
print(f"session_id={session.id}")
158+
```
159+
160+
## Step 6 — Stream events and watch the trajectory
161+
162+
```python
163+
async for evt in c.sessions.events(session.id):
164+
print(f" seq={evt.seq:>3} {evt.type:<24} actor={evt.actor}")
165+
if evt.type in ("session.succeeded", "session.failed", "session.cancelled"):
166+
print(f"\nTERMINAL: {evt.type}")
167+
break
168+
```
169+
170+
What we expect to see, in order:
171+
1. `session.created` — actor=user
172+
2. `session.started` — actor=agent (orchestrator picked up + sandbox spawned)
173+
3. `session.message` events from `actor=agent` carrying the LLM output
174+
chunks (token-by-token if streaming is wired) or full text
175+
4. `session.succeeded` — actor=system
176+
177+
If we see `session.failed` or `session.cancelled`, capture the
178+
`error_message` and the last few events for triage.
179+
180+
## Step 7 — Pull artifacts (if any)
181+
182+
```python
183+
artifacts = await c.sessions.artifacts(session.id)
184+
for a in artifacts:
185+
print(f" {a.name} ({a.size_bytes} bytes, {a.content_type})")
186+
```
187+
188+
For a pure-text response, OpenCode may or may not write an artifact —
189+
the `session.message` events should already contain the answer. Artifact
190+
verification is a bonus check, not a gate.
191+
192+
## Step 8 — Hardening before the next session
193+
194+
A single proof-of-life run validates the path. Before relying on it,
195+
fix anything that the run exposed:
196+
197+
- If OpenCode rejects `fireworks/accounts/...` model IDs → patch the
198+
renderer's model formatter
199+
- If bootstrap doesn't run → either invoke it via `sb.exec` post-spawn
200+
(Modal) or move auth.json render into the spawner directly (mirrors
201+
what we already did for opencode.json)
202+
- If the orchestrator picks up the session but never spawns a sandbox →
203+
trace through `orchestrator_worker.py` claim → spawn path; likely
204+
related to the spawner factory not seeing `SANDBOX_PROVIDER=modal`
205+
- If the agent runs but takes >10min → likely Modal cold start; bake
206+
OpenCode + OmO into the image rather than relying on the bootstrap
207+
to install them
208+
209+
## Hand-off package for next session
210+
211+
This file is the input. Next session should:
212+
213+
1. Confirm orchestrator + smoke account are still live (ScheduleWakeup
214+
usually keeps them warm; if not, re-run `setup_prod_smoke_account.py`)
215+
2. Set `FIREWORKS_API_KEY` on Railway
216+
3. Land the renderer patch from step 2 + push
217+
4. Run a single Python script that does steps 3-7 in sequence
218+
5. Capture the event stream + final status to
219+
`.sisyphus/evidence/agent-proof-of-life-<date>.json`
220+
6. Whatever fails, iterate from step 8
221+
222+
## Spec reference
223+
224+
- spec §3 sessions API (create / events / cancel / reply)
225+
- spec §4 broker (credential alias delivery to sandbox)
226+
- spec §5 environment versions (immutable credential binding)
227+
- spec §14 OmO config schema
228+
- spec §18 §3 primitive patterns (this is Pattern B — sync wait)
229+
230+
## Discipline notes (from the user)
231+
232+
- **Terminal first, UIs second.** Provider mgmt + GitHub auth + onboarding
233+
must work as CLI flows before anyone builds a frontend page for them.
234+
This proof-of-life script is itself the first such CLI.
235+
- **Bare minimum is shipped.** Resist building extras (Slack bot, Chrome
236+
extension, voice, etc.) until the agent loop has been observed end-to-
237+
end at least once. Stuff built on top of an unobserved loop is built on
238+
faith.

0 commit comments

Comments
 (0)