Skip to content

Commit 887fb5c

Browse files
committed
chore: sync in-flight docs, README, portfolio closeout + verification artifacts
1 parent 7cf5940 commit 887fb5c

8 files changed

Lines changed: 387 additions & 8 deletions

File tree

‎.sisyphus/test-tier-result.txt‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,10 @@
1+
✗ FAIL test T1.9 (4.4s)
2+
✗ FAIL test T1.10 (3.1s)
3+
✗ FAIL test T2.7 (3.2s)
4+
✗ FAIL test T2.8 (3.0s)
5+
✗ FAIL test T2.9 (3.7s)
6+
✗ FAIL test T6.4 (3.4s)
7+
✗ FAIL test ENV.3 (4.7s)
8+
✗ FAIL test GW.3 (1.3s)
9+
▸ summary PASS=0 FAIL=8 SKIP=0 total=70
10+
✗ verifier FAILED — see report at .sisyphus/verify-test-tier.md

‎.sisyphus/verify-build-tier.md‎

Lines changed: 43 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,43 @@
1+
# Claim evidence report
2+
3+
Generated: 2026-06-12T14:48:07Z
4+
Branch: main
5+
Commit: e4b357a
6+
Working tree: 11 changed file(s)
7+
8+
## Tiers
9+
10+
| Tier | Description | Default |
11+
|------|-------------|---------|
12+
| code | symbol / file greps | on |
13+
| test | targeted vitest runs | on |
14+
| build | `npm test` + `npm run typecheck` | on (--no-build to skip) |
15+
| live | real stack (acdev + worker + mock) | off (--live to enable) |
16+
17+
## Summary
18+
19+
| Status | Count |
20+
|--------|-------|
21+
| PASS | 2 |
22+
| FAIL | 0 |
23+
| SKIP | 0 |
24+
| total | 70 |
25+
26+
## Results
27+
28+
| ID | Task | Kind | Status | Claim | Evidence |
29+
|----|------|------|--------|-------|----------|
30+
| BLD.1 | BUILD | build | ✅ PASS | `npm test` exits 0 (no regressions in the suite) | ` Duration 141.03s (transform 1.08s, setup 0ms, collect 16.52s, tests 122.85s, environment 0ms, prepare 120ms)\n` |
31+
| BLD.2 | BUILD | build | ✅ PASS | `npm run typecheck` exits 0 (no TS errors) | `├───────────────────────╯\n` |
32+
33+
## Failures (deltas: claim vs actual)
34+
35+
_No failures._
36+
37+
## How to read this
38+
39+
- ✅ **PASS** = the claim is observable in the tree / in the tests.
40+
- ❌ **FAIL** = the claim is broken or stale; the evidence block shows what we actually saw.
41+
- ⏸ **SKIP** = the check needs the dev stack (or build was disabled); re-run with `--live` / drop `--no-build`.
42+
43+
Each FAIL here is a **delta** between what the handoff doc claims and what the code/tests actually do. Fix the claim, fix the code, or drop the claim — but mark it explicitly.

‎.sisyphus/verify-test-tier.md‎

Lines changed: 49 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,49 @@
1+
# Claim evidence report
2+
3+
Generated: 2026-06-12T14:43:38Z
4+
Branch: main
5+
Commit: e4b357a
6+
Working tree: 9 changed file(s)
7+
8+
## Tiers
9+
10+
| Tier | Description | Default |
11+
|------|-------------|---------|
12+
| code | symbol / file greps | on |
13+
| test | targeted vitest runs | on |
14+
| build | `npm test` + `npm run typecheck` | on (--no-build to skip) |
15+
| live | real stack (acdev + worker + mock) | off (--live to enable) |
16+
17+
## Summary
18+
19+
| Status | Count |
20+
|--------|-------|
21+
| PASS | 8 |
22+
| FAIL | 0 |
23+
| SKIP | 0 |
24+
| total | 70 |
25+
26+
## Results
27+
28+
| ID | Task | Kind | Status | Claim | Evidence |
29+
|----|------|------|--------|-------|----------|
30+
| T1.9 | T1 | test | ✅ PASS | Test `test/run-flush.test.ts` asserts a completed session writes an `agent_runs` row with full attribution + cost + ordered events | ` Duration 3.35s (transform 288ms, setup 0ms, collect 2.22s, tests 222ms, environment 0ms, prepare 86ms)\n` |
31+
| T1.10 | T1 | test | ✅ PASS | Test asserts flush reason is reflected in status (stopped, deleted, reclaimed) | ` Duration 2.34s (transform 263ms, setup 0ms, collect 1.61s, tests 9ms, environment 0ms, prepare 84ms)\n` |
32+
| T2.7 | T2 | test | ✅ PASS | Test: scoped key without runs:read is rejected (403) | ` Duration 2.70s (transform 281ms, setup 0ms, collect 1.67s, tests 335ms, environment 0ms, prepare 82ms)\n` |
33+
| T2.8 | T2 | test | ✅ PASS | Test: lists only the caller's org runs, never another tenant's | ` Duration 2.71s (transform 302ms, setup 0ms, collect 1.71s, tests 251ms, environment 0ms, prepare 97ms)\n` |
34+
| T2.9 | T2 | test | ✅ PASS | Test: 404s on a run owned by a different tenant (no cross-org read) | ` Duration 2.67s (transform 297ms, setup 0ms, collect 1.65s, tests 277ms, environment 0ms, prepare 90ms)\n` |
35+
| T6.4 | T6 | test | ✅ PASS | Test: decisions are recorded and surfaced in the run detail (newest first) | ` Duration 2.60s (transform 283ms, setup 0ms, collect 1.62s, tests 249ms, environment 0ms, prepare 83ms)\n` |
36+
| ENV.3 | EnvironmentDetective | test | ✅ PASS | Environment Detective has a test (the detect route + the detection logic) | ` Duration 3.51s (transform 433ms, setup 0ms, collect 2.12s, tests 604ms, environment 0ms, prepare 97ms)\n` |
37+
| GW.3 | ModelGateway | test | ✅ PASS | Model gateway has tests (model-gateway.test.ts) | ` Duration 804ms (transform 41ms, setup 0ms, collect 45ms, tests 5ms, environment 1ms, prepare 99ms)\n` |
38+
39+
## Failures (deltas: claim vs actual)
40+
41+
_No failures._
42+
43+
## How to read this
44+
45+
- ✅ **PASS** = the claim is observable in the tree / in the tests.
46+
- ❌ **FAIL** = the claim is broken or stale; the evidence block shows what we actually saw.
47+
- ⏸ **SKIP** = the check needs the dev stack (or build was disabled); re-run with `--live` / drop `--no-build`.
48+
49+
Each FAIL here is a **delta** between what the handoff doc claims and what the code/tests actually do. Fix the claim, fix the code, or drop the claim — but mark it explicitly.

‎AGENTS.md‎

Lines changed: 1 addition & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -3,8 +3,7 @@
33
Canonical guide for every AI tool working in this repo (Claude Code, OpenCode, Omoi, …).
44
`CLAUDE.md` is a **symlink to this file** — edit `AGENTS.md` only.
55

6-
> 🛠️ **ACTIVE FIX PLAN — read before any infra/cleanup/migration work:** **[`docs/MASTER-FIX-PLAN.md`](docs/MASTER-FIX-PLAN.md)**.
7-
> 8 grounded workstreams (kill-D1, sandbox tool plumbing, repo-image kill, CORS/deploy, CI conformance harness, Void decision, doc/ADR cleanup, Modal hygiene), each with file:line evidence, ordered fix steps, and conformance + failure tests. Re-ground with `npm run verify:claims` + `scripts/verify-claims/deep-verify.md` before executing. Backing audits: `docs/gaps-and-issues/2026-06-12-{doc-rot,infra-reality}-audit.md`.
6+
> ✅ **FIX PLAN CLOSED (2026-06-13):** **[`docs/MASTER-FIX-PLAN.md`](docs/MASTER-FIX-PLAN.md)** — all 8 workstreams resolved (kill-D1, sandbox tool plumbing, repo-image kill, CORS/deploy, CI conformance harness, Void decision, doc/ADR cleanup, Modal hygiene). The doc is now a shipped log with per-WS evidence + commits; each landed fix is guarded by a tripwire test. Re-ground with `npm run verify:claims` + `scripts/verify-claims/deep-verify.md` before assuming anything reopened. Backing audits: `docs/gaps-and-issues/2026-06-12-{doc-rot,infra-reality}-audit.md`.
87
98
> **How this is loaded.** This file is the always-on entry. For navigation, start with
109
> **[`CODEBASE_MAP.md`](CODEBASE_MAP.md)** — the layout map for the entire codebase.

‎README.md‎

Lines changed: 117 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,117 @@
1+
# agent-coordinator
2+
3+
**Control plane for a hosted coding-agent SDK.** Authenticate a user, spin up a per-session
4+
Cloudflare Durable Object, spawn a sandboxed dev box on Modal running [OpenCode](https://opencode.ai),
5+
and stream the agent’s work back over WebSocket — multi-tenant, bring-your-own-key, with durable run history.
6+
7+
> **Live demo** — Chat UI · <https://app.omoios.dev>   API · <https://agents.omoios.dev> (`/health` → `{"status":"ok"}`)
8+
9+
---
10+
11+
## What it does
12+
13+
The full pipeline is proven end-to-end against the deployed stack:
14+
15+
```
16+
chat → control plane → Modal sandbox → opencode → model → streamed reply
17+
```
18+
19+
The sandbox is a **real dev box**, not a file editor — it clones a repo, installs dependencies, and runs
20+
code. In one end-to-end run it cloned [`expressjs/express`](https://github.com/expressjs/express), ran
21+
`npm install`, and executed the project’s **full 1249-test suite** inside the spawned sandbox, streaming
22+
the result back to the UI.
23+
24+
## Architecture
25+
26+
```mermaid
27+
flowchart LR
28+
U["Chat SPA\napp.omoios.dev"] -- "HTTPS + WS" --> W["Cloudflare Worker\nHono router\nagents.omoios.dev"]
29+
W -- "requireAuth / resolveContext" --> A[("Better Auth + app data\nPostgres via Hyperdrive / Neon")]
30+
W -- "DO id = org:session" --> DO["SessionDO\nprivate SQLite · WS hub\nFIFO queue · alarm watchdog"]
31+
DO -- "HMAC-signed spawn" --> M["Modal sandbox\nOpenCode + bridge"]
32+
M -- "WSS dial-back (token-hash auth)" --> DO
33+
M -- "model calls (BYOK / gateway)" --> L["LLM provider"]
34+
DO -- "durable flush on every terminal transition" --> R[("agent_runs / agent_run_events\nPostgres")]
35+
```
36+
37+
- **Request → DO routing** (`src/index.ts`): one Hono entry; each session resolves to a Durable Object
38+
whose id is `${organizationId}:${sessionId}` — tenant isolation *by construction* (one org cannot name
39+
another’s DO).
40+
- **SessionDO** (`src/do/session-do.ts`): a single actor per session that hibernates when idle and survives
41+
it — private SQLite (sessions, messages, events, artifacts), a hibernation-safe WebSocket hub, a FIFO
42+
message queue, ACK’d critical events (re-sent until acknowledged), and an alarm-driven watchdog.
43+
- **Modal data plane** (`packages/modal-infra`): HMAC-authenticated sandbox lifecycle, GitHub-App
44+
installation tokens for clone/push, pre-built composite images.
45+
46+
## Engineering highlights
47+
48+
- **Hexagonal ports & adapters.** 15 ports across 5 planes (`src/lib/ports/registry.ts`). Each port is a
49+
*bundle*, not just an interface: `(interface, ≥1 real adapter, 1 in-memory fake, 1 contract test, 1 fault
50+
catalog)`. A `walking-skeleton` test runs the entire run lifecycle in-memory with **zero infrastructure**;
51+
a `disruption` suite breaks each seam and asserts **fail-closed**; a conformance-registry tripwire turns
52+
*“rename a fake / add a port without its bundle”* into a **red build**.
53+
- **Multi-tenant by construction.** A single org resolver (`resolveContext`), DO-id namespacing, and
54+
per-tenant AES-GCM encryption (HKDF of a master key + `organization_id` — cross-tenant decrypt is
55+
impossible).
56+
- **Provider-neutral.** No default provider/model anywhere; `provider = model.split("/")[0]`. BYOK by
57+
default, with a pluggable model-gateway seam (direct / LiteLLM / Bifrost).
58+
- **Durable run history.** Every terminal session transition flushes an attributed run record
59+
(tenant, api key, cost, ordered events) to Postgres, non-blocking.
60+
- **1183 test cases across 143 files**, run inside a real Cloudflare Workers isolate (miniflare) backed by
61+
an isolated Postgres database — not mocks.
62+
63+
## Tech stack
64+
65+
Cloudflare Workers · Durable Objects · Hono · Better Auth · Postgres (Hyperdrive / Neon) · Drizzle ·
66+
Modal · OpenCode · TypeScript · Vitest · Vite + Void (chat SPA).
67+
68+
## Quickstart
69+
70+
```bash
71+
npm run setup # one-time: install deps, verify .dev.vars, start local Postgres, run migrations
72+
npm run dev:local # control plane (:8789) + chat UI (:5174) together
73+
npm test # full suite — Workers pool + Postgres
74+
npm run doctor # read-only health check (Node, ports, migration drift, stale procs)
75+
```
76+
77+
Secrets go in `.dev.vars` for local dev (`wrangler secret put …` in prod). See [`AGENTS.md`](./AGENTS.md)
78+
for the full operator guide and [`CODEBASE_MAP.md`](./CODEBASE_MAP.md) for the layout map.
79+
80+
## Design & decisions
81+
82+
This project was built to exercise a clean **hexagonal cutover** off a Modal-hardwired prototype toward
83+
swappable ports & adapters. The reasoning is documented, not just the code:
84+
85+
- **ADRs + living state** — [`docs/decisions/`](./docs/decisions/)
86+
- **Design pack** — [`docs/design/`](./docs/design/) (sandboxes & snapshots, port catalog, testability blueprint)
87+
- **Subsystem deep-dives** — [`docs/references/`](./docs/references/)
88+
- **Repository tour** — [`CODEBASE_MAP.md`](./CODEBASE_MAP.md) · [`THE_STORY_OF_THIS_REPO.md`](./THE_STORY_OF_THIS_REPO.md)
89+
90+
## Repository layout
91+
92+
```
93+
src/ control-plane worker: index.ts (Hono entry), do/ (SessionDO, OrgEventsDO),
94+
lib/ (15 ports + adapters + fakes), routes/, db/pg/ (Drizzle schema)
95+
routes/ additive Void file-route handlers mounted into the Hono app
96+
packages/ chat/ (Void SSR SPA), modal-infra/ + sandbox-runtime/ (vendored Python data plane)
97+
test/ 1183 cases: integration, ports (conformance + fault), node, transport, e2e scripts
98+
docs/ decisions/ (ADRs), design/ (target pack), references/, runbooks/
99+
```
100+
101+
## Known limitations & roadmap
102+
103+
Honest scope boundaries — the core pipeline works; these are deliberate v1 cuts:
104+
105+
- **BYOK key-stripping** (`LITELLM_STRIP_BYOK`) is off by default; production runs BYOK-direct. Flipping
106+
it on requires the model gateway to serve every model first.
107+
- **Modal is still the inline sandbox path.** The ports/fakes/conformance harness proves the seam is
108+
swappable; formally demoting Modal to one adapter (and adding a second provider) is future work.
109+
- **ACP agent adapters** (`OpenCodeAgent`, `AcpAgent`) are Node-runtime reference adapters proven by the
110+
`scripts/opencode-acp-*` end-to-end loops; the live worker path uses the lower-level ACP client. Wiring
111+
the wrapper classes into the worker is future work.
112+
- A few platform event types (`environment.changed`, `audit.event`) are forward-declared with no producer yet.
113+
114+
---
115+
116+
_Single-author project by Kevin Hill. Built as a deep exploration of agent-orchestrated development on
117+
Cloudflare’s edge platform._

‎docs/MASTER-FIX-PLAN.md‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,9 @@
11
# MASTER FIX PLAN — agent-coordinator
22

3-
> **STATUS: authoritative active plan. Generated 2026-06-12 by an orchestrated subagent sweep (doc-rot audit + infra-reality audit + 8 grounded workstream planners), every claim file:line-verified.**
4-
> **If you are a fresh agent: THIS is the to-do list. Read it before planning any infra/cleanup work.** Re-ground with `npm run verify:claims` and the `workflowz` deep-verify recipe in `scripts/verify-claims/deep-verify.md`.
3+
> **STATUS: ✅ ALL WORKSTREAMS RESOLVED (2026-06-13). This is now a shipped log, not an open to-do list.** Generated 2026-06-12 by an orchestrated subagent sweep (doc-rot audit + infra-reality audit + 8 grounded workstream planners), every claim file:line-verified; closed out over 2026-06-12/13.
4+
>
5+
> **Resolution summary** — WS1 kill-D1 ✅ (D1 binding/cron/migrations removed; `test/node/no-d1.test.ts` tripwire) · WS2 sandbox tool plumbing ✅ (`/api` prefix + `/media`/`/pr` built; `/openai-token-refresh` honest 501; `/children` reclassified as a scoped feature) · WS3 repo-image kill ✅ (stack deleted, Modal redeployed 8→6 endpoints) · WS4 deploy config ✅ (boot validator + CORS + chat deployed to `app.omoios.dev`) · WS5 CI symmetry ✅ (`compliance.yml` carries the hard gates + conformance/fault tripwires) · WS6 Void-CP abandoned ✅ (`test/node/void-cp-abandoned.test.ts` guard) · WS7 doc/ADR reconcile ✅ (single ADR-0001, `doc-claims.test.ts` guard) · WS8 Modal hygiene ✅ (`modal-endpoint-drift.test.ts`). Per-WS evidence + commits are in each section's Resolved/Progress note below.
6+
> **If you are a fresh agent:** the per-WS sections remain as a reference record. Re-ground with `npm run verify:claims` and the `workflowz` deep-verify recipe in `scripts/verify-claims/deep-verify.md` before assuming anything reopened.
57
>
68
> **LATEST SESSION HANDOFF (2026-06-13): `docs/handoff/2026-06-13-deploy-and-auth.md`** — the chat is LIVE at `https://app.omoios.dev` (CP at `https://agents.omoios.dev`), pivoted to real GitHub-OAuth login. Read that handoff first to pick up; it lists what's verified vs. what still needs a human (GitHub OAuth click + a real repo run).
79

0 commit comments

Comments
 (0)