Skip to content

Commit 2fcffb1

Browse files
feat: day 3 agent loop - llm chat+fallback, /agent/run API, AI-free tests
1 parent 4fdc8eb commit 2fcffb1

13 files changed

Lines changed: 917 additions & 1 deletion

‎docs/day3-agent-loop.md‎

Lines changed: 85 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,85 @@
1+
# Day 3 — Agent loop: line-level report
2+
3+
Standing-rule report: what, why, what changed, and exactly where (file:lines).
4+
5+
## What
6+
7+
Built the agent's brain (the loop), upgraded the LLM "phone", exposed the loop
8+
as an HTTP API, pinned it with AI-free tests, and proved it live. The system
9+
now goes: task → model thinks → real tool call in a Docker sandbox → answer —
10+
all visible step by step.
11+
12+
## Why
13+
14+
- **The loop is the product:** everything after this (sessions, tools, memory,
15+
the frontend) hangs off `run_agent`. Day 3 made it real and testable.
16+
- **Resilience planted early:** fallback provider (Gemini/etc.) if OpenRouter
17+
dies; `achat` keeps the event loop responsive; max-10-step guard is the
18+
robot's seatbelt.
19+
- **Traceability = future observability:** every step is returned in the API
20+
response — Day 45's replays, forked from Day 3's JSON.
21+
22+
## What changed (file → lines)
23+
24+
### `orchestrator/agent.py` (new, 111 lines) — the loop
25+
| Lines | Content |
26+
|---|---|
27+
| 8 | `MAX_STEPS = 10` guard |
28+
| 10-18 | system prompt: sandbox rules, tool, `<answer>`/`<bash>` contract |
29+
| 20-33 | `run_bash` tool schema |
30+
| 36-45 | `_summarize`: command output → short context for the model |
31+
| 48-111 | `run_agent`: message loop, tool execution, `<bash>` fallback, `<answer>` extraction, max-steps return |
32+
33+
### `orchestrator/llm.py` (edited) — the phone
34+
| Lines | Content |
35+
|---|---|
36+
| 1 | `asyncio` import |
37+
| 21-23 | fallback provider settings (`LLM_FALLBACK_*`) |
38+
| 26-43 | `_completion` (shared provider call) |
39+
| 59-88 | `chat`: full message with `tool_calls` + automatic fallback |
40+
| 91-96 | `achat`: async twin, runs LLM call off the event loop |
41+
| 46-56 | `ask_llm` unchanged (ping) |
42+
43+
### `orchestrator/main.py` (edited) — the door
44+
| Lines | Content |
45+
|---|---|
46+
| 5 | `from agent import run_agent` |
47+
| 18-19 | `AgentRunRequest {task}` |
48+
| 47-51 | `POST /agent/run`: 422 on empty task, else run and return `{ok, answer, steps}` |
49+
50+
### Tests (new) — no AI, no network, no Docker
51+
| File | Lines | Covers |
52+
|---|---|---|
53+
| `test_agent.py` | 22-34 | immediate answer (0 tools) |
54+
| `test_agent.py` | 37-66 | real tool call → step recorded |
55+
| `test_agent.py` | 69-101 | `<bash>` tag fallback |
56+
| `test_agent.py` | 103-121 | max-steps guard fires |
57+
| `test_llm.py` | 1-81 | fallback routing + achat (4 tests) |
58+
| `test_agent_api.py` | 1-29 | `/agent/run` wiring + 422 |
59+
60+
### Docs (new)
61+
| File | Topic |
62+
|---|---|
63+
| `docs/day3-step1-agent-loop.md` | step 1 report |
64+
| `docs/day3-step2-llm-upgrade.md` | step 2 report |
65+
| `docs/day3-step3-agent-run.md` | step 3 report |
66+
| `docs/day3-step4-agent-tests.md` | step 4 report |
67+
| `docs/day3-step5-live-demo.md` | live demo + reproduce |
68+
| `docs/day3-agent-loop.md` | this report |
69+
70+
## Verified
71+
72+
- `ruff check` / `ruff format --check` — clean
73+
- `mypy` — clean
74+
- `pytest` — **13 passed** (no network, no Docker)
75+
- Live (Step 5): 2-step tool run (`ls -la /` → `python3 --version` → answer)
76+
and 0-tool run (`2+2` → `4`); 4 free-tier calls
77+
- CI (GitHub Actions): orchestrator ruff+mypy+pytest, gateway clippy+test,
78+
frontend build
79+
80+
## How it works in one sentence
81+
82+
`POST /agent/run` takes a task; `run_agent` loops: model (via `llm.achat`)
83+
thinks → if it asks, `sandbox.sandbox_exec` runs the command in the Docker
84+
sandbox and the result is fed back → until an `<answer>` appears or the
85+
10-step guard trips; every tool call is kept in `steps[]` and returned as-is.

‎docs/day3-step1-agent-loop.md‎

Lines changed: 102 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,102 @@
1+
# Day 3, Step 1 - The agent loop engine (`orchestrator/agent.py`)
2+
3+
The brain learns to press its own button. This file turns the orchestrator from a
4+
remote-controlled robot into one that thinks: decide a command -> run it in the sandbox ->
5+
read the output -> repeat -> answer.
6+
7+
## The topic (ELI5)
8+
9+
An **agent loop** is: THINK -> ACT -> LOOK -> THINK -> ... -> DONE.
10+
11+
- THINK: the LLM decides the next shell command (or says "I'm done")
12+
- ACT: we run that command in the sandbox (the `/exec` button from Step 2)
13+
- LOOK: we show the output back to the LLM
14+
- Repeat until the LLM answers, or we hit the safety limit (10 steps)
15+
16+
The LLM can press the button two ways:
17+
1. **Tool calling (primary):** the LLM returns a structured
18+
`{"name": "run_bash", "arguments": {"cmd": "ls -la /"}}` message (the industry-standard
19+
way - what Devin / Claude Code do).
20+
2. **Tag language (fallback):** some free models can't do tool calling, so we also teach:
21+
`<bash>command</bash>` = press the button, `<answer>text</answer>` = I'm done.
22+
Either way, the loop works. This resilience is the interview-gold part.
23+
24+
## What I did, file by file (with line ranges)
25+
26+
### NEW `orchestrator/agent.py` (111 lines) - the loop engine
27+
- Lines 1-6: imports. Note lines 5-6: `import llm` and `import sandbox` (module imports,
28+
NOT `from x import y`). Why: tests monkeypatch `llm.chat` and `sandbox.sandbox_exec`;
29+
module-style imports let the patch take effect at call time. (Learned the hard way -
30+
function imports silently ignored the patch and hit the real network.)
31+
- Line 8: `MAX_STEPS = 10` - the "unplug the robot" limit.
32+
- Lines 10-18: `SYSTEM_PROMPT` - the rules of the game: one tool, network limited to
33+
GitHub/PyPI/npm, prefer short commands, `<answer>` protocol, `<bash>` fallback.
34+
- Lines 20-33: `BASH_TOOL` - the JSON "button spec" in OpenAI tool-calling format
35+
(`type: function`, name `run_bash`, one required arg `cmd`).
36+
- Lines 36-45: `_summarize()` - turns an ExecResult into a short readable string for the LLM
37+
(handles timeouts, stderr-only, normal stdout, no output).
38+
- Lines 48-53: `run_agent(task, max_steps=10)` - builds the message history
39+
`[system, user task]` and an empty `steps` trace log.
40+
- Lines 55-59: the loop head - one LLM call per iteration via `llm.chat(messages, tools=...)`.
41+
`# noqa: BLE001` is intentional: LLM providers throw a zoo of exception types (verified
42+
live - no common base class exists), and the loop must survive any of them.
43+
- Lines 62-89: **tool-calling path** - if the model returned `tool_calls`, parse the JSON
44+
args (bad JSON -> empty cmd -> a helpful error result), run it with `sandbox.sandbox_exec`,
45+
record the step, append the result as a `role: "tool"` message, continue the loop.
46+
- Lines 91-99: **tag fallback path** - no tool_calls but `<bash>cmd</bash>` present: run the
47+
last one, append the output as a plain `user` message, continue.
48+
- Lines 101-105: **done path** - extract `<answer>...</answer>`, or treat any remaining text
49+
as the final answer; return `{"ok": True, "answer": ..., "steps": [...]}`.
50+
- Lines 107-111: **safety exit** - 10 steps without an answer -> `{"ok": False, "error": "max steps ..."}`.
51+
52+
### `orchestrator/llm.py` (was 30 lines, now 51)
53+
- Line 2: added `from typing import Any`.
54+
- Lines 34-51: NEW `chat(messages, tools=None, timeout_s=120)` - like `ask_llm` but takes the
55+
full message history and optional tools, and returns the whole assistant message as a dict
56+
(via `model_dump()`), including `tool_calls` when the model wants a tool.
57+
`ask_llm` (lines 21-31) is unchanged and still powers `/llm/ping`.
58+
- Effect: the phone now speaks the tool-calling dialect; the agent loop uses it.
59+
60+
### NEW `orchestrator/test_agent.py` (118 lines) - CI-safe fake-brain tests
61+
- Lines 6-18: `_tool_msg()` helper - builds a fake tool-calling message.
62+
- Lines 21-31: `test_answers_without_tools` - model answers immediately -> ok, no steps.
63+
- Lines 33-57: `test_uses_bash_tool` - model requests `ls -la /`, gets a fake result,
64+
then answers -> 1 tool step recorded.
65+
- Lines 59-87: `test_tag_fallback_when_no_tool_calling` - `<bash>echo hi</bash>` path.
66+
- Lines 89-101: `test_max_steps_guard` - model loops forever -> ok False, "max steps", 10 steps.
67+
- Lines 103-106: `run_agent_sync()` - small helper wrapping `asyncio.run`.
68+
- Key detail: `fake_chat` fakes are **sync** (llm.chat is sync) and `fake_exec` fakes are
69+
**async** (sandbox_exec is awaited) - mismatching these makes tests explode with
70+
coroutine errors (learned the hard way, twice).
71+
72+
## Why (the decisions)
73+
- **Why a trace (`steps` list)?** Every action is recorded end-to-end. That's Day 45's
74+
observability planted early - you can see exactly what the robot did.
75+
- **Why module imports?** Testability (see above).
76+
- **Why MAX_STEPS = 10?** Free-tier models are rate-limited (~50 calls/day); 10 steps is a
77+
generous ceiling that also prevents infinite loops and cost blowups.
78+
- **Why both tool calling AND tags?** Free OpenRouter models rotate; the loop must keep
79+
working no matter which model the auto-router picks.
80+
- **Why catch-all with noqa?** Verified: litellm exception hierarchy has no common base.
81+
A blind catch with an explanation is the correct engineering choice here.
82+
83+
## What changed / what it brings
84+
- The orchestrator can now run an autonomous multi-step agent task with one function call:
85+
`run_agent("list files and tell me the python version")`.
86+
- No new dependencies (uses `llm.py` + `sandbox.py` + stdlib `json`/`re`).
87+
- Test count: orchestrator now 7 tests (3 API + 4 agent), all CI-safe (no network, no Docker).
88+
89+
## Proof (test run, 9 Aug 2026)
90+
```
91+
$ uv run pytest -q
92+
7 passed, 1 warning in 2.46s
93+
94+
$ uv run ruff check . -> All checks passed!
95+
$ uv run ruff format --check . -> 7 files already formatted
96+
$ uv run mypy agent.py ... -> Success: no issues found
97+
```
98+
99+
## What's next (Day 3, Steps 2-5)
100+
- Step 2: expose the loop as `POST /agent/run` in `orchestrator/main.py`
101+
- Step 5: LIVE demo - the robot solves "list files and tell me the Python version" for real,
102+
using the OpenRouter key and the Docker sandbox.

‎docs/day3-step2-llm-upgrade.md‎

Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,71 @@
1+
# Day 3 — Step 2: Upgrade the phone (`orchestrator/llm.py`)
2+
3+
## Topic
4+
5+
`orchestrator/llm.py` is the "phone" of the agent: the only module allowed to
6+
talk to an LLM provider (OpenRouter today, via LiteLLM). This step upgraded it
7+
from a single sync call into a resilient, async-capable layer.
8+
9+
Half of this step already landed during Day 3 Step 1 (the `chat(messages,
10+
tools=None)` function that returns the full message with `tool_calls`, used by
11+
the agent loop). This step finished the job.
12+
13+
## What I did
14+
15+
1. **Added fallback provider settings** to `LLMSettings`
16+
(`llm.py:21-23`): `LLM_FALLBACK_PROVIDER`, `LLM_FALLBACK_API_KEY`,
17+
`LLM_FALLBACK_MODEL`. These were already documented in
18+
`infra/.env.example` but were never implemented.
19+
2. **Extracted `_completion()`** (`llm.py:26-43`): the actual LiteLLM call,
20+
parameterized by provider/key/model, so both `chat` and the fallback path
21+
share one code path.
22+
3. **Automatic fallback in `chat()`** (`llm.py:59-88`): if the primary call
23+
throws (provider down, rate limit, bad key), it retries once with the
24+
fallback provider — but only if `LLM_FALLBACK_API_KEY` is configured.
25+
Otherwise it re-raises the original error.
26+
4. **Added `achat()`** (`llm.py:91-96`): async twin of `chat` that runs the
27+
blocking LiteLLM call in a worker thread via `asyncio.to_thread`, so the
28+
server event loop is never blocked.
29+
5. **Agent loop now uses the async phone** (`agent.py:57`):
30+
`await llm.achat(messages, tools=[BASH_TOOL])`.
31+
6. **New tests** (`test_llm.py`, 4 offline tests): primary path passes tools +
32+
key through; fallback fires when primary fails and dials the backup; no
33+
fallback configured means the error propagates; `achat` returns the same
34+
result off the event loop.
35+
36+
## Why
37+
38+
- **Resilience:** `run_agent`/`/agent/run` depend on one provider. If OpenRouter
39+
is down or rate-limited, the whole agent dies. The fallback (e.g. Gemini via
40+
a second key) keeps the loop alive for a fraction of the code cost — it was
41+
already promised by `.env.example`, now it actually works.
42+
- **Event-loop hygiene:** FastAPI's endpoints are async; `run_agent` awaited
43+
a sync LiteLLM call, which would freeze the entire server for up to 120s on
44+
one slow request. `achat` moves that off-thread.
45+
- **Tests stay offline:** the fallback logic is pure routing around a
46+
monkeypatched `litellm.completion` — no network, no keys needed.
47+
48+
## What change it brings
49+
50+
- If you put `LLM_FALLBACK_PROVIDER=gemini`, `LLM_FALLBACK_API_KEY=...`,
51+
`LLM_FALLBACK_MODEL=gemini-2.5-flash` in `infra/.env`, agent requests
52+
automatically use Gemini when OpenRouter fails.
53+
- The agent loop no longer blocks the FastAPI event loop.
54+
- Behavior with a single key is unchanged (verified with a real call: reply
55+
`ok`, no tool_calls).
56+
57+
## Where (file → lines)
58+
59+
| File | Lines | Change |
60+
|---|---|---|
61+
| `orchestrator/llm.py` | 1-9 | imports + env path (asyncio added) |
62+
| `orchestrator/llm.py` | 12-23 | `LLMSettings` fallback fields |
63+
| `orchestrator/llm.py` | 26-43 | new `_completion()` (extracted) |
64+
| `orchestrator/llm.py` | 46-56 | `ask_llm` unchanged |
65+
| `orchestrator/llm.py` | 59-88 | `chat` with fallback logic |
66+
| `orchestrator/llm.py` | 91-96 | new `achat()` |
67+
| `orchestrator/agent.py` | 57 | loop uses `await llm.achat(...)` |
68+
| `orchestrator/test_llm.py` | 1-81 | new file, 4 offline tests |
69+
70+
Verified: `ruff check` clean, `mypy` clean, `pytest` 11 passed
71+
(3 sandbox + 4 agent + 4 llm), plus 1 live call to the real API.

‎docs/day3-step3-agent-run.md‎

Lines changed: 69 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,69 @@
1+
# Day 3 — Step 3: Expose the loop to the world (`orchestrator/main.py`)
2+
3+
## Topic
4+
5+
The agent loop (`agent.py`) was a library function. This step gives it a door:
6+
`POST /agent/run` on the orchestrator API. Anyone can now send a task and get
7+
back the full trace of the agent's thinking + the final answer.
8+
9+
## What I did
10+
11+
1. **Imported the loop** (`main.py:5`): `from agent import run_agent`.
12+
2. **Request model** (`main.py:18-19`): `AgentRunRequest` with a single `task`
13+
field, validated by Pydantic.
14+
3. **The endpoint** (`main.py:47-51`):
15+
- Empty/whitespace task → HTTP `422` (fail fast, no LLM call wasted).
16+
- Otherwise `await run_agent(req.task)` → returns the loop's dict directly:
17+
`{ok, answer, steps}` (or `{ok: false, error, steps}` on failure).
18+
4. **API tests** (`test_agent_api.py`, new file):
19+
- `test_agent_run_wiring` — monkeypatches `main.run_agent` (the name bound
20+
in main's namespace, not the source module) so no network is touched;
21+
asserts 200 + `ok`/`answer`/`steps` shape.
22+
- `test_agent_run_rejects_empty_task` — asserts 422 + message.
23+
5. **Fixed the environment** for the live demo: the Rust sandbox-gateway
24+
(port 8080) was built but not running — the first demo attempt failed with
25+
`httpx.ConnectError` at the gateway. Restarted the prebuilt binary
26+
(`sandbox-gateway/target/debug/sandbox-gateway`).
27+
28+
## Why
29+
30+
- **Traceability by design:** the response keeps every `steps[]` entry —
31+
each tool call, its command, and the raw sandbox output. That's the seed of
32+
Day 45's observability (replay "what did the agent do?"): today it's just
33+
JSON in an HTTP response, later it becomes stored sessions + live streams.
34+
- **The demo proved the whole stack:** HTTP → FastAPI → agent loop → LiteLLM
35+
→ real model → real tool call → Docker sandbox → extracted `<answer>`.
36+
37+
## What change it brings
38+
39+
- `curl -X POST localhost:8000/agent/run -H 'Content-Type: application/json' -d '{"task": "..."}'`
40+
now drives the full agent. Response example (real run):
41+
42+
```json
43+
{
44+
"ok": true,
45+
"answer": "The directories in the root directory (`/`) are: bin, boot, dev, ...",
46+
"steps": [
47+
{
48+
"type": "tool",
49+
"cmd": "ls -la /",
50+
"result": { "exit_code": 0, "stdout": "total 56\ndrwxr-xr-x ...", "duration_ms": 3746 }
51+
}
52+
]
53+
}
54+
```
55+
56+
- Bad input is rejected with 422 before any API calls are spent.
57+
58+
## Where (file → lines)
59+
60+
| File | Lines | Change |
61+
|---|---|---|
62+
| `orchestrator/main.py` | 5 | import `run_agent` |
63+
| `orchestrator/main.py` | 18-19 | `AgentRunRequest` model |
64+
| `orchestrator/main.py` | 47-51 | `POST /agent/run` endpoint (422 guard + call) |
65+
| `orchestrator/test_agent_api.py` | 1-29 | new file, 2 offline API tests |
66+
67+
Verified: `ruff` clean, `mypy` clean, `pytest` **13 passed**, plus a live
68+
end-to-end run over real HTTP (server on :8010, real LLM, real Docker sandbox —
69+
1 tool step, answer with directory list).

‎docs/day3-step4-agent-tests.md‎

Lines changed: 65 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,65 @@
1+
# Day 3 — Step 4: Tests that need no AI (`orchestrator/test_agent.py`)
2+
3+
## Topic
4+
5+
The agent loop must be testable **without any AI** — no API key, no network,
6+
no Docker. Step 4 is the test suite that proves the loop's behavior by
7+
swapping the real brain (LLM) and hands (sandbox) for scripted fakes.
8+
9+
## What I did
10+
11+
A monkeypatched **fake brain** (`llm.chat`) plus a **fake hand**
12+
(`sandbox.sandbox_exec`) drive `run_agent` in 4 offline tests:
13+
14+
1. **Answers immediately** (`test_agent.py:22-34`) — fake brain replies
15+
`<answer>42</answer>` on the first turn. Asserts `ok`, `answer == "42"`,
16+
`steps == []` (no tool was ever needed).
17+
2. **Uses the tool** (`test_agent.py:37-66`) — fake brain emits a real
18+
`tool_calls` entry on turn 1 (via the `_tool_msg` helper, `:8-19`), then
19+
`<answer>files listed</answer>` on turn 2. Fake hand executes `ls -la /`.
20+
Asserts the tool result landed in `steps[0]` with the right `cmd`.
21+
3. **`<bash>` tag fallback** (`test_agent.py:69-101`) — fake brain ignores tool
22+
calling and answers in plain text with `<bash>echo hi</bash>`; the loop must
23+
parse the tag and still execute the command. Asserts `steps[0]["cmd"] == "echo hi"`.
24+
4. **Never finishes → max-steps guard** (`test_agent.py:103-121`) — fake brain
25+
loops forever issuing tool calls. Asserts `ok is False`, error mentions
26+
"max steps", and exactly `MAX_STEPS` steps were recorded (no infinite loop).
27+
28+
`run_agent_sync` (`:124-125`) wraps the async loop in `asyncio.run` so tests
29+
stay plain sync functions.
30+
31+
## Why
32+
33+
- **CI-safe by construction:** the loop's only two I/O boundaries are
34+
monkeypatched in every test — the imports are just `asyncio`, `json`,
35+
`typing` and `agent`. No `httpx` socket ever opens, no Docker container ever
36+
starts, no key needed. Tests run anywhere, in seconds, for free.
37+
- **Deterministic:** the fake brain always behaves the same, so the loop's
38+
logic (tool-call routing, tag fallback, guard rails) is tested in
39+
isolation — exactly what you can't get from a real, flaky model.
40+
- **Guard rail proven:** the max-steps test pins the loop's one real safety
41+
property — a runaway model cannot run forever.
42+
43+
## What change it brings
44+
45+
- `pytest` now validates the core loop offline: 4 tests, ~0 API cost, 0
46+
network calls, 0 Docker calls.
47+
- Any future edit to the loop (new tool, different prompt, guard changes) is
48+
caught in seconds without touching the wallet.
49+
- This is the seed of the project's test discipline: fakes at the boundaries,
50+
real integrations proven separately with 1-2 live calls.
51+
52+
## Where (file → lines)
53+
54+
| File | Lines | Change |
55+
|---|---|---|
56+
| `orchestrator/test_agent.py` | 1-5 | imports (loop constants + stdlib only) |
57+
| `orchestrator/test_agent.py` | 8-19 | `_tool_msg` fake assistant tool-call message |
58+
| `orchestrator/test_agent.py` | 22-34 | test: immediate answer, no tools |
59+
| `orchestrator/test_agent.py` | 37-66 | test: real tool call → result in steps |
60+
| `orchestrator/test_agent.py` | 69-101 | test: `<bash>` tag fallback path |
61+
| `orchestrator/test_agent.py` | 103-121 | test: max-steps guard fires |
62+
| `orchestrator/test_agent.py` | 124-125 | `run_agent_sync` async wrapper |
63+
64+
Verified: `pytest` — 13 passed total, including these 4, with no network or
65+
Docker involved (all boundaries monkeypatched).

0 commit comments

Comments
 (0)