|
| 1 | +# Day 3, Step 1 - The agent loop engine (`orchestrator/agent.py`) |
| 2 | + |
| 3 | +The brain learns to press its own button. This file turns the orchestrator from a |
| 4 | +remote-controlled robot into one that thinks: decide a command -> run it in the sandbox -> |
| 5 | +read the output -> repeat -> answer. |
| 6 | + |
| 7 | +## The topic (ELI5) |
| 8 | + |
| 9 | +An **agent loop** is: THINK -> ACT -> LOOK -> THINK -> ... -> DONE. |
| 10 | + |
| 11 | +- THINK: the LLM decides the next shell command (or says "I'm done") |
| 12 | +- ACT: we run that command in the sandbox (the `/exec` button from Step 2) |
| 13 | +- LOOK: we show the output back to the LLM |
| 14 | +- Repeat until the LLM answers, or we hit the safety limit (10 steps) |
| 15 | + |
| 16 | +The LLM can press the button two ways: |
| 17 | +1. **Tool calling (primary):** the LLM returns a structured |
| 18 | + `{"name": "run_bash", "arguments": {"cmd": "ls -la /"}}` message (the industry-standard |
| 19 | + way - what Devin / Claude Code do). |
| 20 | +2. **Tag language (fallback):** some free models can't do tool calling, so we also teach: |
| 21 | + `<bash>command</bash>` = press the button, `<answer>text</answer>` = I'm done. |
| 22 | + Either way, the loop works. This resilience is the interview-gold part. |
| 23 | + |
| 24 | +## What I did, file by file (with line ranges) |
| 25 | + |
| 26 | +### NEW `orchestrator/agent.py` (111 lines) - the loop engine |
| 27 | +- Lines 1-6: imports. Note lines 5-6: `import llm` and `import sandbox` (module imports, |
| 28 | + NOT `from x import y`). Why: tests monkeypatch `llm.chat` and `sandbox.sandbox_exec`; |
| 29 | + module-style imports let the patch take effect at call time. (Learned the hard way - |
| 30 | + function imports silently ignored the patch and hit the real network.) |
| 31 | +- Line 8: `MAX_STEPS = 10` - the "unplug the robot" limit. |
| 32 | +- Lines 10-18: `SYSTEM_PROMPT` - the rules of the game: one tool, network limited to |
| 33 | + GitHub/PyPI/npm, prefer short commands, `<answer>` protocol, `<bash>` fallback. |
| 34 | +- Lines 20-33: `BASH_TOOL` - the JSON "button spec" in OpenAI tool-calling format |
| 35 | + (`type: function`, name `run_bash`, one required arg `cmd`). |
| 36 | +- Lines 36-45: `_summarize()` - turns an ExecResult into a short readable string for the LLM |
| 37 | + (handles timeouts, stderr-only, normal stdout, no output). |
| 38 | +- Lines 48-53: `run_agent(task, max_steps=10)` - builds the message history |
| 39 | + `[system, user task]` and an empty `steps` trace log. |
| 40 | +- Lines 55-59: the loop head - one LLM call per iteration via `llm.chat(messages, tools=...)`. |
| 41 | + `# noqa: BLE001` is intentional: LLM providers throw a zoo of exception types (verified |
| 42 | + live - no common base class exists), and the loop must survive any of them. |
| 43 | +- Lines 62-89: **tool-calling path** - if the model returned `tool_calls`, parse the JSON |
| 44 | + args (bad JSON -> empty cmd -> a helpful error result), run it with `sandbox.sandbox_exec`, |
| 45 | + record the step, append the result as a `role: "tool"` message, continue the loop. |
| 46 | +- Lines 91-99: **tag fallback path** - no tool_calls but `<bash>cmd</bash>` present: run the |
| 47 | + last one, append the output as a plain `user` message, continue. |
| 48 | +- Lines 101-105: **done path** - extract `<answer>...</answer>`, or treat any remaining text |
| 49 | + as the final answer; return `{"ok": True, "answer": ..., "steps": [...]}`. |
| 50 | +- Lines 107-111: **safety exit** - 10 steps without an answer -> `{"ok": False, "error": "max steps ..."}`. |
| 51 | + |
| 52 | +### `orchestrator/llm.py` (was 30 lines, now 51) |
| 53 | +- Line 2: added `from typing import Any`. |
| 54 | +- Lines 34-51: NEW `chat(messages, tools=None, timeout_s=120)` - like `ask_llm` but takes the |
| 55 | + full message history and optional tools, and returns the whole assistant message as a dict |
| 56 | + (via `model_dump()`), including `tool_calls` when the model wants a tool. |
| 57 | + `ask_llm` (lines 21-31) is unchanged and still powers `/llm/ping`. |
| 58 | +- Effect: the phone now speaks the tool-calling dialect; the agent loop uses it. |
| 59 | + |
| 60 | +### NEW `orchestrator/test_agent.py` (118 lines) - CI-safe fake-brain tests |
| 61 | +- Lines 6-18: `_tool_msg()` helper - builds a fake tool-calling message. |
| 62 | +- Lines 21-31: `test_answers_without_tools` - model answers immediately -> ok, no steps. |
| 63 | +- Lines 33-57: `test_uses_bash_tool` - model requests `ls -la /`, gets a fake result, |
| 64 | + then answers -> 1 tool step recorded. |
| 65 | +- Lines 59-87: `test_tag_fallback_when_no_tool_calling` - `<bash>echo hi</bash>` path. |
| 66 | +- Lines 89-101: `test_max_steps_guard` - model loops forever -> ok False, "max steps", 10 steps. |
| 67 | +- Lines 103-106: `run_agent_sync()` - small helper wrapping `asyncio.run`. |
| 68 | +- Key detail: `fake_chat` fakes are **sync** (llm.chat is sync) and `fake_exec` fakes are |
| 69 | + **async** (sandbox_exec is awaited) - mismatching these makes tests explode with |
| 70 | + coroutine errors (learned the hard way, twice). |
| 71 | + |
| 72 | +## Why (the decisions) |
| 73 | +- **Why a trace (`steps` list)?** Every action is recorded end-to-end. That's Day 45's |
| 74 | + observability planted early - you can see exactly what the robot did. |
| 75 | +- **Why module imports?** Testability (see above). |
| 76 | +- **Why MAX_STEPS = 10?** Free-tier models are rate-limited (~50 calls/day); 10 steps is a |
| 77 | + generous ceiling that also prevents infinite loops and cost blowups. |
| 78 | +- **Why both tool calling AND tags?** Free OpenRouter models rotate; the loop must keep |
| 79 | + working no matter which model the auto-router picks. |
| 80 | +- **Why catch-all with noqa?** Verified: litellm exception hierarchy has no common base. |
| 81 | + A blind catch with an explanation is the correct engineering choice here. |
| 82 | + |
| 83 | +## What changed / what it brings |
| 84 | +- The orchestrator can now run an autonomous multi-step agent task with one function call: |
| 85 | + `run_agent("list files and tell me the python version")`. |
| 86 | +- No new dependencies (uses `llm.py` + `sandbox.py` + stdlib `json`/`re`). |
| 87 | +- Test count: orchestrator now 7 tests (3 API + 4 agent), all CI-safe (no network, no Docker). |
| 88 | + |
| 89 | +## Proof (test run, 9 Aug 2026) |
| 90 | +``` |
| 91 | +$ uv run pytest -q |
| 92 | +7 passed, 1 warning in 2.46s |
| 93 | +
|
| 94 | +$ uv run ruff check . -> All checks passed! |
| 95 | +$ uv run ruff format --check . -> 7 files already formatted |
| 96 | +$ uv run mypy agent.py ... -> Success: no issues found |
| 97 | +``` |
| 98 | + |
| 99 | +## What's next (Day 3, Steps 2-5) |
| 100 | +- Step 2: expose the loop as `POST /agent/run` in `orchestrator/main.py` |
| 101 | +- Step 5: LIVE demo - the robot solves "list files and tell me the Python version" for real, |
| 102 | + using the OpenRouter key and the Docker sandbox. |
0 commit comments