A from-scratch OpenAI tool-calling agent built around a reducer-style
state machine. No langchain, no gradio, no streamlit — the kind
of thing you'd build in a live AI Engineer interview.
| Branch | Purpose |
|---|---|
master |
Starter project — work through the checkpoints below. |
solution |
Completed reference implementation (agent loop, server, streaming). |
git checkout solution # see the finished code
git checkout master # back to the exerciseBuild an agent that satisfies user prompts by chaining tool calls. The "agent" is just three pieces:
| Piece | Role |
|---|---|
| System prompt | The spec of the state machine. |
TOOL_REGISTRY dict |
The reducer table — a map of name → fn. |
| The agent loop | The dispatcher — runs one step at a time. |
Each tool is a small reducer function: given some arguments, it returns structured JSON that gets appended back to the conversation as new context. The LLM looks at the updated state and decides the next action (call another tool, or stop and answer).
"What is the word for today's day of the week, in Spanish?"
A single round trip can't answer this — the model has no clock and no dictionary. With our reducer table it becomes a 3-step state machine:
user prompt
│
▼
[ACTION] get_nanotime() ─► {nanotime: 1_748_887_xxx_xxxxxxxxx}
│
▼
[ACTION] day_of_week_from_nanotime(nanotime) ─► {day_of_week: "Tuesday"}
│
▼
[ACTION] translate(text="Tuesday",
target_language="spanish") ─► {translated: "martes"}
│
▼
"martes."
Other queries route through the same machine:
- "What time is it in nanoseconds?" → 1 call (
get_nanotime). - "How do you say goodbye in French?" → 1 call (
translate). - "What day is it today, in German?" → 3 calls (time → day → translate).
- "What is unobtainium in French?" → 1 call, tool returns an
errorpayload, the model gracefully says it doesn't know.
The "agentic" magic is just while tool_calls: dispatch; append; ask again. Everything else is glue.
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ main.py │────────▶│ agent.py │────────▶│ OpenAI API │
│ (CLI loop) │ │ (agent loop)│◀────────│ │
└─────────────┘ └──────┬──────┘ └─────────────┘
│ httpx (HTTP)
▼
┌─────────────┐
│ server.py │ ← FastAPI + uvicorn
│ /tools/* │
└──────┬──────┘
▼
┌─────────────┐
│ tools.py │ ← reducer table + JSON schemas
└─────────────┘
The LLM never runs code. It returns a tool_calls payload that
describes the call it wants. Our agent POSTs to our own FastAPI server,
gets the result, appends it to the conversation, and loops.
- macOS / Linux
uv(brew install uvon macOS)- Python 3.13 (uv will install it if missing)
- An OpenAI API key
cd ~/Documents/projects/agentic-from-scratch
# uv already created .venv and installed deps when the project was scaffolded.
# If you ever pull this fresh or wipe .venv, run:
uv sync
# Make your .env from the example and fill in your key:
cp .env.example .env
# then edit .env and put your real OPENAI_API_KEY inWhy
uv? It replacespython -m venv,pip,pip-tools, andpip install -r requirements.txtwith a single fast tool.uv syncreadspyproject.toml+uv.lockand gives you a reproducible.venv/.uv run <cmd>runs<cmd>inside that venv without you having to activate it.
If you'd rather see what's happening:
source .venv/bin/activate # classic venv activation
python -V # should print 3.13.x
which python # should point inside .venv
deactivate # exit the venvYou'll need two terminals:
Terminal 1 — the tool server (FastAPI + uvicorn):
uv run uvicorn server:app --reload --port 8000server:appmeans "importappfromserver.py".--reloadwatches files and hot-restarts uvicorn on changes.- Visit http://localhost:8000/docs for an auto-generated Swagger UI.
Terminal 2 — the chat REPL:
uv run python main.pyUntil you implement the loop you'll see NotImplementedError — that's
expected. Use it as your TODO list.
uv run pytest -qWork through these in order. Each step is small, runs in seconds, and leaves you with a thing you can actually exercise.
uv run uvicorn server:app --port 8000boots without error.curl localhost:8000/healthzreturns{"status":"ok"}.curl localhost:8000/toolsreturns the registry list with the one worked schema (get_nanotime).uv run pytest -qshows 2 passing tests.
- Implement
day_of_week_from_nanotime(nanotime)— convert ns → seconds, build a UTC datetime, return{"day_of_week": "<name>"}. - Implement
translate(text, target_language)— normalize both inputs, look up_TRANSLATIONS[lang][word], return errors as data (not exceptions) so the model can recover. - Add the two missing schemas to
TOOL_SCHEMASso the LLM can call them. - Uncomment the TODO tests in
tests/test_tools.pyand make them pass.
- Finish
POST /tools/{name}per the docstring (map exceptions to HTTP status codes). - Verify each transition with curl:
curl -s -XPOST localhost:8000/tools/get_nanotime -H 'content-type: application/json' -d '{}' curl -s -XPOST localhost:8000/tools/day_of_week_from_nanotime -H 'content-type: application/json' -d '{"nanotime": 1748887200000000000}' curl -s -XPOST localhost:8000/tools/translate -H 'content-type: application/json' -d '{"text":"Tuesday","target_language":"spanish"}'
The big one. Implement in this order:
_call_model()— one OpenAI call, normalize the assistant message into a re-sendable dict (includingtool_callsif present)._execute_tool_call()— JSON-decodearguments, POST to the server, return arole="tool"message withtool_call_idset.chat()— the loop itself. Append user message, then repeatedly: call the model, append its reply, execute every tool call, append each result, repeat — until the model returns a message with notool_calls. Cap withself.max_steps.
Then run uv run python main.py and ask:
- "What time is it in nanoseconds?" → 1 call
- "How do you say
thanksin German?" → 1 call - "What is today's day of the week, in Spanish?" → 3 calls
- "Same question but in French and German, please." → up to 5 calls
Watch the server logs in Terminal 1 — you should see one POST per tool call, in the order the agent decided to make them.
- More tools:
format_date(nanotime, fmt),add_phrase(language, english, translation)(mutates_TRANSLATIONSat runtime so the agent can teach itself a word and reuse it next turn). - Streaming: switch to
stream=Trueand print tokens as they arrive. - Parallel tools: execute multiple
tool_callsin one assistant turn concurrently withconcurrent.futures.ThreadPoolExecutor. - Trace log: print every model call + tool call as JSON so you can replay/debug a conversation.
- Async end-to-end:
AsyncOpenAI+httpx.AsyncClient, expose the agent itself as a FastAPI route atPOST /chat. - Real translation tool: replace the dictionary with a call to an
actual translation API (kept behind the same
translateschema — the rest of the system doesn't change).
- Tools = reducer functions, registry = reducer table, loop = dispatcher. You're really just running a tiny event-sourced state machine where the LLM is the policy.
- The model is a planner, not an executor. Tool calls are requests the model makes; the trust boundary is in our code. The HTTP hop makes that boundary explicit and lets us reuse the same tools from multiple agents or processes.
- Schemas live next to implementations.
tools.pyowns the Python callable and the JSON schema the model sees, so they can't drift apart. max_stepsguards against runaway loops. Cheap, essential, and the first thing a reviewer will look for.- Errors are data. When a tool returns
{"error": ...}instead of raising, the model can read it on its next turn and either retry or admit a limitation — the conversation never crashes.