A minimal autonomous coding agent that actually works. No frameworks, no magic — just a while loop, tool use, and a rich TUI.
The core agent loop is small enough to read in one sitting (roughly ~70 lines of actual logic — the exact count depends on how you slice it). Most of the codebase is the terminal UI: of ~1,000 lines of Python, more than half live in the interface.
Status: agent-42 is an exploratory spike — a deliberately small, single-author project built to expose the bare mechanism of a coding agent, not to be a production tool. It has no test suite and pins no dependency versions (see Caveats). Read it, learn from it, fork it. Treat it as a teaching artifact, not a supported product.
Most agent frameworks bury the actual mechanism under layers of abstraction — chains, planners, memory modules, orchestrators. agent-42 does the opposite: it exposes the loop.
while True:
response = llm.stream(messages)
if response has tool_calls:
result = execute_tool(tool_call)
messages.append(tool_result(result))
continue
display(response.text)
messages.append(user_input())The model calls a tool, gets the result, reasons again. That's interleaved thinking — and it's all you need for an agent that writes code, runs it, reads files, fixes bugs, and iterates autonomously.
The intelligence is in the model, not in the code.
- One loop, two front-ends — the same
run_turndrives both the Textual TUI (app.py) and a plain CLI fallback (python agent.py). The loop is decoupled from the UI through a dict of callbacks (on_chunk,on_tool_call,on_tool_result, …), so the interface is fully pluggable. - Native tool-use loop, no orchestrator —
run_turn(inagent.py) is the rawreason → act → observepattern with no framework on top. LangChain appears only as a provider adapter inllm.py. - Three tools, plain dicts —
bash,read_file,write_file, defined as OpenAI function-calling dicts. Dispatch is a singlematch/case. No decorators, no registry. - Sandboxed shell — the
bashtool runs viadocker execin a container started withnetwork_mode: none. See What is and isn't sandboxed for the important caveat about the file tools. - Two-level context management —
pruneclears old tool outputs while protecting recent turns;compactsummarizes the whole conversation through the LLM when you approach the context limit. - Streaming with debounced rendering — chunks accumulate incrementally, and the TUI re-renders markdown on a debounce (every ~50 ms / ~200 chars) instead of repainting on every token.
- Multi-provider — any OpenAI-compatible endpoint. Three providers ship pre-wired (Anthropic, OpenAI, Z.AI); add more in
config.py.
- Python 3.10+ — the code uses
match/caseandX | Nonetype hints. - Docker — required only for the
bashtool. Without Docker,read_fileandwrite_filestill work; only shell execution breaks.
Dependencies are not pinned — there is no requirements.txt, pyproject.toml, or lockfile. Install the latest of each package (see below) and be aware that a future breaking release of langchain or textual could require adjustments.
git clone https://github.com/JoaoHenriqueBarbosa/agent-42.git
cd agent-42
python -m venv venv
source venv/bin/activate
pip install langchain-openai python-dotenv textual
docker compose up -d # builds the ubuntu:24.04 sandbox, network disabled
cp .env.example .env
# Fill in your provider keys (see the note below)Heads up on
.env: despite the "at least one provider is required" comment in.env.example,config.pycurrently reads every provider key withos.environ["..."](direct access). If any of the three keys —ANTHROPIC_API_KEY,OPENAI_API_KEY,ZAI_API_KEY— is missing, importingconfig.pyraisesKeyErrorand the app won't start. Until this is fixed, set all three (a placeholder value is enough for providers you don't intend to select).
Run the TUI:
python app.pySelect a provider with the arrow keys, hit Enter, and start coding.
Or the CLI fallback (no Textual UI, just a > prompt):
python agent.pyAny OpenAI-compatible API works out of the box:
| Provider | Default model | Endpoint | Context |
|---|---|---|---|
| Anthropic | claude-sonnet-4-20250514 | api.anthropic.com/v1/ |
200k |
| OpenAI | gpt-4o-mini | api.openai.com/v1 |
128k |
| Z.AI | GLM-4.5-air | api.z.ai/api/coding/paas/v4 |
128k |
Each provider is configured with <PROVIDER>_API_KEY, <PROVIDER>_BASE_URL, and <PROVIDER>_MODEL environment variables. To use a local model (Ollama, LM Studio, …), point one of these providers' BASE_URL at your local endpoint, or add a new entry in config.py — these aren't pre-configured out of the box.
| Tool | What it does | Runs in |
|---|---|---|
bash |
Runs a shell command (working dir /workspace, 30s timeout) |
Docker sandbox (no network) |
read_file |
Reads a file with line numbers, optional start_line/end_line |
Host |
write_file |
Creates or overwrites a file, auto-creating parent directories | Host |
Tools are plain Python dicts in OpenAI function-calling format. No decorators, no abstractions. Dispatch is a match/case in execute_tool.
This distinction matters, so it's stated plainly:
bashis sandboxed. It executes throughdocker execinside a container started withnetwork_mode: none. Shell commands cannot reach the network and run isolated from the host.read_fileandwrite_filerun on the host. They operate directly on the host filesystem, rooted at theworkspace/directory and guarded only by a_safe_pathprefix check — they do not go through the container. The isolation guarantee applies to shell execution, not to file reads and writes.
If you point the agent at a directory you care about, remember it can write to files under the configured workspace on your real machine.
graph TD
TUI["Textual TUI<br/>(app.py)"] <--> Agent["agent.py<br/>async run_turn<br/>tool dispatch: match/case"]
CLI["CLI fallback<br/>(agent.py main)"] <--> Agent
Agent <--> LLM["LLM provider<br/>(any OpenAI-compatible)"]
Agent --> Bash["bash"]
Agent --> ReadFile["read_file<br/>(host)"]
Agent --> WriteFile["write_file<br/>(host)"]
Bash --> Docker["Docker sandbox<br/>(no network)"]
agent-42/
├── app.py # Textual entrypoint — composes TUI, runs agent as async worker
├── agent.py # run_turn() agent loop (callback-decoupled) + CLI main()
├── ui.py # Textual widgets: ChatView, ChatMessage, ToolWidget, ChatInput, StatusFooter
├── ui_cli.py # Plain print/input UI for the CLI mode
├── tools.py # 3 tool dicts + execute_tool dispatch + tool_bash/read/write
├── context.py # get_token_count, prune, compact — context management
├── llm.py # make_llm(provider) with ChatOpenAI + bind_tools; streaming helpers
├── config.py # Provider configuration from .env
├── prompts.py # Loads system_prompt.txt / compact_prompt.txt
├── system_prompt.txt # Agent persona and working conventions
├── compact_prompt.txt# Summarization template for auto-compaction
├── styles.tcss # Textual CSS
├── Dockerfile # ubuntu:24.04 sandbox image
└── compose.yml # Brings up the agent-42-sandbox container
The context layer (context.py) keeps a long conversation inside the model's window through two mechanisms:
prunereplaces old tool outputs with[output cleared], preserving roughly the last ~40k tokens and the last two user turns. It only acts when there are enough prunable tokens to be worth it (~20k+).compactkicks in when the conversation overflows (limit minus a ~20k buffer): it asks the LLM to summarize the entire history, then restarts the message list from that summary.
Token counts come from the provider's usage_metadata when available, with a heuristic fallback: estimate_tokens computes total_chars // CHARS_PER_TOKEN (with CHARS_PER_TOKEN = 4). That fallback is a plain linear map in one dimension, f(x) = x / 4 — scaling character count by a constant. Because it's linear, it's additive: f(a + b) = f(a) + f(b), which is exactly why the code can sum the character lengths of every message first and divide once at the end, rather than estimating each message separately and adding up. There's no clamping or saturation anywhere in that path, so the linearity holds end to end.
Why a while loop? Because that's what an agent is. The model decides, acts, observes, decides again. Any layer on top of this that doesn't add new capability is dead weight.
Why LangChain at all? One reason: ChatOpenAI normalizes message formats across providers. If LangChain dies tomorrow, the migration is swapping .stream() for HTTP calls. Nothing else changes.
Why Docker for bash? The model runs arbitrary shell commands. Docker with network_mode: none gives isolation for free. The model doesn't know it's in a container.
Why Textual? A coding agent deserves better than print(). Markdown rendering, streaming, collapsible tool output, input history — all without leaving the terminal.
agent-42 is a native tool use agent loop — the practical evolution of ReAct (Yao et al., 2022).
ReAct formalized reason → act → observe → reason again, but relied on prompting hacks (Thought:, Action:, Observation: tokens). With native tool use APIs, the pattern became infrastructure:
| ReAct concept | Native implementation |
|---|---|
| Thought | Extended thinking / internal reasoning |
| Action | tool_use block (typed, structured) |
| Observation | tool_result message back to the model |
agent-42 is this loop with zero translation layers. The model speaks directly to the system.
- ReAct: Synergizing Reasoning and Acting in Language Models — Yao et al., 2022. arXiv:2210.03629
- Toolformer: Language Models Can Teach Themselves to Use Tools — Schick et al., 2023. arXiv:2302.04761
- Tool Use — Anthropic Docs
- Function Calling — OpenAI Docs
There is no build step and no test suite — this is a small, flat, script-style project. To hack on it:
git clone https://github.com/JoaoHenriqueBarbosa/agent-42.git
cd agent-42
python -m venv venv && source venv/bin/activate
pip install langchain-openai python-dotenv textual
docker compose up -d
python app.py # or: python agent.pyAll Python modules live at the repository root; there is no package or subpackage to install. If you add tests, pytest is the natural fit — contributions that introduce a test suite are especially welcome (see CONTRIBUTING.md).
Called out explicitly, because honesty beats surprise:
- No tests, no CI. There is currently zero automated test coverage and no CI pipeline.
- Unpinned dependencies. No lockfile or version constraints; a future release of a dependency could break the install.
- Not a published package. There's no
pyproject.toml/setup.py; you run it from a clone, notpip install agent-42. - All three provider keys must be set today (see the
.envnote above), even though only one provider is used at a time. - File tools run on the host — only
bashis sandboxed (see What is and isn't sandboxed). - Roadmap items are not done. Tool cancellation and session persistence (below) are not implemented.
- Context compaction
- Textual TUI with markdown streaming and tool widgets
- Ctrl+C to cancel running tools
- Session persistence
- Allow a single provider to boot without all keys set
- System prompt refinement
Contributions are welcome. See CONTRIBUTING.md for setup and the PR flow, CODE_OF_CONDUCT.md for community expectations, and SECURITY.md for reporting vulnerabilities.
Released under the MIT License.


