An open-source AI assistant that thinks before it acts — and stops before it sends.
A harness is a governance and reliability control plane around an autonomous agent: the agent has intelligence, the harness has authority. The agent proposes; the harness decides what's allowed to happen next; evidence decides whether the result is accepted. Aielia — the Build A Harness personal assistant — routes every turn through an 11-layer implementation of that control plane, governing what the agent believes, what it's allowed to do, how it catches its own mistakes, and what it learns. A quick fact lookup stays light. Sending an email, paying an invoice, running a shell command, or deleting a file stops for your approval first — and a classifier that errors out requires approval rather than sailing through as safe.
The assistant is the front door. Underneath it is a full visual harness builder — draw the same 11 layers on a canvas, compile to LangGraph / CrewAI / Mastra / MS Agent Framework, trace every decision in Langfuse.
npx @buildaharness/personal-assistantFirst run walks you through picking a model — reuse an existing claude CLI
login (no API key), or paste an Anthropic / OpenAI / OpenRouter key. Then just
talk to it.
import { LLMClient } from '@buildaharness/runtime'
import { PersonalAssistant } from '@buildaharness/personal-assistant'
const aielia = new PersonalAssistant({ llmClient: new LLMClient({ proxyUrl, authToken }) })
await aielia.turn('What time zone is Tokyo in?')
// { status: 'ok', reply: '…', riskLevel: 'LOW', stepsUsed: 1 }
await aielia.turn('Send an email to my boss saying I quit.')
// { status: 'needs_approval', reason: '…', riskLevel: 'HIGH' } — no LLM call made
await aielia.turn('Send an email to my boss saying I quit.', { approved: true })
// approved — proceeds and runs the harness normallyTry it in your browser → buildaharness.com/try — bring your own key (stored only in your browser), or see the approval gate fire before you add one.
One core, three front ends: terminal CLI, browser (@buildaharness/chat-ui),
and native desktop (@buildaharness/desktop).
The /harness-comparison page
maps the three most-used open agents (Hermes Agent, Kilo Code, OpenClaw) against
this architecture. None of them ships both a tiered Control State resolver and
a reviewer/output gate. Aielia ships:
- Live per-tool-call ControlState gate — every read-only tool call is checked
against a per-turn
ControlStatebefore it runs (deterministic ALLOW / DENY / REQUIRE_APPROVAL, not advisory), so a developing failure pattern can trip a real deny mid-turn. - Fail-safe risk classification — a classifier error or unparseable response
returns
UNKNOWN → requires approval, never a silent default to low-risk. - Reviewer Pass — a 3-lens review (consistency, adversarial, abstraction fit) and output-contract validation run before a reply goes out.
- Typed fact provenance — only facts you actually state promote to durable memory by default; model-inferred facts stay session-scoped until confirmed.
- AnswerClaim — replies distinguish "verified against evidence" from "found this but couldn't independently confirm it," surfaced in the chat's "Why?" panel.
- Crash-safe mid-turn resume — a turn that dies mid-flight resumes from its last checkpoint instead of silently starting over; a checkpoint that keeps crashing on replay is discarded automatically after two attempts.
- Untrusted-content boundary — web results and shell output are wrapped as data the model is instructed never to follow as commands.
Full write-up: packages/personal-assistant/README.md.
A workflow routes prompts from node to node. A harness governs belief, permission, self-correction, and learning. Build A Harness delivers the complete 11-layer architecture as a visual builder.
Canvas → flow.json → LangGraph · CrewAI · Mastra · MS Agent Framework → Langfuse
The spec is the contract. The canvas is the editor. The adapters are the compilers.
| Simple Agent Loop | Full Harness — Implemented |
|---|---|
| Input / Caller | Caller State — constraints · clarification |
| ↓ | World Model — beliefs · contradictions · generation_id |
| LLM Call | Reasoning — evidence · hypotheses (4 sources) · VOI |
| ↓ | Control ← key — 5-tier resolver · ALLOW/DENY permission · NORMAL/CAUTIOUS/RECOVERY mode |
| Tool Call ↺ loop | Planning — task graph (6-state) · parallel concurrency |
| ↓ | Execution + Verification — VOI gate · 9 layers |
| Output | Recovery + Memory — 6 strategies · compression |
| Learning — experience store · warm start (optional) | |
| Output & Reviewer Pass — contract · 3-lens review | |
| prompt in → answer out | 27 nodes · 11 layers · 759 harness-layer tests |
|
Canvas & execution layer
|
Reasoning & control layer
|
The full node palette and schema-sync mechanics live in docs/nodes.md; the field-level FlowSpec reference is docs/flowspec.md.
All four runtimes compile from the same flow.json — no rewriting. /compile
checks the target runtime's actual capabilities first: a FlowSpec requiring
something the runtime doesn't support (durable checkpointing, token streaming)
fails fast with a clear error instead of silently degrading.
| Runtime | Language | HITL | Key integration |
|---|---|---|---|
| LangGraph | Python | interrupt() |
@observe · harness child spans |
| CrewAI | Python | — | context_from → Task.context · tier-aware memory |
| Mastra | TypeScript | suspend()/resume() |
Node.js sidecar |
| MS Agent Framework | Python | _HitlPause |
AgentGroupChat native · OTel → Langfuse |
Compile: POST /compile?runtime=langgraph. Deploy as a REST endpoint,
MCP tool, or A2A agent in one step.
Self-hosted Langfuse starts with docker compose up — no extra config.
Per-node child spans across all four runtimes, token/latency/cost per node via
LiteLLM, a live View trace → link in the canvas after each run, and managed
prompts via the Langfuse prompt API (prompt_ref on any llm_call node).
Just want the assistant? Nothing to clone:
npx @buildaharness/personal-assistant # terminal
# or open https://buildaharness.com/try # browser, bring your own keyBuilding and compiling harnesses needs the full stack (canvas + adapter API
- Langfuse):
./scripts/setup-env.sh # generate secrets, write .env
docker compose up # start all 12 services| Service | URL |
|---|---|
| Canvas | http://localhost:3000 |
| Adapter API | http://localhost:8000/health |
| Langfuse | http://localhost:3001 |
Without Docker
./scripts/setup-env.sh && source adapter/.venv/bin/activate
npm install && npm run dev # canvas → localhost:3000
cd adapter && python main.py # adapter → localhost:8000Running tests
npm test # Vitest — validates 5 reference flows
pytest adapter/tests/ -v # adapter unit + integration
pytest adapter/tests/test_maf_adapter.py -v # MAF suite (42 tests)New here? Start with docs/getting-started.md · Startup errors? docs/troubleshooting.md · Real-time collaboration: docs/collab.md · On-prem / Kubernetes: docs/deployment.md
The assistant reaches a model directly (Anthropic, OpenAI, OpenRouter, or a
claude CLI login). The full stack routes every call through LiteLLM — add
the key to .env:
| Provider | Env var | Example models |
|---|---|---|
| OpenAI | OPENAI_API_KEY |
gpt-4o, gpt-4o-mini |
| Anthropic | ANTHROPIC_API_KEY |
claude-sonnet, claude-opus |
| Ollama (local) | — | mistral, qwen3, qwen2.5-coder |
No API key? Install Ollama, run
ollama pull mistral, then./scripts/setup-ollama.sh— tests all four frameworks with no paid account.
Full setup: docs/llm-setup.md
npm install @buildaharness/canvasimport { BuildAHarnessCanvas } from '@buildaharness/canvas'
import '@buildaharness/canvas/styles.css'
<BuildAHarnessCanvas
initialSpec={mySpec}
onSpecChange={(updated) => save(updated)}
execStats={runState.nodeStats}
theme="dark"
/>Full props reference: packages/canvas/README.md
| docs/getting-started.md | Fresh clone → secrets → LLM → first run |
| docs/nodes.md | The 27-node palette + schema-sync mechanics |
| docs/flowspec.md | FlowSpec v1.0.0 — all 27 node types, edges, fields |
| docs/architecture.md | System design, service interactions, data flows — the 11 layers as implementations of 5 primitives (State, Evidence, Policy, Effect, Recovery) |
| docs/api.md | REST API reference — compile, execute, deploy, HITL resume |
| docs/llm-setup.md | LLM provider setup — OpenAI, Anthropic, Ollama, custom |
| docs/qdrant.md | Qdrant vector store — seeding, collections, production |
| docs/env-vars.md | All environment variables across all services |
| docs/collab.md | Real-time collaboration — Yjs setup and internals |
| docs/deployment.md | Docker, Helm, SSO/OIDC |
| docs/troubleshooting.md | Common startup errors |
| docs/threat-model.md | The assistant's trust boundaries + accepted non-goals |
| SECURITY.md | Supported versions + how to report a vulnerability |
| CONTRIBUTING.md | How to contribute |
Apache 2.0 — see LICENSE.