A local workbench for recording, replaying, comparing, and inspecting tool-using agent runs without sending trace data to a hosted service.
-
Updated
Aug 4, 2026 - Python
Agent harnesses are the runtime scaffolding around AI agents. They usually combine context delivery, tool interfaces, planning state, memory, sandboxes, permissions, evaluation, and observability so agents can complete longer tasks reliably. Agent harnesses are especially common in coding agents, research agents, and multi-agent workflows where repeatability, safety, and traceability matter.
A local workbench for recording, replaying, comparing, and inspecting tool-using agent runs without sending trace data to a hosted service.
Open-source framework for reproducible A/B evaluation of agent skills and instructions
Codex skill for designing and scaffolding durable agent harness projects, with a Ralph Loop preset and upgrade-friendly doctrine.
Interlinked markdown wiki on agent harnesses, orchestration, formal methods, and adjacent research.
Self-updating intelligence dashboard tracking the AI model and tooling landscape — frontier models, agent harnesses, self-hosting, strategy.