Skip to content

About

Local fault-injection lab for tool-using AI agents. Test timeouts, duplicate delivery, stale reads and recovery with deterministic assertions, replayable evidence, MCP and a local inspector.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

AgentFaultLab

The tool timed out. Did the agent create the ticket twice?

AgentFaultLab is a local test lab for agents that change business state through tools. It runs a synthetic helpdesk, injects failures at explicit transaction boundaries, and evaluates the records and notifications that actually exist. Requests, deliveries, committed effects, and agent-visible observations stay separate in the evidence.

The first domain covers creating a ticket, assigning it with entity-version checks, and notifying a customer. Every effect is simulated. No real support system or email account is contacted.

Start with the keyless demo

Install uv and Git. The tested Python baseline is 3.12; uv can install it. Dependency installation requires internet access. The installed demo and core tests require no provider key or network service.

git clone https://github.com/vahapogut/agent-fault-lab.git
cd agent-fault-lab
uv sync --locked
uv run afl demo --output .afl/demo

Open .afl/demo/naive/report.html and .afl/demo/resilient/report.html in a browser. The reports work offline. The same executed evidence is stored as run.json, JUnit XML, and local SQLite history.

Scripted controller Behavior under a lost creation response Expected gate
builtin:naive Retries with a new operation key and creates a duplicate fail
builtin:resilient Reuses stable keys under the declared idempotency contract pass
builtin:noop Does nothing and fails the completion rules fail

These are acceptance fixtures, not measured LLM performance. The demo exits successfully only when the real engine produces the declared outcomes. Ordinary failed runs still exit with failure.

uv run afl run scenarios/helpdesk/lost-create-response.yaml --agent builtin:naive --output .afl/naive
uv run afl run scenarios/helpdesk/lost-create-response.yaml --agent builtin:resilient --output .afl/resilient
uv run afl replay .afl/naive/run.json
uv run afl compare .afl/naive/run.json .afl/resilient/run.json

Replaying the failed run should report reproduction: verified and original_gate: fail. It re-executes recorded actions without calling a model. It does not turn the original agent failure into a passing test.

What the lab measures

  • Completion: exact request ticket/notification counts and requested assignment.
  • Safety: at-most rules, correct recipient, approval, assignment before notification, unrelated ticket preservation, and authorized effects.
  • Execution: completion, agent/harness errors, cancellation, deadline and budget exhaustion.
  • Fault coverage: planned, matched and activated failures, including explicit suppression reasons.
  • Measurement: routed calls, deliveries, virtual time and real duration. Unobserved provider usage and cost remain unknown.

A run passes only when its assertions pass, execution completes, and every required fault activates. completed means the agent loop ended; the business task can still fail. A fail gate records an assertion violation, while error identifies a harness failure. Missing required fault coverage or unfinished execution makes the result inconclusive unless an observed violation takes precedence. Historical safety violations remain visible after later corrective actions. A final-state check does not erase an early notification.

Failure contracts

Fault Simulated boundary
timeout_before_commit No write or successful operation record; generic timeout
timeout_after_commit Write and operation record survive; the same generic timeout
duplicate_delivery Two deliveries for one agent call; ordinary server idempotency applies
rate_limit Typed pre-write denial with virtual retry delay, optionally repeated
stale_read A real earlier committed snapshot; authoritative state is unchanged

Idempotency is an explicit supported or unsupported tool contract. Business request keys are not unique database constraints. The simulator deliberately allows duplicate requests so incorrect recovery is observable.

CLI

Use uv run afl --help or command-specific --help.

uv run afl init my-scenarios
uv run afl schema --output schemas/scenario.schema.json
uv run afl validate scenarios/helpdesk/lost-create-response.yaml
uv run afl campaign campaigns/offline.yaml --output .afl/campaign
uv run afl inspect .afl/demo/naive/run.json
uv run afl export .afl/demo/naive/run.json --format html --output .afl/report.html

Exit codes: 0 success/pass (or verified replay), 1 observed assertion failure/regression, 2 invalid configuration/corruption/harness error, 3 inconclusive execution or missing required coverage. Campaign summaries retain negative outcomes and show numerators, denominators, and raw trials.

Connect an existing agent

The optional integrations are available locally. Their standard tests use synthetic data and fake provider responses.

Local inspector

Use Node.js 22.16 or a compatible newer release to build the frontend:

uv sync --locked --all-extras
cd web
npm ci
npm run build
cd ..
uv run afl serve

Open the local inspector to launch scenarios, inspect history, compare runs, open violation evidence, examine observation/effect pairs and state changes, and export reports. Run diagnostics explain recorded failures, limits and incomplete coverage. Registered MCP client usage appears separately from the server's unknown usage when its evidence passes consistency checks. All displayed outcomes come from executed runs. Local API and UI details.

AgentFaultLab inspector showing two executed scripted fixtures

The screenshot shows actual local fixture runs: a duplicate-producing retry fails, while the stable-key controller passes. These are synthetic acceptance results.

Optional model adapters and MCP

The openai adapter uses the official SDK and requires explicit live opt-in, a chosen model and a process-local API key. Offline protocol tests exercise the real engine. No live OpenAI call has been executed. Configure the adapter.

The optional together adapter uses Chat Completions with a separate Together key and model selection. The 2026-09-26 validation retained 28 actual trials: five smokes across three models and two prompt versions, an 18-trial Qwen/Qwen3.5-9B campaign (13 pass, 2 fail, 3 inconclusive), two passing Qwen MCP runs, and three follow-up MCP trials for stale reads. The lab detected both duplicate-ticket failures. A separate, explicit readback-before-notify-v1 policy passed its no-fault control and stale-read trial: Qwen received an older unassigned ticket, encountered a version conflict, read the current assignment and notified once. The unchanged base-policy follow-up still omitted the read and remains inconclusive, as do the original three stale cases. These are small, configuration-specific observations; results from different models, prompts, policies, profiles and transports are not pooled into a model ranking. Together configuration and all validation attempts and evidence record the actual results and limits.

uv run --extra mcp afl mcp scenarios/helpdesk/lost-create-response.yaml --output .afl/mcp

This serves one run over MCP stdio using the official v2 SDK. Clients can read the public context and call the six simulated tools. finish_run permanently closes the run before returning evaluation. MCP client configuration and boundaries.

Trusted Python adapter

An adapter receives only the public task, actor capabilities and ordinary tool catalog. Route its calls through await context.call_tool(name, arguments) and backoff through await context.sleep(ms). The agent never receives the fault schedule, scenario filename, private world or grading rules.

from agent_fault_lab.adapters import AgentContext, AgentResult


async def run(context: AgentContext) -> AgentResult:
    observation = await context.call_tool(
        "get_customer", {"customer_id": context.task["customer_id"]}
    )
    # Feed the observation into your own reasoning loop and route further tool calls.
    return AgentResult(output=f"Customer lookup succeeded: {observation.ok}")

This small example only reads a customer; it cannot pass the default effectful task. Execute your own installed module:callable only with --trust-local-agent. Trusted Python adapters run in a supervised child process, while the parent owns world state. This enforces deadlines and closes late effects; it is not a hostile-code sandbox. Unrouted tools and external actions are outside the test.

Development and evidence

uv sync --locked --all-extras
uv run ruff check .
uv run ruff format --check .
uv run mypy src/agent_fault_lab
uv run pytest -m "not live"
uv run afl demo --output .afl/verification-demo

Frontend checks run in web/: npm run typecheck, npm test, npm run build, and npm run test:e2e. Install the browser once with npx playwright install chromium (on Linux CI, use --with-deps). Browser tests start an isolated local backend and exercise real stored evidence.

See PROGRESS.md for executed checks and remaining work. See the implementation specification, scenario format, tool contracts, replay semantics, measurements, threat model, and contribution guide.

Only sequential helpdesk workflows are modeled. No distributed races, production traffic interception, automatic OpenAPI business inference, or universal reliability guarantee is claimed. Related work is documented here. AgentFaultLab is the chosen project name; package, trademark and domain availability are not asserted.

Author and license

Abdulvahap Ogut (@vahapogut). Original code is licensed under Apache-2.0. Third-party dependencies retain their respective licenses.

About

Local fault-injection lab for tool-using AI agents. Test timeouts, duplicate delivery, stale reads and recovery with deterministic assertions, replayable evidence, MCP and a local inspector.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages