Find paths that bypass your AI agent safety controls.
You added an approval before production. This scanner checks whether another modeled path can reach the same production consequence without crossing the boundary you expected.
EXPECTED
AI agent
-> approval:P
-> production deploy ✅
COUNTEREXAMPLE
AI agent
-> shell
-> gh workflow run deploy.yml
-> production deploy ❌
Missing boundary: approval:P
Install from a checkout of this repository:
python -m pip install .Then check the production-deploy boundary already modeled by the MVP:
consequence-boundary . --consequence production_deploy --boundary PWhen a certain bypass path exists, the command reports
COUNTEREXAMPLE_FOUND, returns the path plus source evidence, and exits
with code 1 so the same check can become a CI gate.
Example shape:
{
"status": "COUNTEREXAMPLE_FOUND",
"consequence": "production_deploy",
"expected_boundary": "P",
"path": [
"root:agent",
"effect:shell.exec",
"workflow:deploy.yml",
"consequence:production_deploy"
]
}This is not another authorization engine. It tests a different question: after you add an approval or policy boundary, is there another modeled route to the same real-world consequence that bypasses it?
Current scope is deliberately narrow: Python + GitHub Actions consequence
paths, with conservative outcomes. COVERED_WITHIN_MODEL is not a claim of
repository-wide completeness, and UNKNOWN is never a safety guarantee.
Please open a GitHub issue and include only:
- the returned status:
COUNTEREXAMPLE_FOUND,COVERED_WITHIN_MODEL, orUNKNOWN; - whether the reported path was real or a false positive;
- whether this check would be useful enough to keep in CI.
Do not include private source code, credentials, or proprietary repository details.
The repository also contains agent-action-guard, a deterministic local
policy gate for proposed AI-agent actions. It evaluates one action proposal
against one policy document and emits one canonical JSON decision record.
It evaluates proposals only. It never executes actions.
Exit code 0 means a decision record was produced — not that permission was granted. The JSON
decisionfield is the only authority. Never writeagent-action-guard && do-the-thing.
Systems that let AI agents act need an auditable answer to the question "what is this agent allowed to do, and what evidence supports that decision?" before anything happens. That answer must be boring: the same inputs must produce the same bytes, no matter how the policy file is ordered, whitespaced, or escaped, and anything malformed must fail closed rather than fail open.
v0.1 is deliberately small: a single offline evaluation of one action JSON file against one policy JSON file, producing one decision record on stdout. The complete semantics are fixed by ADR-0001, the sole normative authority for this version (rules N1–N39). If anything in this README appears to disagree with ADR-0001, ADR-0001 wins.
- Strict fail-closed validation — unknown fields, duplicate JSON members, duplicate rule IDs, wrong types, empty strings, non-UTF-8 bytes, a UTF-8 BOM, and unpaired surrogates all reject the input (exit 3) instead of being silently tolerated.
- Exact code-point matching — case-sensitive, no trimming, no case
folding, no Unicode normalization, no regex, no globs, no coercion.
Omitted match members are unconstrained;
"side_effect": nullpresent in a match is distinct from omitting it. - Effect precedence — DENY > REQUIRE_APPROVAL > eligible ALLOW, then default DENY. No first-match ordering exists.
- Side-effect eligibility safeguard — an ALLOW rule that omits
side_effectcan never authorize an action whoseside_effectistrueornull; authors must write the value explicitly to authorize side-effecting or unknown-side-effect actions. - Rule-order independence — reordering the policy's
rulesarray changes neither the decision nor a single output byte. - Canonical byte-reproducible JSON — compact separators, fixed member order, code-point-sorted rule-ID groups, direct UTF-8 (no ASCII escaping of non-ASCII), no BOM, exactly one trailing LF.
- No network, no database, no LLM, no persistence, no execution — the only outputs are the stdout record, stderr diagnostics, and the exit code.
Runs on the Python 3.10+ standard library. No dependencies.
From the repository root:
python3 -m agent_action_guard evaluate --action action.json --policy policy.json{
"action_id": "act-2026-001",
"actor": "agent-1",
"tool": "fs",
"operation": "read",
"resource": "/data/report.txt",
"side_effect": false
}All six fields are required; side_effect must be exactly true,
false, or null.
{
"policy_id": "policy-baseline",
"rules": [
{
"rule_id": "allow-fs-read",
"effect": "ALLOW",
"match": {
"tool": "fs",
"operation": "read",
"side_effect": false
}
}
]
}A match object may constrain any subset of actor, tool,
operation, resource, and side_effect; an empty match object
field-matches every action.
One compact line on stdout, terminated by a single LF (shown wrapped here for readability only — the real output has no inner whitespace or line breaks):
{"action_id":"act-2026-001","policy_id":"policy-baseline","decision":"ALLOW","matched_rule_ids":{"DENY":[],"REQUIRE_APPROVAL":[],"ALLOW":["allow-fs-read"]},"decisive_reason":"matched_eligible_allow","default_deny_used":false}decision |
Meaning |
|---|---|
ALLOW |
An eligible ALLOW rule field-matched and nothing stronger did. |
DENY |
A DENY rule field-matched, or the default deny applied. |
REQUIRE_APPROVAL |
A REQUIRE_APPROVAL rule field-matched and no DENY did. |
decisive_reason |
default_deny_used |
Meaning |
|---|---|---|
matched_deny |
false |
At least one DENY rule field-matched. |
matched_require_approval |
false |
No DENY matched; a REQUIRE_APPROVAL rule did. |
matched_eligible_allow |
false |
No DENY or REQUIRE_APPROVAL matched; an eligible ALLOW did. |
default_deny_no_match |
true |
No rule field-matched at all. |
default_deny_allow_ineligible |
true |
Only ALLOW rules field-matched, and none was eligible. |
default_deny_used is fully determined by decisive_reason: it is
true exactly for the two default_deny_* codes.
matched_rule_ids always contains all three groups (DENY,
REQUIRE_APPROVAL, ALLOW), each a possibly-empty array of every rule
ID that field-matched with that effect, sorted by Unicode code point.
Ineligible ALLOW rules stay listed in the ALLOW group; eligibility
affects only decision and decisive_reason.
| Code | Meaning |
|---|---|
| 0 | A complete decision record was written to stdout — for ALLOW, DENY, and REQUIRE_APPROVAL alike. Not permission. |
| 1 | Unexpected internal error. |
| 2 | CLI argument or usage error only. |
| 3 | File access, decoding, JSON syntax, or input validation error — missing, unreadable, non-UTF-8, BOM-prefixed, malformed, or schema-invalid files. |
On every non-zero exit, stdout is exactly empty; diagnostics go to stderr and their wording is not a stable contract.
This branch also contains the first usable MVP of the Consequence Boundary Completeness work.
It answers a narrower question than the policy evaluator above:
From this repository entrypoint, can the current model trace a path to this GitHub workflow consequence?
Install the repository as a local CLI package:
python -m pip install .Then check the existing consequence-boundary property:
consequence-boundary . --consequence production_deploy --boundary PA proven bypass exits with code 1 and reports
COUNTEREXAMPLE_FOUND, including the alternate path and its evidence.
COVERED_WITHIN_MODEL is deliberately not a repository-wide safety
guarantee.
For lower-level root-to-target path tracing, the experimental path CLI remains available:
python -m retest_lab scan --repo . --root main --target deploy.ymlThe function: and workflow: prefixes are optional. The equivalent
explicit form is:
python -m retest_lab scan \
--repo . \
--root function:main \
--target workflow:deploy.ymlExample output:
PROVEN
root: function:main
target: workflow:deploy.yml
path:
function:main
-> effect:shell.exec [invokes; certain; app.py:42]
-> workflow:deploy.yml [gh_workflow_dispatch; certain; app.py:42]
Machine-readable output is available with --json:
python -m retest_lab scan \
--repo . \
--root main \
--target deploy.yml \
--jsonThe scanner returns one of three statuses:
| Status | Meaning |
|---|---|
PROVEN |
A modeled path from the requested root to target exists and every edge on that path is supported as certain by the bounded extractor. |
POSSIBLE |
A modeled path exists, but at least one edge depends on unresolved runtime state or incomplete static information. |
UNKNOWN |
The current model did not find a path. This does not mean the path is impossible. |
Important: PROVEN is a statement about the returned modeled path, not a
claim that the repository has been analyzed completely. POSSIBLE must never
be promoted to a proven bypass, and UNKNOWN must never be interpreted as a
safety guarantee.
The current pinned real-world acceptance corpus contains 25 cases: 15 known-positive paths and 10 negative controls. The current gate requires zero missed positives and zero certain or possible false positives before CI passes.
Run the full test suite from the repository root:
python3 -m unittest discover -s tests -t . -vThe suite currently contains 113 tests covering the model, the strict loader, the pure decision engine, the canonical serializer, and the CLI contract — including golden-byte and shuffle-invariance checks.
- ADR-0001: Scope and decision semantics — the sole normative authority (N1–N39).
- Problem statement
- Implementation plan
- No execution, retrying, persistence, or transmission of any action.
- No regex, glob, or expression-language matching.
- No policy composition, batch evaluation, or stdin input.
- No
schema_versionfield — ADR-0001 is the version authority; any semantic change requires a superseding ADR. - No per-decision exit codes: exit status signals evaluation success only, never the verdict.
- No hosted service, account system, billing, or remote execution.