A controller that plans, selects a tool, observes, and decides whether the evidence is adequate — with the four production controls that separate an agent from an unbounded loop: budgets, typed tool schemas, least privilege, and explicit terminal states.
Built around the principle that agency must earn its cost against a measured baseline.
Standard library only, fully deterministic.
python -m agentic evaluate # agent vs single-pass baseline
python -m agentic budget # what each budget dimension does when it binds
python -m agentic multihop --steps # full trajectory for one question
python -m unittest discover -v # 30 testsflowchart TD
Q["Question + caller scopes"] --> G{"Ambiguous?"}
G -->|yes| CL["clarify → needs_clarification"]
G -->|no| P["Plan next action"]
P --> B{"Budget allows<br/>this call?"}
B -->|no| AB["abstained"]
B -->|yes| CH["Charge cost + latency<br/><b>before</b> the call"]
CH --> T{"Which tool?"}
T -->|"named asset"| SQL["sql_lookup<br/><i>read-only, scoped</i>"]
T -->|"relationship"| KG["graph_query<br/><i>typed, directional</i>"]
T -->|"narrative"| VEC["vector_search"]
SQL --> O["Observe evidence"]
KG --> O
VEC --> O
SQL -->|refused| ESC["escalated"]
O --> S{"Evidence<br/>adequate?"}
S -->|no, budget remains| P
S -->|yes| A["answered + citations"]
S -->|no, budget spent| AB
classDef ok fill:#bbf7d0,stroke:#15803d,color:#1f2937
classDef bad fill:#fecaca,stroke:#b91c1c,color:#1f2937
classDef warn fill:#fde68a,stroke:#b45309,color:#1f2937
class A ok
class AB,ESC bad
class CL,B warn
Every run ends in one of five declared states — answered, needs_clarification,
abstained, escalated, failed — and a test asserts the benchmark never produces
anything else. "It kept going" is not a terminal state.
{
"tasks": 5,
"agent_success": 1.0,
"baseline_success": 0.2,
"marginal_value_of_agency": 0.8,
"cost_multiple": 0.82,
"tool_selection_accuracy": 1.0,
"trajectory_accuracy": 1.0,
"termination_accuracy": 1.0
}
| Question | Status | Tools actually run |
|---|---|---|
| How often should filter E-217 be replaced? | answered |
vector_search |
| Which part supersedes E-217? | answered |
graph_query, vector_search |
| What firmware is on C7-114 and how often is its filter replaced? | answered |
sql_lookup, vector_search |
| What do technicians inspect on this compressor? | needs_clarification |
clarify |
| What firmware is on asset C7-114? (unscoped caller) | escalated |
— |
Two honest readings of this table:
- On row 1 agency added nothing. The baseline answers it too, at one tool call. A test asserts that explicitly. A well-filtered lookup is faster, cheaper and easier to audit, and the agent should not be credited for questions the simple path already handles.
cost_multipleof 0.82 is misleading and is left in deliberately. The agent looks cheaper than the baseline only because two of five tasks terminate early without answering (clarify costs nothing, the refused call never runs). Averaging cost across runs that did different amounts of work is exactly the kind of misleading aggregate that operational metrics produce — the number needs slicing by terminal status before it means anything.
Four independent ceilings, because each fails for a different reason:
graph LR
subgraph B["Budget dimensions"]
S["max_steps<br/>planning loops"]
T["max_tool_calls<br/>call count"]
C["max_cost_units<br/>spend"]
L["max_latency_ms<br/>wall clock"]
end
S --> R["Whichever binds first<br/>→ abstained"]
T --> R
C --> R
L --> R
Capping steps alone still permits three calls to an expensive tool; capping cost alone permits a hundred cheap calls that never converge. Cost and latency are charged before the call is made, so an over-budget call is never issued — a test asserts the counter is unchanged after a rejected charge.
generous status=answered traj=['sql_lookup', 'vector_search']
one tool call status=abstained traj=['sql_lookup']
tiny cost cap status=abstained traj=[]
two steps status=answered traj=['sql_lookup', 'vector_search']
Tools carry typed argument schemas and validate both directions — missing and unexpected
arguments are rejected. sql_lookup is a parameterised read-only projection, not SQL
access: the agent never composes a query string.
Authorization lives in the backend, not in the controller. sql_lookup raises
ToolError for a caller without the support or engineering scope, and the run
terminates escalated rather than answering from whatever else it happened to find. An
agent that filtered after the fact would have already read the row.
read_only is enforced at call time — anything else requires approval, which is what stops
an autonomous loop from taking a consequential action nobody sanctioned.
| Bug | Why it mattered |
|---|---|
Cues were matched against stemmed terms but written unstemmed — supersedes stems to supersed, which does not start with supersede |
The relation cue never fired; every relationship question silently fell through to vector search |
replac was a supersession cue |
"How often is the filter replaced" — a maintenance-interval question — was routed to the parts graph |
| Trajectory counted tools that were skipped for budget | The trace claimed the agent used sql_lookup and graph_query under a one-call budget. Trajectory now reflects only executed calls, so evaluation cannot credit work that never happened |
The planner is rule-based, not an LLM. _plan inspects the question for an asset id, a
relation cue, and what has already been observed. This is a deliberate scope choice: what
the project demonstrates is the control surface around an agent — budgets, typed schemas,
least privilege, terminal states, trajectory evaluation — and none of that gets easier to
test by making the planner nondeterministic. _plan is the single swap point for a real
LLM planner; everything around it stays.
The consequence is honest too: a rule-based planner scores 1.0 on a five-task benchmark it was written against. That number measures the harness, not planning ability. A real evaluation needs a held-out task set the planner has not seen, an LLM planner, and adversarial tasks where the correct behaviour is to stop.