Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Agentic RAG

A controller that plans, selects a tool, observes, and decides whether the evidence is adequate — with the four production controls that separate an agent from an unbounded loop: budgets, typed tool schemas, least privilege, and explicit terminal states.

Built around the principle that agency must earn its cost against a measured baseline.

Standard library only, fully deterministic.

python -m agentic evaluate            # agent vs single-pass baseline
python -m agentic budget              # what each budget dimension does when it binds
python -m agentic multihop --steps    # full trajectory for one question
python -m unittest discover -v        # 30 tests

Control loop

flowchart TD
    Q["Question + caller scopes"] --> G{"Ambiguous?"}
    G -->|yes| CL["clarify → needs_clarification"]
    G -->|no| P["Plan next action"]

    P --> B{"Budget allows<br/>this call?"}
    B -->|no| AB["abstained"]
    B -->|yes| CH["Charge cost + latency<br/><b>before</b> the call"]

    CH --> T{"Which tool?"}
    T -->|"named asset"| SQL["sql_lookup<br/><i>read-only, scoped</i>"]
    T -->|"relationship"| KG["graph_query<br/><i>typed, directional</i>"]
    T -->|"narrative"| VEC["vector_search"]

    SQL --> O["Observe evidence"]
    KG --> O
    VEC --> O
    SQL -->|refused| ESC["escalated"]

    O --> S{"Evidence<br/>adequate?"}
    S -->|no, budget remains| P
    S -->|yes| A["answered + citations"]
    S -->|no, budget spent| AB

    classDef ok fill:#bbf7d0,stroke:#15803d,color:#1f2937
    classDef bad fill:#fecaca,stroke:#b91c1c,color:#1f2937
    classDef warn fill:#fde68a,stroke:#b45309,color:#1f2937
    class A ok
    class AB,ESC bad
    class CL,B warn
Loading

Every run ends in one of five declared states — answered, needs_clarification, abstained, escalated, failed — and a test asserts the benchmark never produces anything else. "It kept going" is not a terminal state.

Does agency earn its cost?

{
  "tasks": 5,
  "agent_success": 1.0,
  "baseline_success": 0.2,
  "marginal_value_of_agency": 0.8,
  "cost_multiple": 0.82,
  "tool_selection_accuracy": 1.0,
  "trajectory_accuracy": 1.0,
  "termination_accuracy": 1.0
}
Question Status Tools actually run
How often should filter E-217 be replaced? answered vector_search
Which part supersedes E-217? answered graph_query, vector_search
What firmware is on C7-114 and how often is its filter replaced? answered sql_lookup, vector_search
What do technicians inspect on this compressor? needs_clarification clarify
What firmware is on asset C7-114? (unscoped caller) escalated —

Two honest readings of this table:

  • On row 1 agency added nothing. The baseline answers it too, at one tool call. A test asserts that explicitly. A well-filtered lookup is faster, cheaper and easier to audit, and the agent should not be credited for questions the simple path already handles.
  • cost_multiple of 0.82 is misleading and is left in deliberately. The agent looks cheaper than the baseline only because two of five tasks terminate early without answering (clarify costs nothing, the refused call never runs). Averaging cost across runs that did different amounts of work is exactly the kind of misleading aggregate that operational metrics produce — the number needs slicing by terminal status before it means anything.

Budgets

Four independent ceilings, because each fails for a different reason:

graph LR
    subgraph B["Budget dimensions"]
        S["max_steps<br/>planning loops"]
        T["max_tool_calls<br/>call count"]
        C["max_cost_units<br/>spend"]
        L["max_latency_ms<br/>wall clock"]
    end
    S --> R["Whichever binds first<br/>→ abstained"]
    T --> R
    C --> R
    L --> R
Loading

Capping steps alone still permits three calls to an expensive tool; capping cost alone permits a hundred cheap calls that never converge. Cost and latency are charged before the call is made, so an over-budget call is never issued — a test asserts the counter is unchanged after a rejected charge.

generous         status=answered     traj=['sql_lookup', 'vector_search']
one tool call    status=abstained    traj=['sql_lookup']
tiny cost cap    status=abstained    traj=[]
two steps        status=answered     traj=['sql_lookup', 'vector_search']

Least privilege

Tools carry typed argument schemas and validate both directions — missing and unexpected arguments are rejected. sql_lookup is a parameterised read-only projection, not SQL access: the agent never composes a query string.

Authorization lives in the backend, not in the controller. sql_lookup raises ToolError for a caller without the support or engineering scope, and the run terminates escalated rather than answering from whatever else it happened to find. An agent that filtered after the fact would have already read the row.

read_only is enforced at call time — anything else requires approval, which is what stops an autonomous loop from taking a consequential action nobody sanctioned.

Three bugs the tests caught

Bug Why it mattered
Cues were matched against stemmed terms but written unstemmed — supersedes stems to supersed, which does not start with supersede The relation cue never fired; every relationship question silently fell through to vector search
replac was a supersession cue "How often is the filter replaced" — a maintenance-interval question — was routed to the parts graph
Trajectory counted tools that were skipped for budget The trace claimed the agent used sql_lookup and graph_query under a one-call budget. Trajectory now reflects only executed calls, so evaluation cannot credit work that never happened

Honest scope

The planner is rule-based, not an LLM. _plan inspects the question for an asset id, a relation cue, and what has already been observed. This is a deliberate scope choice: what the project demonstrates is the control surface around an agent — budgets, typed schemas, least privilege, terminal states, trajectory evaluation — and none of that gets easier to test by making the planner nondeterministic. _plan is the single swap point for a real LLM planner; everything around it stays.

The consequence is honest too: a rule-based planner scores 1.0 on a five-task benchmark it was written against. That number measures the harness, not planning ability. A real evaluation needs a held-out task set the planner has not seen, an LLM planner, and adversarial tasks where the correct behaviour is to stop.

About

Bounded agentic RAG controller with tool budgets, least privilege and explicit terminal states

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages