Skip to content

fix(llminternal): re-evaluate toolsets on every step in toolProcessor - #785

Open
nuthalapativarun wants to merge 3 commits into
google:mainfrom
nuthalapativarun:fix/757-toolset-re-evaluate-per-step
Open

nuthalapativarun wants to merge 3 commits into
google:mainfrom
nuthalapativarun:fix/757-toolset-re-evaluate-per-step

Conversation

@nuthalapativarun

Copy link
Copy Markdown
Contributor

Link to Issue or Description of Change

Description

Problem:
toolProcessor in internal/llminternal/tools_processor.go cached f.Tools after the first call and short-circuited with if f.Tools != nil { return } on every subsequent step. Since a Flow is created once per Runner.Run() and reused across all runOneStep() iterations, Toolset.Tools(ctx) was only ever called on the first step of a run.

Any toolset whose Tools(ctx) output depends on session state modified by an earlier tool call could not surface new tools within the same Runner.Run(). The ctx parameter on Tools(ctx) implies dynamic per-step evaluation, but the caching made it effectively static for the duration of the run.

Root cause:

// tools_processor.go — before
func toolProcessor(...) iter.Seq2[*session.Event, error] {
    return func(yield func(*session.Event, error) bool) {
        if f.Tools != nil {   // ← cached after first call; never re-evaluated
            return
        }
        ...
        f.Tools = tools
    }
}

Solution:
Remove the if f.Tools != nil { return } guard so Toolset.Tools() is called before every model step. Static Tools from the agent config and dynamic toolset tools are rebuilt into f.Tools on each call, which is correct because runOneStep creates a fresh *model.LLMRequest each iteration anyway.

// tools_processor.go — after
func toolProcessor(...) iter.Seq2[*session.Event, error] {
    return func(yield func(*session.Event, error) bool) {
        // No cache guard — toolsets re-evaluated every step.
        ...
        f.Tools = tools
    }
}

Testing Plan

Unit Tests:

  • Added TestToolProcessorReEvaluatesToolsetsEachStep in internal/llminternal/base_flow_test.go with a dynamicToolset that returns no tools on the first call and one tool on the second call — verifying that toolProcessor picks up the change on the next step.
  • All existing unit tests pass locally.
$ go test ./internal/llminternal/... -run TestToolProcessorReEvaluatesToolsetsEachStep -v
=== RUN   TestToolProcessorReEvaluatesToolsetsEachStep
--- PASS: TestToolProcessorReEvaluatesToolsetsEachStep (0.00s)
PASS
ok  	google.golang.org/adk/internal/llminternal	0.392s

$ go test ./internal/llminternal/...
ok  	google.golang.org/adk/internal/llminternal	0.909s
ok  	google.golang.org/adk/internal/llminternal/googlellm	0.196s

Manual End-to-End (E2E) Tests:

The bug is deterministic: any agent with a Toolset that uses session state to control which tools are returned will reproduce it. The activate_vehicle_tools example from issue #757 exercises the exact path — activate_vehicle_tools sets state, and conditionalToolset.Tools(ctx) checks that state to surface a vehicle tool. After this fix, the activated tools appear on the next runOneStep() within the same Runner.Run().

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules.

@nuthalapativarun

Copy link
Copy Markdown
Contributor Author

Hi, following up on this PR. Happy to make any changes needed — let me know if there's anything blocking review. Thanks!

@nuthalapativarun

Copy link
Copy Markdown
Contributor Author

Hi, just following up on this PR. Happy to make any adjustments needed — let me know if there's anything blocking review. Thanks!

@nuthalapativarun
nuthalapativarun force-pushed the fix/757-toolset-re-evaluate-per-step branch from 752ec36 to 94e1371 Compare May 17, 2026 16:02
@nuthalapativarun

Copy link
Copy Markdown
Contributor Author

Hi, just following up on this PR. Happy to make any adjustments needed — let me know if there's anything blocking review. Thanks!

@nuthalapativarun
nuthalapativarun force-pushed the fix/757-toolset-re-evaluate-per-step branch from 94e1371 to 7ab69bf Compare June 7, 2026 01:08
@nuthalapativarun

Copy link
Copy Markdown
Contributor Author

Hi, just following up on this PR. Happy to make any adjustments if something needs to change — let me know if there's anything blocking review. Thanks!

@nuthalapativarun
nuthalapativarun force-pushed the fix/757-toolset-re-evaluate-per-step branch from 7ab69bf to c9b98da Compare July 25, 2026 04:32

@karolpiotrowicz karolpiotrowicz left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The change itself is right, and it matches how adk-python has always behaved. _process_agent_tools there re-resolves every tool union, toolsets included, on each step (base_llm_flow.py:458, awaited from _preprocess_async at line 1435, inside the step loop), and the comment at lines 540-541 states the reason outright: "Tool sets can change between model steps, so the cache is refreshed each time." Driving a state-gated toolset through runner.Runner.Run, step 2's tool list is [activate] on main and [activate search_vehicles] with this applied, so the fix works at the public entry point. Sorry this sat unreviewed for so long.

Two things before it can merge.

The test passes on main, so it does not pin the fix

The fixture at base_flow_test.go:804 is &State{Toolsets: []tool.Toolset{ts}}, which leaves State.Tools nil, and toolsByCall[0] is nil too. Appending zero elements to a nil slice returns nil, so f.Tools is still nil after the first call and the guard you removed — if f.Tools != nil { return } — would never have fired here. I put the merged test file onto an otherwise unmodified main and it passes:

$ go test -run TestToolProcessorReEvaluatesToolsetsEachStep ./internal/llminternal/
--- PASS: TestToolProcessorReEvaluatesToolsetsEachStep (0.00s)
ok  google.golang.org/adk/v2/internal/llminternal  0.117s

Giving the agent one static tool makes f.Tools non-nil after the first call, which is what the guard needs in order to trip. With this the test fails on main and passes with your change, and ./internal/llminternal/... stays green:

-	agentState := &State{Toolsets: []tool.Toolset{ts}}
+	baseTool := &mockFunctionTool{name: "base_tool"}
+	agentState := &State{Tools: []tool.Tool{baseTool}, Toolsets: []tool.Toolset{ts}}
@@
-	if len(f.Tools) != 0 {
-		t.Errorf("after call 1: got %d tools, want 0", len(f.Tools))
+	if len(f.Tools) != 1 {
+		t.Errorf("after call 1: got %d tools, want 1", len(f.Tools))
 	}
@@
-	if len(f.Tools) != 1 || f.Tools[0].Name() != extraTool.Name() {
+	if len(f.Tools) != 2 || f.Tools[1].Name() != extraTool.Name() {

Rebase

The branch is 104 commits behind and conflicts in internal/llminternal/base_flow_test.go. It is only that both sides append tests at the end of the file, so keeping both blocks resolves it, but it does need doing.

Non-blocking, worth knowing

  • Toolsets that do I/O now pay per step. Toolset.Tools() goes from one call per run to one per model call — 11 calls for an 11-model-call run, measured. That matters for mcptoolset, where every call is a live paginated ListTools with no caching, and the loop at tools_processor.go:39 walks toolsets serially where Python runs them under asyncio.gather specifically to overlap those listings. Nothing to change in this PR — caching belongs in mcptoolset — but it is a consequence someone will hit.
  • A toolset that fails mid-run now aborts the invocation. The error return is unchanged, but it was previously only reachable before the first model call. A toolset erroring on its third evaluation used to leave the run to finish, and now ends it with earlier steps' side effects already persisted. The same holds for a toolset that starts returning a name already taken by a static tool, which surfaces as duplicate tool: "...". Python fails closed the same way, so this looks like the intended semantics rather than something to fix — it is just worth a line in the description so it is not a surprise.
  • The comment overreaches slightly on the live path. tools_processor.go:27-30 says the list "must be rebuilt before each model call", but RunLive calls preprocess once at base_flow.go:336, outside the reconnect loop, so toolsets are resolved once per live session both before and after this change. Issue #757's symptom survives there. That is a separate gap, not something this PR needs to close — the comment just shouldn't imply it already has. The comment it replaced described ContentRequestProcessor, so this is still a clear improvement.

toolProcessor cached f.Tools after the first call and returned early on
subsequent steps within the same Flow.Run(). This prevented toolsets
whose Tools() output depends on session state from updating their tool
list after a tool call modified that state earlier in the same run.

Remove the f.Tools guard so toolsets are re-evaluated before every
model call, matching the intent of the Toolset.Tools(ctx) signature.

Fixes google#757
@nuthalapativarun
nuthalapativarun force-pushed the fix/757-toolset-re-evaluate-per-step branch from c9b98da to e8054d2 Compare September 13, 2026 21:04
@nuthalapativarun

Copy link
Copy Markdown
Contributor Author

Rebased onto current main — resolved a straightforward conflict in internal/llminternal/base_flow_test.go where main had independently appended a different set of tests (thought-only-turn termination/reset tests and the partial-last-event regression test) right after the same insertion point this PR's TestToolProcessorReEvaluatesToolsetsEachStep/dynamicToolset additions land at; both sets of tests are now present, no logic changes needed. go build ./... and go test ./internal/llminternal/... are green. Force-pushed the rebase. This PR hasn't had a first review yet — would appreciate a look when you have a chance. Thanks!

@karolpiotrowicz karolpiotrowicz left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One thing still blocks: the new test passes without the fix, so it would not catch the cache guard coming back. Before merge, the test needs to fail on main and pass with your change. The fixture-line comment has a change that does this. The full reasoning is in my earlier review, which I suspect was easy to miss, since your last comment mentions no review yet.

The branch is behind main again, with no conflicts at the moment.

Comment thread internal/llminternal/base_flow_test.go Outdated
Resolves the conflict in internal/llminternal/base_flow_test.go, where
both sides appended tests at the end of the file. Both blocks are kept.

Also makes TestToolProcessorReEvaluatesToolsetsEachStep fail without the
fix, as requested in review: the agent now has one static tool, so
f.Tools is non-nil after the first call and the removed cache guard
would have returned early. Verified red with main's tools_processor.go
and green with this change.
@wolo-lab

wolo-lab commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

@karolpiotrowicz I pushed a commit to this branch (caf6660) that merges current main and applies your fixture change. The conflict in base_flow_test.go was only both sides appending tests, so both blocks are kept.

With the static base_tool in place, the test now fails against the tools_processor.go from main:

--- FAIL: TestToolProcessorReEvaluatesToolsetsEachStep
    base_flow_test.go:1287: toolset.Tools() called 1 times, want 2

It passes with the change in this PR, and CI is green on the new head. I also updated the first-call comment and the second-call failure message, since both still described the old fixture with no static tool.

@nuthalapativarun heads-up that your branch moved, in case you have local changes on it.

@wolo-lab wolo-lab left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

wolo-lab pushed a commit to wolo-lab/adk-go that referenced this pull request Oct 3, 2026
…google#785)

toolProcessor cached f.Tools after the first call and returned early on
subsequent steps within the same Flow.Run(). This prevented toolsets
whose Tools() output depends on session state from updating their tool
list after a tool call modified that state earlier in the same run.

Remove the f.Tools guard so toolsets are re-evaluated before every
model call.

Fixes google#757
wolo-lab added a commit to wolo-lab/adk-go that referenced this pull request Oct 3, 2026
Experiments behind the review of the native lazy tool catalog proposal:

- turnproof: a discovered tool is absent from the generation that found
  it, so model-triggered search costs one extra model call.
- dupproof: packing a tool that Tools() already returned aborts the run
  with a duplicate-tool error.
- lazyclean: a catalog written purely through Toolset.Tools(ctx) only
  works if toolsets are re-evaluated per model step (google#757).
- presearch: searching from UserContent before the first generation
  removes the extra model call when the search hits.
- rank: recall of five rankers on long user messages over a 68-tool
  catalog, with and without ARD-style representativeQueries.

lazyclean and presearch depend on google#785 and fail without it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Toolset.Tools(ctx) evaluated once per Runner.Run(), not per model step — state-driven toolsets broken

3 participants